EDBT 2026 Demo / reviewers in the wild / expert
Nathaniel Bleier
dblp:270/3843 · also Nathan Bleier
· DBLP profile ↗
13ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0001-7791-2065ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 8 first-author · 11 since 2021Software engineering, systems software and programming languages · 8 · 5 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Circuit-Aware Analysis of Arithmetic Error Detection CodesabstractIn modern CMOS VLSI circuits, arithmetic datapaths are increasingly vulnerable to radiation-induced soft errors and circuit-aging, leading to application crashes or silent data corruption. However, existing studies lack a gate-level framework for evaluating the efficacy of arithmetic error-detection codes in the presence of logic and timing masking. We focus on arithmetic codes—such as AN codes and Redundant Residue Number System codes—that natively support arithmetic operations, unlike conventional linear block codes. We prioritize detection over correction given the low cost of re-execution and the rarity of radiation strikes relative to compute throughput. Because 48-bit values and pointers typically suffice in contemporary systems, we adopt a 48+16-bit (data+redundancy) organization. We introduce a statistical framework that models radiation strikes by injecting LET-based pulse-widths into netlists and emulates circuit-aging through delay scaling. The toolchain enables designers to screen arithmetic code schemes before tape-out and tailor resilience for deployments ranging from sea-level cosmic-ray exposure to the high-radiation conditions such as particle accelerators, spacecraft, and defense electronics. By exposing how code parameters and circuit topology interact at the gate level, the framework fosters tightly coupled code-hardware co-design. Cheng Chiu, Keyon Mazandarani, Nathaniel Bleier |
DATE | 3 |
| 2026 | æSIP: μArch-Aware ASIP-ISA Co-Design via Program Synthesis, Equality Saturation, and External Don't Cares
Haoran Jin, Jirong Yang, Barry Lyu, Ruijie Gao, Nathaniel Bleier |
ISCA | 5 |
| 2025 | Architecting Space Microdatacenters: A System-level ApproachabstractServer-based computing in space has been recently proposed due to potential benefits in terms of capability, latency, security, sustainability, and cost. Despite this, there has been no work asking the question: how should we architect systems for server-based computing in space when considering overall cost. This paper presents a Total Cost of Ownership (TCO)-based approach to architecture of server-based computing systems for space (Space Microdatacenters - SμDC) for processing data produced by low Earth orbit (LEO)-based Earth observation (EO) satellites. We show that power of compute is the primary factor in determining SμDC TCO, though the dependence is sublinear. Second, the impact of compute mass, monetary cost, and communication on TCO is relatively insignificant. Third, architectures with the highest $\frac{\text { FLOPs }}{{\mathrm {W}}}$ provide much higher performance per TCO ${\$}$ even if they have poor $\frac{\mathrm{FLOPs}}{\$}$. We leverage these insights to advocate extreme heterogeneity designs for SμDCs. These designs reduce SμDC TCO by 116× in spite of poor $\frac{\mathrm{FLOPs}}{\$}$ characteristics. We also show that (a) collaborative compute constellations — constellations in which EO satellites are also equipped with compute hardware — further improve SμDC TCO by 1.31 to 1.74×, (b) a distributed architecture reduces TCO by 10% over a monolithic architecture, and (c) low monetary cost of compute can be leveraged to provide near zero cost compute overprovisioning which improves an SμDC’s availability significantly and supports graceful degradation. Overall, this is the first paper on cost-aware architecture and optimization of a SμDC. Nathaniel Bleier, Rick Eason, Michael Lembeck, Rakesh Kumar 0002 |
HPCA | 1 |
| 2025 | Prompt, Fab, Flex: Agentic LLMs for Flexible Electronics DesignabstractFlexible Electronics (FE) have emerged as a promising platform for extreme edge applications that demand attributes tailored to the application domain, such as ultra-low cost, low power consumption, mechanical flexibility, biocompatibility, and environmental sustainability. While advances in printed and flexible device technologies have demonstrated the feasibility of sensing, computing, and communication on deformable substrates, the design and implementation of FE-based systems remain limited by traditional Electronic Design Automation (EDA) workflows, which are complex, time-intensive, and largely inaccessible to non-experts. In parallel, recent progress in Large Language Models (LLMs) has enabled automation across multiple stages of integrated circuit design; however, existing approaches exclusively target conventional silicon technologies and are not designed to address the unique constraints of FE. This work introduces the first LLM-driven framework for end-to-end hardware design automation in flexible electronics. The proposed methodology supports Register-Transfer Level (RTL) generation, logic synthesis, and cross-layer Power–Performance–Area (PPA) Design Space Exploration (DSE) for bespoke Machine Learning (ML) classifiers. Experimental results demonstrate the feasibility and effectiveness of the approach in generating resource-efficient hardware designs optimized for FE, thereby lowering barriers to adoption and accelerating the development of personalized, application-specific FEs. Farshad Firouzi, Bahareh J. Farahani, Polykarpos Vergos, Deepesh Sahoo, Nathaniel Bleier, Krishnendu Chakrabarty |
ICCAD | 5 |
| 2023 | Exploiting Short Application Lifetimes for Low Cost Hardware Encryption in Flexible ElectronicsabstractMany emerging flexible electronics [1] applications require hardware-based encryption, but it is unclear if practical hardware-based encryption is possible for flexible applications due to stringent power requirements of these applications and high area and power overheads of flexible technologies relative to silicon CMOS technologies. In this work, we observe that the lifetime of many flexible applications is so small that often one key suffices for the entire lifetime. This means that, instead of generating keys and round keys in hardware, we can generate the round keys offline, and instead store these round keys directly on the engine post fabrication in an on-chip programmable read-only memory. This eliminates the need for hardware for dynamic generation of round keys, which significantly reduces encryption overhead, while still allowing engines to have unique keys. This significant reduction in encryption overhead allows us to demonstrate the first practical flexible encryption engines. To prevent an adversary from reading out the stored round keys, we scramble the round keys before storing them in the ROM; camouflage cells are used to unscramble the keys before feeding them to logic. In spite of the unscrambling overhead, our encryption engines consume 27.4% lower power than the already heavily area and power-optimized baselines, while being 21.9% smaller on average. Nathaniel Bleier, Muhammad Husnain Mubarik, Suman Balaji, Francisco Rodriguez, Antony Sou, Rakesh Kumar 0002 |
DATE | 1 |
| 2023 | Programmable Olfactory ComputingabstractWhile smell is arguably the most visceral of senses, olfactory computing has been barely explored in the mainstream. We argue that this is a good time to explore olfactory computing since a) a large number of driver applications are emerging, b) odor sensors are now dramatically better, and c) non-traditional form factors such as sensor, wearable, and xR devices that would be required to support olfactory computing are already getting widespread acceptance. Through a comprehensive review of literature, we identify the key algorithms needed to support a wide variety of olfactory computing tasks. We profiled these algorithms on existing hardware and identified several characteristics, including the preponderance of fixed-point computation, and linear operations, and real arithmetic; a variety of data memory requirements; and opportunities for data-level parallelism. We propose Ahromaa, a heterogeneous architecture for olfactory computing targeting extremely power and energy constrained olfactory computing workloads and evaluate it against baseline architectures of an MCU, a state-of-art CGRA, and an MCU with packed SIMD. Across our algorithms, Ahromaa's operating modes outperform the baseline architectures by 1.36, 1.22, and 1.1× in energy efficiency when operating at MEOP. We also show how careful design of data memory organization can lead to significant energy savings in olfactory computing, due to the limited amount of data memory many olfactory computing kernels require. These improvements to the data memory organization lead to additional 4.21, 4.37, and 2.85× improvements in energy efficiency on average. Nathaniel Bleier, Abigail Wezelis, Lav R. Varshney, Rakesh Kumar 0002 |
ISCA | 1 |
| 2023 | Space MicrodatacentersabstractEarth observation (EO) has been a key task for satellites since the first time a satellite was put into space. The temporal and spatial resolution at which EO satellites take pictures has been increasing to support space-based applications, but this increases the amount of data each satellite generates. We observe that future EO satellites will generate so much data that this data cannot be transmitted to Earth due to the limited capacity of communication that exists between space and Earth. We show that conventional data reduction techniques such as compression [126] and early discard [41] do not solve this problem, nor does a direct enhancement of today’s RF-based infrastructure [133, 153] for space-Earth communication. We explore an unorthodox solution instead - moving to space the computation that would have happened on the ground. This alleviates the need for data transfer to Earth. We analyze ten non-longitudinal RGB and hyperspectral image processing Earth observation applications for their computation and power requirements and discover that these requirements cannot be met by the small satellites that dominate today’s EO missions. We make a case for space microdatacenters - large computational satellites whose primary task is to support in-space computation of EO data. We show that one 4KW space microdatacenter can support the computation need of a majority of applications, especially when used in conjunction with early discard. We do find, however, that communication between EO satellites and space microdatacenters becomes a bottleneck. We propose three space microdatacenter-communication co-design strategies – k − list-based network topology, microdatacenter splitting, and moving space microdatacenters to geostationary orbit – that alleviate the bottlenecks and enable effective usage of space microdatacenters. Nathaniel Bleier, Muhammad Husnain Mubarik, Gary R. Swenson, Rakesh Kumar 0002 |
MICRO | 1 |
| 2022 | FlexiCores: low footprint, high yield, field reprogrammable flexible microprocessorsabstractFlexible electronics is a promising approach to target applications whose computational needs are not met by traditional silicon-based electronics due to their conformality, thinness, or cost requirements. A microprocessor is a critical component for many such applications; however, it is unclear whether it is feasible to build flexible processors at scale (i.e., at high yield), since very few flexible microprocessors have been reported and no yield data or data from multiple chips has been reported. Also, prior manufactured flexible systems were not field-reprogrammable and were evaluated either on a simple set of test vectors or a single program. A working flexible microprocessor chip supporting complex or multiple applications has not been demonstrated. Finally, no prior work performs a design space of flexible microprocessors to optimize area, code size, and energy of such microprocessors. Nathaniel Bleier, Calvin Lee 0004, Francisco Rodriguez, Antony Sou, Rakesh Kumar 0002 |
ISCA | 1 |
| 2022 | Rethinking programmable wearable processorsabstractEarables such as earphones [15, 16, 73], hearing aids [28], and smart glasses [2, 14] are poised to be a prominent programmable computing platform in the future. In this paper, we ask the question: what kind of programmable hardware would be needed to support earable computing in future? To understand hardware requirements, we propose EarBench, a suite of representative emerging earable applications with diverse sensor-based inputs and computation requirements. Our analysis of EarBench applications shows that, on average, there is a 13.54×-3.97× performance gap between the computational needs of EarBench applications and the performance of the microprocessors that several of today's programmable earable SoCs are based on; more complex microprocessors have unacceptable energy efficiency for Earable applications. Our analysis also shows that EarBench applications are dominated by a small number of digital signal processing (DSP) and machine learning (ML)-based kernels that have significant computational similarity. We propose SpEaC --- a coarse-grained reconfigurable spatial architecture - as an energy-efficient programmable processor for earable applications. SpEaC targets earable applications efficiently using a) a reconfigurable fixed-point multiply-and-add augmented reduction tree-based substrate with support for vectorized complex operations that is optimized for the earable ML and DSP kernel code and b) a tightly coupled control core for executing other code (including non-matrix computation, or non-multiply or add operations in the earable DSP kernel code). Unlike other CGRAs that typically target general-purpose computations, SpEaC substrate is optimized for energy-efficient execution of the earable kernels at the expense of generality. Across all our kernels, SpEaC outperforms programmable cores modeled after M4, M7, A53, and HiFi4 DSP by 99.3×, 32.5×, 14.8×, and 9.8× respectively. At 63 mW in 28 nm, the energy efficiency benefits are 1.55 ×, 9.04×, 68.3 ×, and 32.7 × respectively; energy efficiency benefits are 15.7 × -- 1087 × over a low power Mali T628 MP6 GPU. Nathaniel Bleier, Muhammad Husnain Mubarik, Srijan Chakraborty, Shreyas Kishore, Rakesh Kumar 0002 |
ISCA | 1 |
| 2021 | Property-driven Automatic Generation of Reduced-ISA HardwareabstractAs the diversity of computing workloads and customers continues to increase, so does the need to customize hardware at low cost for different computing needs. This work focuses on automatic customization of a given hardware, available as a soft or firm IP, through eliminating unneeded or undesired instruction set architecture (ISA) instructions. We present a property-based framework for automatically generating reduced-ISA hardware. Our framework directly operates on a given arbitrary RTL or gate-level netlist, uses property checking to identify gates that are guaranteed to not toggle if only a reduced ISA needs to be supported, and automatically eliminates these untoggleable gates to generate a new design. We show a 14% gate count reduction when the Ibex [19] core is optimized using our framework for the instructions required by a set of embedded (MiBench) workloads. Reduced-ISA versions generated by our framework that support a limited set of ISA extensions and which cannot be generated using Ibex’s parameterization options provide 10%47% gate count reduction. For an obfuscated Cortex M0 netlist optimized to support the instructions in the MiBench benchmarks, we observe a 20% area reduction and 18% gate count reduction compared to the baseline core, demonstrating applicability of our framework to obfuscated designs. We demonstrate the scalability of our approach by applying our framework to a 100,000-gate RIDECORE [21] design, showing a 14%17% gate count reduction. Nathaniel Bleier, John Sartori, Rakesh Kumar 0002 |
DAC | 1 |
| 2021 | Printed Stochastic Computing Neural NetworksabstractPrinted electronics (PE) offers flexible, extremely low-cost, and on-demand hardware due to its additive manufacturing process, enabling emerging ultra-low-cost applications, including machine learning applications. However, large feature sizes in PE limit the complexity of a machine learning classifier (e.g., a neural network (NN)) in PE. Stochastic computing Neural Networks (SC-NNs) can reduce area in silicon technologies, but still require complex designs due to unique implementation tradeoffs in PE. In this paper, we propose a printed mixed-signal system, which substitutes complex and power-hungry conventional stochastic computing (SC) components by printed analog designs. The printed mixed-signal SC consumes only 35% of power consumption and requires only 25% of area compared to a conventional 4-bit NN implementation. We also show that the proposed mixed-signal SC-NN provides good accuracy for popular neural network classification problems. We consider this work as an important step towards the realization of printed SC-NN hardware for near-sensor-processing. Dennis Weller, Nathaniel Bleier, Michael Hefenbrock, Jasmin Aghassi-Hagmann, Michael Beigl, Rakesh Kumar 0002, Mehdi Baradaran Tahoori |
DATE | 2 |
| 2020 | Printed MicroprocessorsabstractPrinted electronics holds the promise of meeting the cost and conformality needs of emerging disposable and ultra-low cost margin applications. Recent printed circuits technologies also have low supply voltage and can, therefore, be battery-powered. In this paper, we explore the design space of microprocessors implemented in such printing technologies - these printed microprocessors will be needed for battery-powered applications with requirements of low cost, conformality, and programmability. To enable this design space exploration, we first present the standard cell libraries for EGFET and CNT-TFT printed technologies - to the best of our knowledge, these are the first synthesis and physical design ready standard cell libraries for any low voltage printing technology. We then present an area, power, and delay characterization of several off-the-shelf low gate count microprocessors (Z80, light8080, ZPU, and openMSP430) in EGFET and CNT-TFT technologies. Our characterization shows that several printing applications can be feasibly targeted by battery-powered printed microprocessors. However, our results also show the need to significantly reduce area and power of such printed microprocessors. We perform a design space exploration of printed microprocessor architectures over multiple parameters - datawidths, pipeline depth, etc. We show that the best cores outperform pre-existing cores by at least one order of magnitude in terms of power and area. Finally, we show that printing-specific architectural and low-level optimizations further improve area and power characteristics of low voltage battery-compatible printed microprocessors. Program-specific ISA, for example, improves power, and area by up to 4.18x and 1.93x respectively. Crosspoint-based instruction ROM outperforms a RAM-based design by 5.77x, 16.8x, and 2.42x respectively in terms of power, area, and delay. Nathaniel Bleier, Muhammad Husnain Mubarik, Farhan Rasheed, Jasmin Aghassi-Hagmann, Mehdi Baradaran Tahoori, Rakesh Kumar 0002 |
ISCA | 1 |
| 2020 | Printed Machine Learning ClassifiersabstractA large number of application domains have requirements on cost, conformity, and non-toxicity that silicon-based computing systems cannot meet, but that may be met by printed electronics. For several of these domains, a typical computational task to be performed is classification. In this work, we explore the hardware cost of inference engines for popular classification algorithms (Multi-Layer Perceptrons, Support Vector Machines (SVMs), Logistic Regression, Random Forests and Binary Decision Trees) in EGT and CNT-TFT printed technologies and determine that Decision Trees and SVMs provide a good balance between accuracy and cost. We evaluate conventional Decision Tree and SVM architectures in these technologies and conclude that their area and power overhead must be reduced. We explore, through SPICE and gate-level hardware simulations and multiple working prototypes, several classifier architectures that exploit the unique cost and implementation tradeoffs in printed technologies - a) Bespoke printed classifers that are customized to a model generated for a given application using specific training datasets, b) Lookup-based printed classifiers where key hardware computations are replaced by lookup tables, and c) Analog printed classifiers where some classifier components are replaced by their analog equivalents. Our evaluations show that bespoke implementation of EGT printed Decision Trees has 48.9× lower area (average) and 75.6× lower power (average) than their conventional equivalents; corresponding benefits for bespoke SVMs are 12.8× and Decision outperform 12.7× respectively. Lookup-based Trees their non-lookup bespoke equivalents by 38% and 70%; lookup-based SVMs are better by 8% and 0.6%. Analog printed Decision Trees provide 437× area and 27× power benefits over digital bespoke counterparts; analog SVMs yield 490× area and 12× power improvements. Our results and prototypes demonstrate feasibility of fabricating and deploying battery and self-powered printed classifiers in the application domains of interest. Muhammad Husnain Mubarik, Dennis Weller, Nathaniel Bleier, Matthew Tomei, Jasmin Aghassi-Hagmann, Mehdi Baradaran Tahoori, Rakesh Kumar 0002 |
MICRO | 3 |