Nathan Youngblood

dblp:213/7757 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
2since 2021 · last 2025
0000-0003-2552-9376ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 61% Interconnection networks and networks-on-chip · 39%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
LightML: A Photonic Accelerator for Efficient General Purpose Machine Learning · ISCA 2025
Interconnection networks and networks-on-chip › optical interconnection networks
optical crossbar
0.912025
LightML: A Photonic Accelerator for Efficient General Purpose Machine Learning · ISCA 2025
Hardware accelerators and domain-specific architectures
photonic accelerator
0.912025
LightML: A Photonic Accelerator for Efficient General Purpose Machine Learning · ISCA 2025
Interconnection networks and networks-on-chip › switch architecture
buffer design
0.312025
LightML: A Photonic Accelerator for Efficient General Purpose Machine Learning · ISCA 2025

Methods — techniques the papers use, named apart from their topics

matrix multiplication · 0.9convolutional layers · 0.9
YearPublicationVenuePosition
2025 LightML: A Photonic Accelerator for Efficient General Purpose Machine Learning
abstract
The rapid integration of AI technologies into everyday life across sectors such as healthcare, autonomous driving, and smart home applications requires extensive computational resources, placing strain on server infrastructure and incurring significant costs.We present LightML, the first system-level photonic crossbar design, optimized for high-performance machine learning applications.This work provides the first complete memory and buffer architecture carefully designed to support the high-speed photonic crossbar, achieving over 80% utilization.LightML also introduces solutions for key ML functions, including large-scale matrix multiplication (MMM), element-wise operations, non-linear functions, and convolutional layers.Delivering 325 TOP/s at only 3 watts, LightML offers significant improvements in speed and power efficiency, making it ideal for both edge devices and dense data center workloads.
Sadra Rahimi Kari, Xin Xin 0008, Nathan Youngblood, Youtao Zhang, Jun Yang 0002
ISCA4
2024 OFHE: An Electro-Optical Accelerator for Discretized TFHE
abstract
This paper presents OFHE, an electro-optical accelerator designed to process Discretized TFHE (DTFHE) operations, which encrypt multi-bit messages and support homomorphic multiplications, lookup table operations and full-domain functional bootstrappings. While DTFHE is more efficient and versatile than other fully homomorphic encryption schemes, it requires 32-, 64-, and 128-bit polynomial multiplications, which can be time-consuming. Existing TFHE accelerators are not easily upgradable to support DTFHE operations due to limited datapaths, a lack of datapath bit-width reconfigurability, and power inefficiencies when processing FFT and inverse FFT (IFFT) kernels. Compared to prior TFHE accelerators, OFHE addresses these challenges by improving the DTFHE operation latency by 8.7%, the DTFHE operation throughput by 57%, and the DTFHE operation throughput per Watt by 94%.
Mengxin Zheng, Cheng Chu, Qian Lou, Nathan Youngblood, Sajjad Moazeni, Lei Jiang 0001
ISLPED4
2020 LightBulb: A Photonic-Nonvolatile-Memory-based Accelerator for Binarized Convolutional Neural Networks
abstract
Although Convolutional Neural Networks (CNNs) have demonstrated the state-of-the-art inference accuracy in various intelligent applications, each CNN inference involves millions of expensive floating point multiply-accumulate (MAC) operations. To energy-efficiently process CNN inferences, prior work proposes an electro-optical accelerator to process power-of-2 quantized CNNs by electro-optical ripple-carry adders and optical binary shifters. The electro-optical accelerator also uses SRAM registers to store intermediate data. However, electro-optical ripple-carry adders and SRAMs seriously limit the operating frequency and inference throughput of the electro-optical accelerator, due to the long critical path of the adder and the long access latency of SRAMs. In this paper, we propose a photonic nonvolatile memory (NVM)-based accelerator, Light-Bulb, to process binarized CNNs by high frequency photonic XNOR gates and popcount units. LightBulb also adopts photonic racetrack memory to serve as input/output registers to achieve high operating frequency. Compared to prior electro-optical accelerators, on average, LightBulb improves the CNN inference throughput by 17× ~ 173× and the inference throughput per Watt by 17.5 × ~ 660×.
Farzaneh Zokaee, Qian Lou, Nathan Youngblood, Weichen Liu 0001, Yiyuan Xie, Lei Jiang 0001
DATE3