Danial Zuberi

dblp:398/7159 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0002-0192-9703ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Processor architecture and microarchitecture · 55% Hardware accelerators and domain-specific architectures · 34% Reconfigurable computing and FPGAs · 10%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Operating systems › resource management › process management › CPU scheduling
user-space scheduling
0.912025
The Benefits and Limitations of User Interrupts for Preemptive Userspace Scheduling · NSDI 2025
Processor architecture and microarchitecture › exception handling
interrupt handling
0.912025
Extended User Interrupts (xUI): Fast and Flexible Notification without Polling · ASPLOS (2) 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.912025
Greater than the Sum of its LUTs: Scaling Up LUT-based Neural Networks with AmigoLUT · FPGA 2025
Machine learning › Efficient and distributed learning › large-scale learning
model scaling
0.312025
Greater than the Sum of its LUTs: Scaling Up LUT-based Neural Networks with AmigoLUT · FPGA 2025
Processor architecture and microarchitecture › instruction set architecture
processor extensions
0.312025
Extended User Interrupts (xUI): Fast and Flexible Notification without Polling · ASPLOS (2) 2025

Methods — techniques the papers use, named apart from their topics

lookup table mapping · 1.7ensemble · 1.7gem5 simulation · 0.9
YearPublicationVenuePosition
2025 Extended User Interrupts (xUI): Fast and Flexible Notification without Polling
abstract
Extended user interrupts (xUI) is a set of processor extensions that builds on Intel's UIPI model of user interrupts, for enhanced performance and flexibility. This paper deconstructs Intel's current UIPI design through analysis and measurement, and uses this to develop an accurate model of its timing. It then introduces four novel enhancements to user interrupts: tracked interrupts, hardware safepoints, a kernel bypass timer, and interrupt forwarding. xUI is modeled in gem5 simulation and evaluated on three use cases -- preemption in a high-performance user-level runtime, IO notification in a layer3 router using DPDK, and IO notification in a synthetic workload with a streaming accelerator modeled after Intel's Data Streaming Accelerator. This work shows that xUI offers the performance of shared memory polling with the efficiency of asynchronous notification.
Berk Aydogmus, Linsong Guo, Danial Zuberi, Tal Garfinkel, Dean M. Tullsen, Amy Ousterhout, Mohammadkazem Taram
ASPLOS (2)3
2025 Greater than the Sum of its LUTs: Scaling Up LUT-based Neural Networks with AmigoLUT
abstract
Applications like high-energy physics and cybersecurity require extremely high throughput and low latency neural network (NN) inference. Lookup-table-based NNs address these constraints by implementing NNs as lookup tables (LUTs), achieving inference latency on the order of nanoseconds. Since LUTs are a fundamental FPGA building block, LUT-based NNs efficiently map to FPGAs. LogicNets (and its successors) form one class of LUT-based NNs that target FPGAs, mapping neurons directly to LUTs to meet low latency constraints with minimal resources. However, it is difficult to build larger, more performant LUT-based NNs like LogicNets because LUT usage increases exponentially with respect to neuron fan-in (i.e., number of synapses X synapse bitwidth). A large LUT-based NN quickly runs out of LUTs on an FPGA. Our work AmigoLUT addresses this issue by creating ensembles of smaller LUT-based NNs that scale linearly with respect to the number of models. AmigoLUT improves the scalability of LUT-based NNs, reaching higher throughput with up to an order of magnitude fewer LUTs than the largest LUT-based NNs.
Olivia Weng, Marta Andronic, Danial Zuberi, Caleb Geniesse, George A. Constantinides, Nicholas J. Fraser, Javier M. Duarte, Ryan Kastner
FPGA3
2025 The Benefits and Limitations of User Interrupts for Preemptive Userspace Scheduling
Linsong Guo, Danial Zuberi, Tal Garfinkel, Amy Ousterhout
NSDI2