EDBT 2026 Demo / reviewers in the wild / expert
Annus Zulfiqar
dblp:255/7879
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0003-0612-4939ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Computer networks · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPLIDT: Partitioned Decision Trees for Scalable Stateful Inference at Line Rate
Murayyiam Parvez, Annus Zulfiqar, Roman Beltiukov, Shir Landau Feibish, Walter Willinger, Arpit Gupta, Muhammad Shahbaz 0001 |
NSDI | 2 |
| 2026 | Towards Network-Efficient Cross-Regional Inference via Learned Activation CompressionabstractLarge transformer models are increasingly deployed across geographically distributed GPU clusters due to capacity, cost, and locality constraints. When inference is partitioned across sites, intermediate activations must be transmitted over wide area network (WAN) links at each partition boundary, introducing significant communication overhead. We present Feather, a system that reduces this overhead by compressing intermediate activations before transmission and reconstructing them before downstream layers resume execution. Regan McDonald, Marilyn Rego, Ertza Warraich, Annus Zulfiqar, Muhammad Shahbaz 0001 |
SIGCOMM | 4 |
| 2025 | Gigaflow: Pipeline-Aware Sub-Traversal Caching for Modern SmartNICsabstractThe success of modern public/edge clouds hinges heavily on the performance of their end-host network stacks if they are to support the emerging and diverse tenants' workloads (e.g., distributed training in the cloud to fast inference at the edge). Virtual Switches (vSwitches) are vital components of this stack, providing a unified interface to enforce high-level policies on incoming packets and route them to physical interfaces, containers, or virtual machines. As performance demands escalate, there has been a shift toward offloading vSwitch processing to SmartNICs to alleviate CPU load and improve efficiency. However, existing solutions struggle to handle the growing flow rule space within the NIC, leading to high miss rates and poor scalability. Annus Zulfiqar, Ali Imran 0005, Venkat Kunaparaju, Ben Pfaff, Gianni Antichi, Muhammad Shahbaz 0001 |
ASPLOS (2) | 1 |
| 2025 | NetSparse: In-Network Acceleration of Distributed Sparse Kernels
Gerasimos Gerogiannis, Dimitrios Merkouriadis, Charles Block, Annus Zulfiqar, Filippos Tofalos, Muhammad Shahbaz 0001, Josep Torrellas |
MICRO | 4 |
| 2025 | SpliDT: Partitioned Decision Trees for Scalable Stateful Inference at Line RateabstractMachine learning is increasingly used in programmable data planes, such as switches [4, 12, 13] and smartNICs [1, 16], to enable real-time traffic analysis and security monitoring at line rate. Decision trees (DTs) are particularly well-suited for these tasks due to their interpretability and compatibility with the Reconfigurable Match-Action Table (RMT) architecture. However, current DT implementations require collecting all features upfront, which limits scalability and accuracy due to constrained data plane resources. Murayyiam Parvez, Annus Zulfiqar, Roman Beltiukov, Shir Landau Feibish, Walter Willinger, Arpit Gupta, Muhammad Shahbaz 0001 |
SIGCOMM | 2 |
| 2024 | A Smart Cache for a SmartNIC! Scaling End-Host Networking to 400Gbps and Beyondabstract•Virtual switches optimize performance by caching multi-table lookup traversals to single-table Megaflow cache, which SmartNICs offload directly to hardware •We present Gigaflow: a multi-table sub-traversal cache for SmartNICs, designed to capture a much larger rule space using the same cache size •Open vSwitch caches traversals into Megaflow and can't share sub-traversals among traffic, making the captured rule space proportional to cache size •By caching sub-traversals into a multi-table cache, we can capture 3 orders of magnitude more rule space, attain 51% higher cache hit rate, and 31% lower end-to-end packet latency, with manageable processing overhead Annus Zulfiqar, Ali Imran 0005, Venkat Kunaparaju, Ben Pfaff, Gianni Antichi, Muhammad Shahbaz 0001 |
HCS | 1 |
| 2023 | Homunculus: Auto-Generating Efficient Data-Plane ML Pipelines for Datacenter NetworksabstractSupport for Machine Learning (ML) applications in networking has significantly improved over the last decade. The availability of public datasets and programmable switching fabrics (including low-level languages to program them) presents a full-stack to the programmer for deploying in-network ML. However, the diversity of tools involved, coupled with complex optimization tasks of ML model design and hyperparameter tuning while complying with the network constraints (like throughput and latency), puts the onus on the network operator to be an expert in ML, network design, and programmable hardware. Tushar Swamy, Annus Zulfiqar, Luigi Nardi, Muhammad Shahbaz 0001, Kunle Olukotun |
ASPLOS (3) | 2 |