Alireza Farshin

dblp:201/0743 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
4since 2021 · last 2026
0000-0001-5083-4052ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 first-author · 1 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Memory systems · 46% Distributed systems · 29% Cloud and datacenter computing · 25%
Computer networks
3 papers
Routing and switching · 39% Software-defined and programmable networks · 36% Internet architecture and protocols · 21%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
1.012026
Aggregate Local, Sync Global: A Hierarchical Approach to Efficient Geo-Distributed LLM Training · INFOCOM 2026
Distributed systems › distributed machine learning
distributed training
1.012026
Aggregate Local, Sync Global: A Hierarchical Approach to Efficient Geo-Distributed LLM Training · INFOCOM 2026
Routing and switching › packet switching
packet ordering
0.612022
Packet Order Matters! Improving Application Performance by Deliberately Delaying Packets · NSDI 2022
Internet architecture and protocols
packet scheduling
0.612022
Packet Order Matters! Improving Application Performance by Deliberately Delaying Packets · NSDI 2022
Routing and switching › data plane › router data plane
high-speed packet processing
0.512021
PacketMill: toward per-Core 100-Gbps networking · ASPLOS 2021
Software-defined and programmable networks
programmable data plane
0.512021
PacketMill: toward per-Core 100-Gbps networking · ASPLOS 2021
Software-defined and programmable networks › programmable data plane
software packet processing
0.512021
PacketMill: toward per-Core 100-Gbps networking · ASPLOS 2021
Memory systems › cache management
direct cache access
0.412020
Reexamining Direct Cache Access to Optimize I/O Intensive Applications for Multi-hundred-gigabit Networks · USENIX ATC 2020
Memory systems › cache management
cache isolation
0.412019
Make the Most out of Last Level Cache in Intel Processors · EuroSys 2019
Memory systems
cache management
0.412019
Make the Most out of Last Level Cache in Intel Processors · EuroSys 2019
Memory systems › memory hierarchy › cache hierarchy
last-level cache
0.412019
Make the Most out of Last Level Cache in Intel Processors · EuroSys 2019
Cloud and datacenter computing
datacenter network
0.322021
PacketMill: toward per-Core 100-Gbps networking · ASPLOS 2021
Make the Most out of Last Level Cache in Intel Processors · EuroSys 2019
Cloud and datacenter computing
network i/o
0.112019
Make the Most out of Last Level Cache in Intel Processors · EuroSys 2019

Methods — techniques the papers use, named apart from their topics

hierarchical aggregation · 2.0code optimization · 1.0complex addressing analysis · 0.4DPDK · 0.4
YearPublicationVenuePosition
2026 CheckWork: Enabling Trace-Driven Analysis of Checkpointing Overhead in Distributed ML Training
abstract
Checkpointing is a fundamental mechanism for fault tolerance in large-scale distributed machine learning (ML) training, but it can introduce significant overhead due to interactions with compute and communication phases. Evaluating checkpointing strategies on production clusters is costly, difficult to reproduce, and limits the ability to systematically explore alternative designs. We present CheckWork, a framework that generates checkpoint-aware Chakra execution traces by augmenting training DAGs with checkpoint operations. These traces enable realistic simulation of checkpointing strategies using existing system simulators. Our evaluation shows that the generated traces reproduce checkpointing behaviour previously observed on real test-beds and reported in systems such as CheckFreq and Gemini, capturing both computational overheads and network-level communication interference. Using the traces, we also analyse interference between checkpoint traffic and training communication in multi-job environments and propose a mechanism to mitigate this effect.
Eldar Hasanov, Adel Sefiane, Alireza Farshin, Marios Kogias
APNet3
2026 Aggregate Local, Sync Global: A Hierarchical Approach to Efficient Geo-Distributed LLM Training
Francesco De Luca, Francesco De Nadai, Mariano Scazzariello, Tommaso Caiazzi, Alireza Farshin, Marco Chiesa, Giuseppe Di Battista
INFOCOM5
2022 Packet Order Matters! Improving Application Performance by Deliberately Delaying Packets
Hamid Ghasemirahni, Tom Barbette, George P. Katsikas, Alireza Farshin, Amir Roozbeh, Massimo Girondi, Marco Chiesa, Gerald Q. Maguire Jr., Dejan Kostic
NSDI4
2021 PacketMill: toward per-Core 100-Gbps networking
abstract
We present PacketMill, a system for optimizing software packet processing, which (i) introduces a new model to efficiently manage packet metadata and (ii) employs code-optimization techniques to better utilize commodity hardware. PacketMill grinds the whole packet processing stack, from the high-level network function configuration file to the low-level userspace network (specifically DPDK) drivers, to mitigate inefficiencies and produce a customized binary for a given network function. Our evaluation results show that PacketMill increases throughput (up to 36.4 Gbps -- 70%) & reduces latency (up to 101 us -- 28%) and enables nontrivial packet processing (e.g., router) at ~100 Gbps, when new packets arrive >10× faster than main memory access times, while using only one processing core.
Alireza Farshin, Tom Barbette, Amir Roozbeh, Gerald Q. Maguire Jr., Dejan Kostic
ASPLOS1
2020 Reexamining Direct Cache Access to Optimize I/O Intensive Applications for Multi-hundred-gigabit Networks
Alireza Farshin, Amir Roozbeh, Gerald Q. Maguire Jr., Dejan Kostic
USENIX ATC1
2019 Make the Most out of Last Level Cache in Intel Processors
abstract
In modern (Intel) processors, Last Level Cache (LLC) is divided into multiple slices and an undocumented hashing algorithm (aka Complex Addressing) maps different parts of memory address space among these slices to increase the effective memory bandwidth. After a careful study of Intel's Complex Addressing, we introduce a slice-aware memory management scheme, wherein frequently used data can be accessed faster via the LLC. Using our proposed scheme, we show that a key-value store can potentially improve its average performance ~12.2% and ~11.4% for 100% & 95% GET workloads, respectively. Furthermore, we propose CacheDirector, a network I/O solution which extends Direct Data I/O (DDIO) and places the packet's header in the slice of the LLC that is closest to the relevant processing core. We implemented CacheDirector as an extension to DPDK and evaluated our proposed solution for latency-critical applications in Network Function Virtualization (NFV) systems. Evaluation results show that CacheDirector makes packet processing faster by reducing tail latencies (90-99th percentiles) by up to 119 μs (~21.5%) for optimized NFV service chains that are running at 100 Gbps. Finally, we analyze the effectiveness of slice-aware memory management to realize cache isolation.
Alireza Farshin, Amir Roozbeh, Gerald Q. Maguire Jr., Dejan Kostic
EuroSys1
2019 A modified knowledge-based ant colony algorithm for virtual machine placement and simultaneous routing of NFV in distributed cloud architecture
Alireza Farshin, Saeed Sharifian
J. Supercomput.1
2017 A chaotic grey wolf controller allocator for Software Defined Mobile Network (SDMN) for 5th generation of cloud-based cellular systems (5G)
Alireza Farshin, Saeed Sharifian
Comput. Commun.1
2017 MAP-SDN: a metaheuristic assignment and provisioning SDN framework for cloud datacenters
Alireza Farshin, Saeed Sharifian
J. Supercomput.1