EDBT 2026 Demo / reviewers in the wild / expert
Diman Zad Tootaghaj
dblp:170/3257
· DBLP profile ↗
17ranked-venue papers
9as first author
7since 2021 · last 2026
0000-0003-3664-4484ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 5 first-author · 2 since 2021Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Not A DPU in Name Only! Unleashing RDMA-capable DPUs in Multi-Tenant Serverless Clouds with NADINO
Shixiong Qi, Songyu Zhang, K. K. Ramakrishnan, Diman Zad Tootaghaj, Hardik Soni 0001, Puneet Sharma 0001 |
EuroSys | 4 |
| 2026 | AdaGen: Workload-Adaptive Cluster Scheduler for Latency-Optimal LLM Inference ServingabstractThe inference workloads of Large Language Models (LLMs) pose significant latency and cost challenges due to increasing model sizes and demand for real-time responses. Existing cluster schedulers for multi-instance LLM serving primarily focus on load balancing to optimize memory usage, which is insufficient for workloads with diverse request characteristics. In such cases, the compute layout — the arrangement of tokens across iterations within each instance—plays a crucial role in determining latency. We propose AdaGen, a workload-adaptive cluster scheduler that minimizes latency and thus maximizes SLO attainment by optimizing compute layouts across instances. AdaGen employs a multi-step scheduling strategy: it first classifies requests based on prefill and decode lengths, then balances load, and finally performs selective distributed execution across instances. Each step incrementally refines the scheduling based on the compute layouts derived from the decision of the previous step. To avoid the overhead of actual execution to generate the layouts, AdaGen introduces a novel simulation-based estimator. Extensive experiments using production workloads show that AdaGen achieves up to 3.6× higher SLO attainment and 2× better cost-efficiency compared to the existing systems, while ensuring scalability. Sudipta Saha Shubha, Ayush Goel, Diman Zad Tootaghaj, Khaled Diab 0001, Hardik Soni 0001, K. K. Ramakrishnan, Puneet Sharma 0001, Haiying Shen |
EuroSys | 3 |
| 2026 | Griffin: Coherency-Aware Task Scheduling and Memory Allocation for CXL InterconnectsabstractCXL is an emerging interconnect that has the potential to efficiently realize memory disaggregation. This is because CXL enables the expansion of memory beyond individual hosts, and supports coherent memory sharing among multiple hosts. However, CXL introduces several performance overheads due to the cache coherency protocol for memory sharing, as well as placement constraints for shared data, which, if ignored, can lead to correctness issues. This paper presents the first analysis of the impact of CXL memory sharing and shows that the overheads of hardware-based coherency in CXL interconnects are substantial. We then propose Griffin, a new coherency-aware task and memory allocator for CXL disaggregated memory systems. Griffin introduces new abstractions and algorithms that allow it to prioritize which data is allocated remotely and to which memory node, to efficiently reduce the coherence overheads associated with both the amount of shared data and the load on CXL coherence resources. Our simulation results show that Griffin reduces the total memory time by up to 4.29 × compared to a standard baseline and 1.71 × compared to an advanced baseline. Suyeon Lee, Khaled Diab 0001, Diman Zad Tootaghaj, Lianjie Cao, Puneet Sharma 0001, Ada Gavrilovska |
ICS | 3 |
| 2026 | DynamoServe: A Distributed Tiered Memory System for Multi-tenant LLM ServingabstractThe rapid adoption of large language models (LLMs) has increased the need for efficient multi-tenant inference systems that maximize GPU utilization. However, existing frameworks struggle to scale due to the high memory demands of model weights and key-value (KV) caches. We present DynamoServe, a multi-tenant LLM serving framework that addresses these challenges through three key innovations: (1) leveraging stranded GPU memory to offload model weights and KV caches, (2) mitigating resource fragmentation in multi-workload environments, and (3) improving memory locality through coordinated data placement and demand-driven weight migration across GPUs. Together, these techniques enable high-throughput, low-latency inference. Experiments on state-of-the-art models show that DynamoServe significantly improves memory efficiency without sacrificing latency. Diman Zad Tootaghaj, Khaled Diab 0001, Bob Lantz, Hanjiang Wu, K. K. Ramakrishnan, Md Ashfaqur Rahaman, Ryan Stutsman, Puneet Sharma 0001, Tushar Krishna |
SIGCOMM | 1 |
| 2026 | Are We There Yet? Predicting if Executing Applications are Near Completion
Mohammad Sonji, Mohammed Baydoun, Safaa Diab, Amir Nassereldine, Pedro Bruel, Aditya Dhakal, Rolando P. Hong Enriquez, Gourav Rattihalli, Diman Zad Tootaghaj, Gallig Renaud, Barbara M. Chapman, Fatima K. Abu Salem, Eitan Frachtenberg, Dejan S. Milojicic, Izzat El Hajj |
ICPE | 9 |
| 2025 | Palladium: A DPU-enabled Multi-Tenant Serverless Cloud over Zero-copy Multi-node RDMA FabricsabstractServerless computing offers resource efficiency but suffers from a heavyweight data plane. We present Palladium, a DPU-offloaded serverless data plane enabling distributed zero-copy communication. Palladium uses two-sided RDMA and cross-processor shared memory to mitigate limitations of wimpy DPU cores. Its DPU-enabled network engine (DNE) isolates RDMA resources and manages flows across tenants. By converting HTTP/TCP to RDMA at ingress, Palladium reduces protocol overhead on the critical path. Shixiong Qi, Songyu Zhang, K. K. Ramakrishnan, Diman Zad Tootaghaj, Hardik Soni 0001, Puneet Sharma 0001 |
SIGCOMM | 4 |
| 2024 | Conspirator: SmartNIC-Aided Control Plane for Distributed ML Workloads
Yunming Xiao, Diman Zad Tootaghaj, Aditya Dhakal, Lianjie Cao, Puneet Sharma 0001, Aleksandar Kuzmanovic |
USENIX ATC | 2 |
| 2020 | Homa: An Efficient Topology and Route Management Approach in SD-WAN OverlaysabstractThis paper presents an efficient topology and route management approach in Software-Defined Wide Area Networks (SD-WAN). Traditional WANs suffer from low utilization and lack of global view of the network. Therefore, during failures, topology/service/traffic changes, or new policy requirements, the system does not always converge to the global optimal state. Using Software Defined Networking architectures in WANs provides the opportunity to design WANs with higher fault tolerance, scalability, and manageability. We exploit the correlation matrix derived from monitoring system between the virtual links to infer the underlying route topology and propose a route update approach that minimizes the total route update cost on all flows. We formulate the problem as an integer linear programming optimization problem and provide a centralized control approach that minimizes the total cost while satisfying the quality of service (QoS) on all flows. Experimental results on real network topologies demonstrate the effectiveness of the proposed approach in terms of disruption cost and average disrupted flows. Diman Zad Tootaghaj, Faraz Ahmed, Puneet Sharma 0001, Mihalis Yannakakis |
INFOCOM | 1 |
| 2019 | Modeling, Monitoring and Scheduling Techniques for Network Recovery from Massive Failures
Diman Zad Tootaghaj, Thomas La Porta, Ting He 0001 |
IM | 1 |
| 2019 | Poster: a minimally disruptive network reconfiguration approach in SDNabstractWhen routing flows in a software defined network (SDN), service disruption and inconsistencies can occur during the updates of routing tables leading to degraded QoS or interruption of existing services. We study the problem of rerouting existing flows in an SDN to enable the admission of new flows while minimizing the disruption of existing flows, under link capacity and Quality of Service (QoS) constraints. We formulate the problem as an integer linear programming problem and propose two randomized rounding algorithms with bounded congestion and demand loss to solve this problem. Diman Zad Tootaghaj, Stefan Achleitner, Ting He 0001, Novella Bartolini, Thomas La Porta |
Networking | 1 |
| 2019 | On Progressive Network Recovery From Massive Failures Under UncertaintyabstractNetwork recovery after large-scale failures has tremendous cost implications. While numerous approaches have been proposed to restore critical services after large-scale failures, they mostly assume having full knowledge of failure location, which cannot be achieved in real failure scenarios. Making restoration decisions under uncertainty is often further complicated in a large-scale failure. This paper addresses progressive network recovery under the uncertain knowledge of damages. We formulate the problem as a mixed integer linear programming and show that it is NP-hard. We propose an iterative stochastic recovery algorithm (ISR) to recover the network in a progressive manner to satisfy the critical services. At each optimization step, we make a decision to repair a part of the network and gather more information iteratively, until critical services are completely restored. We propose three different approaches: 1) an iterative shortest path algorithm; 2) an approximate branch and bound (ISR-BB); and 3) an iterative multicommodity LP relaxation (ISR-MULT). Further, we compared our approach with the state-of-the-art centrality-based damage assessment and recovery (CeDAR) and iterative split and prune (ISP) algorithms. Our results show that ISR-BB and ISR-MULT outperform the state-of-the-art ISP and CeDAR algorithms while we can configure our choice of tradeoff between the execution time, the number of repairs (cost), and the demand loss. We show that our recovery algorithm, on average, can reduce the total number of repairs by a factor of about 3 with respect to ISP, while satisfying all critical demands. Diman Zad Tootaghaj, Novella Bartolini, Hana Khamfroush, Thomas La Porta |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2018 | Stochastic Modeling and Optimization of StragglersabstractMapReduce framework is widely used to parallelize batch jobs since it exploits a high degree of multi-tasking to process them. However, it has been observed that when the number of servers increases, the map phase can take much longer than expected. This paper analytically shows that the stochastic behavior of the servers has a negative effect on the completion time of a MapReduce job, and continuously increasing the number of servers without accurate scheduling can degrade the overall performance. We analytically model the map phase in terms of hardware, system, and application parameters to capture the effects of stragglers on the performance. Mean sojourn time (MST), the time needed to sync the completed tasks at a reducer, is introduced as a performance metric and mathematically formulated. Following that, we stochastically investigate the optimal task scheduling which leads to an equilibrium property in a datacenter with different types of servers. Our experimental results show the performance of the different types of schedulers targeting MapReduce applications. We also show that, in the case of mixed deterministic and stochastic schedulers, there is an optimal scheduler that can always achieve the lowest MST. Farshid Farhat, Diman Zad Tootaghaj, Yuxiong He, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das |
IEEE Trans. Cloud Comput. | 2 |
| 2018 | Fast Network Configuration in Software Defined NetworkingabstractSoftware defined networking (SDN) provides a framework to dynamically adjust and re-program the data plane with the use of flow rules. The realization of highly adaptive SDNs with the ability to respond to changing demands or recover after a network failure in a short period of time, hinges on efficient updates of flow rules. We model the time to deploy a set of flow rules by the update time at the bottleneck switch, and formulate the problem of selecting paths to minimize the deployment time under feasibility constraints as a mixed integer linear program (MILP). To reduce the computation time of determining flow rules, we propose efficient heuristics designed to approximate the minimum-deployment-time solution by relaxing the MILP or selecting the paths sequentially. Through extensive simulations we show that our algorithms outperform current, shortest path-based solutions by reducing the total network configuration time up to 55% while having similar packet loss, in the considered scenarios. We also demonstrate that in a networked environment with a certain fraction of failed links, our algorithms are able to reduce the average time to reestablish disrupted flows by 40%. Stefan Achleitner, Novella Bartolini, Ting He 0001, Thomas La Porta, Diman Zad Tootaghaj |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2017 | CAGE: A Contention-Aware Game-Theoretic Model for Heterogeneous Resource AssignmentabstractTraditional resource management systems rely on a centralized approach to manage users running on each resource. The centralized resource management system is not scalable for large-scale servers as the number of users running on shared resources is increasing dramatically and the centralized manager may not have enough information about applications' need. In this paper we propose a distributed game-theoretic resource management approach using market auction mechanism to find optimal strategy in a resource competition game. The applications learn through repeated interactions to choose their action on choosing the shared resources. Specifically, we look into two case studies of cache competition game and main processor and co-processor congestion game. We enforce costs for each resource and derive bidding strategy. Accurate evaluation of the proposed approach show that our distributed allocation is scalable and outperforms the static and traditional approaches. Diman Zad Tootaghaj, Farshid Farhat |
ICCD | 1 |
| 2017 | Controlling Cascading Failures in Interdependent Networks under Incomplete KnowledgeabstractVulnerability due to inter-connectivity of multiple networks has been observed in many complex networks. Previous works mainly focused on robust network design and on recovery strategies after sporadic or massive failures in the case of complete knowledge of failure location. We focus on cascading failures involving the power grid and its communication network with consequent imprecision in damage assessment. We tackle the problem of mitigating the ongoing cascading failure and providing a recovery strategy. We propose a failure mitigation strategy in two steps: 1) Once a cascading failure is detected, we limit further propagation by re-distributing the generator and load's power. 2) We formulate a recovery plan to maximize the total amount of power delivered to the demand loads during the recovery intervention. Our approach to cope with insufficient knowledge of damage locations is based on the use of a new algorithm to determine consistent failure sets (CFS). We show that, given knowledge of the system state before the disruption, the CFS algorithm can find all consistent sets of unknown failures in polynomial time provided that, each connected component of the disrupted graph has at least one line whose failure status is known to the controller. Diman Zad Tootaghaj, Novella Bartolini, Hana Khamfroush, Thomas La Porta |
SRDS | 1 |
| 2011 | Game-theoretic approach to mitigate packet dropping in wireless Ad-hoc networksabstractPerformance of routing is severely degraded when misbehaving nodes drop packets instead of properly forwarding them. In this paper, we propose a Game-Theoretic Adaptive Multipath Routing (GTAMR) protocol to detect and punish selfish or malicious nodes which try to drop information packets in routing phase and defend against collaborative attacks in which nodes try to disrupt communication or save their power. Our proposed algorithm outranks previous schemes because it is resilient against attacks in which more than one node coordinate their misbehavior and can be used in networks which wireless nodes use directional antennas. We then propose a game theoretic strategy, ERTFT, for nodes to promote cooperation. In comparison with other proposed TFT-like strategies, ours is resilient to systematic errors in detection of selfish nodes and does not lead to unending death spirals. Diman Zad Tootaghaj, Farshid Farhat, Mohammad Reza Pakravan, Mohammad Reza Aref |
CCNC | 1 |
| 2011 | Risk of attack coefficient effect on availability of Ad-hoc networksabstractSecurity techniques have been designed to obtain certain objectives. One of the most important objectives all security mechanisms try to achieve is the availability, which insures that network services are available to various entities in the network when required. But there has not been any certain parameter to measure this objective in network. In this paper we consider availability as a security parameter in ad-hoc networks. However this parameter can be used in other networks as well. We also present the connectivity coefficient of nodes in a network which shows how important is a node in a network and how much damage is caused if a certain node is compromised. Diman Zad Tootaghaj, Farshid Farhat, Mohammad Reza Pakravan, Mohammad Reza Aref |
CCNC | 1 |