Ali Munir

dblp:32/9285 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
3since 2021 · last 2026
0000-0001-5148-4306ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 13 · 8 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 2Artificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
10 papers
Software-defined and programmable networks · 31% Network optimization and economics · 28% Datacenter networks · 17%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
Cloud and datacenter computing · 57% Parallel and multicore computing · 26% High-performance computing · 13%
Network and information security
1 paper
Network security · 100%

Topics — the 22 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software-defined and programmable networks
programmable data plane
0.622018
OpenFunction: An Extensible Data Plane Abstraction Protocol for Platform-Independent Software-Defined Middleboxes · IEEE/ACM Trans. Netw. 2018
OpenFunction: An extensible data plane abstraction protocol for platform-independent software-defined middleboxes · ICNP 2016
Datacenter networks
datacenter transport
0.522017
PASE: Synthesizing Existing Transport Strategies for Near-Optimal Data Center Transport · IEEE/ACM Trans. Netw. 2017
Friends, not foes: synthesizing existing transport strategies for data center networks · SIGCOMM 2014
Cloud and datacenter computing
cluster resource management and scheduling
0.412020
Network Scheduling and Compute Resource Aware Task Placement in Datacenters · IEEE/ACM Trans. Netw. 2020
High-performance computing › data transfer
data transfer optimization
0.412020
Optimizing Geo-Distributed Data Analytics with Coordinated Task Scheduling and Routing · IEEE Trans. Parallel Distributed Syst. 2020
Cloud and datacenter computing › datacenter architecture
geo-distributed datacenters
0.412020
Optimizing Geo-Distributed Data Analytics with Coordinated Task Scheduling and Routing · IEEE Trans. Parallel Distributed Syst. 2020
Parallel and multicore computing
task allocation
0.412020
Network Scheduling and Compute Resource Aware Task Placement in Datacenters · IEEE/ACM Trans. Netw. 2020
Parallel and multicore computing
task scheduling
0.412020
Optimizing Geo-Distributed Data Analytics with Coordinated Task Scheduling and Routing · IEEE Trans. Parallel Distributed Syst. 2020
Internet architecture and protocols › traffic management
bandwidth management
0.312018
CoMan: Managing Bandwidth Across Computing Frameworks in Multiplexed Datacenters · IEEE Trans. Parallel Distributed Syst. 2018
Network optimization and economics › resource allocation › bandwidth allocation
in-network bandwidth allocation
0.312018
CoMan: Managing Bandwidth Across Computing Frameworks in Multiplexed Datacenters · IEEE Trans. Parallel Distributed Syst. 2018
Network optimization and economics › resource allocation
bandwidth allocation
0.312017
Multi-tenant multi-objective bandwidth allocation in datacenters using stacked congestion control · INFOCOM 2017
Cloud and datacenter computing › job scheduling
datacenter scheduling
0.212016
Network Scheduling Aware Task Placement in Datacenters · CoNEXT 2016
Cloud and datacenter computing › job scheduling › network-aware scheduling
network-aware task placement
0.212016
Network Scheduling Aware Task Placement in Datacenters · CoNEXT 2016
Network optimization and economics
network scheduling
0.222020
Network Scheduling and Compute Resource Aware Task Placement in Datacenters · IEEE/ACM Trans. Netw. 2020
Network Scheduling Aware Task Placement in Datacenters · CoNEXT 2016
Transport protocols and congestion control
explicit congestion notification
0.212013
Minimizing flow completion times in data centers · INFOCOM 2013
Datacenter networks › datacenter transport
flow completion time optimization
0.212013
Minimizing flow completion times in data centers · INFOCOM 2013
Transport protocols and congestion control
transport protocols
0.212013
Minimizing flow completion times in data centers · INFOCOM 2013
Electronic design automation › high-level synthesis › scheduling
makespan minimization
0.112020
Optimizing Geo-Distributed Data Analytics with Coordinated Task Scheduling and Routing · IEEE Trans. Parallel Distributed Syst. 2020
Network optimization and economics › resource sharing
bandwidth sharing
0.112018
CoMan: Managing Bandwidth Across Computing Frameworks in Multiplexed Datacenters · IEEE Trans. Parallel Distributed Syst. 2018
Network optimization and economics
resource allocation
0.112018
CoMan: Managing Bandwidth Across Computing Frameworks in Multiplexed Datacenters · IEEE Trans. Parallel Distributed Syst. 2018
Transport protocols and congestion control › multipath transport
multipath TCP
0.112017
Multipath TCP traffic diversion attacks and countermeasures · ICNP 2017
Cloud and datacenter computing
performance isolation
0.112017
Multi-tenant multi-objective bandwidth allocation in datacenters using stacked congestion control · INFOCOM 2017
Wireless networking › scheduling
scheduling policy
0.012013
Minimizing flow completion times in data centers · INFOCOM 2013

Methods — techniques the papers use, named apart from their topics

task completion time prediction · 1.4ns-2 simulation · 1.1simulation · 1.0virtual link groups · 0.7optimization · 0.7approximation algorithm · 0.7testbed experimentation · 0.6distributed congestion control · 0.6hypergraph partitioning · 0.4coordination mechanism · 0.4data plane abstraction protocol · 0.3protocol synthesis · 0.3
YearPublicationVenuePosition
2026 Simplifying Prioritization and Scheduling with P2CS
abstract
Modern datacenter networks host diverse services and workloads with varying quality-of-service (QoS) requirements, yet are constrained by hardware limitations. Most notably, the small number of physical priority queues available in commodity switches. Existing scheduling mechanisms, whether end-host or in-network based, struggle to scale under these constraints due to their reliance on global priority information or complex queue management. This paper presents P2CS (Priority-based Probabilistic Congestion Signaling), a lightweight and scalable approach that enables fine-grained traffic prioritization using only a single FIFO queue. P2CS combines priority-aware probabilistic congestion signaling, priority-aware packet dropping, and simple switch-side arbitration to enforce prioritization across flows. P2CS supports a range of scheduling objectives, and requires minimal software changes making it readily deployable in today's datacenter infrastructure. Evaluation on representative workloads, including multi-tenant ML training, HPC, and mixed spray/ECMP traffic, demonstrates that P2CS achieves performance comparable to in-network mechanisms while significantly reducing complexity and cost.
Ali Munir, Xiaolin Pang
SIGCOMM1
2023 Host-Assisted Transport Layer in Data Centers Using Network-Aware Rate Adjustment
abstract
Next generation applications for datacenters, such as Distributed Machine Learning (DML) and Big Data, have complex communication patterns that demand a scalable, stateless and application-aware optimal transport protocol to maximize network utilization and improve application performance. Recent transport protocols either provide limited benefits due to lack of information sharing between application and network; or implement complex stateful mechanisms to improve the application performance. In this paper, we present Omni- Transport Mechanism (Omni-TM) as a message-based congestion control protocol. Omni-TM allows exchanging message information with the network to negotiate the optimal transmission rate without maintaining a per-flow state at the switches (i.e., stateless). Omni- Tmis designed to reach maximum link capacity in one shot. Our simulation results show that Omni- Tmdemonstrates better traffic control decisions (i.e., close to zero queue length while maintaining high link utilization). Furthermore, Omni- Tmreduces Flow Completion Time (FCT) up to 45 % in a realistic workload compared to DCTCP.
Mahmoud Mohamed Bahnasy, S. Hossein Mortazavi, Ali Munir, Hossein Shafieirad, Yashar Ganjali
GLOBECOM3
2022 Special Issue on IFIP Networking 2019
Alex X. Liu, Ali Munir, Jacek Rak, Steve Uhlig, Jordi Domingo-Pascual
Comput. Commun.2
2020 Network Scheduling and Compute Resource Aware Task Placement in Datacenters
abstract
To improve the performance of data-intensive applications, existing datacenter schedulers optimize either the placement of tasks or the scheduling of network flows. The task scheduler strives to place tasks close to their input data (i.e., maximize data locality) to minimize network traffic, while assuming fair sharing of the network. The network scheduler strives to finish flows as quickly as possible based on their sources and destinations determined by the task scheduler, while the scheduling is based on flow properties (e.g., size, deadline, and correlation) and not bound to fair sharing. Inconsistent assumptions of the two schedulers can compromise the overall application performance. In this paper, we propose NEAT+, a task scheduling framework that leverages information from the underlying network scheduler and available compute resources to make task placement decisions. The core of NEAT+ is a task completion time predictor that estimates the completion time of a task under given network condition and a given network scheduling policy. NEAT+ leverages the predicted task completion times to minimize the average completion time of active tasks. Evaluation using ns2 simulations and real-testbed shows that NEAT+ improves application performance by up to 3.7x for the suboptimal network scheduling policies and up to 33% for the optimal network scheduling policy.
Ali Munir, Ting He 0001, Ramya Raghavendra, Franck Le, Alex X. Liu
IEEE/ACM Trans. Netw.1
2020 Optimizing Geo-Distributed Data Analytics with Coordinated Task Scheduling and Routing
abstract
Recent trends show that cloud computing is growing to span more and more globally distributed datacenters. For geo-distributed datacenters, there is an increasingly need for scheduling algorithms to place tasks across datacenters, by jointly considering WAN traffic and computation. This scheduling must deal with situations such as wide-area distributed data, data sharing, WAN bandwidth costs and datacenter capacity limits, while also minimizing makespan. However, this scheduling problem is NP-hard. We propose a new resource allocation algorithm called HPS+, an extension to Hypergraph Partition-based Scheduling. HPS+ models the combined task-data dependencies and data-datacenter dependencies as an augmented hypergraph, and adopts an improved hypergraph partition technique to minimize WAN traffic. It further uses a coordination mechanism to allocate network resources closely following the guidelines of task requirements, for minimizing the makespan. Evaluation across the real China-Astronomy-Cloud model and Google datacenter model show that HPS+ saves the amount of data transfers by upto 53 percent and reduces the makespan by 39 percent compared to existing algorithms.
Laiping Zhao, Ali Munir, Alex X. Liu, Wenyu Qu
IEEE Trans. Parallel Distributed Syst.3
2018 OpenFunction: An Extensible Data Plane Abstraction Protocol for Platform-Independent Software-Defined Middleboxes
Chen Tian 0001, Ali Munir, Alex X. Liu, Yangming Zhao
IEEE/ACM Trans. Netw.2
2018 CoMan: Managing Bandwidth Across Computing Frameworks in Multiplexed Datacenters
abstract
Inefficient bandwidth sharing in a datacenter network, between different application frameworks, e.g., MapReduce and Spark, can lead to inelastic and skewed usage of link bandwidth and increased completion times for the applications. Existing work, however, either solely focuses on managing computation and storage resources or controlling only sending/receiving rate at hosts. In this paper, we present CoMan, a solution that provides global in-network bandwidth management in multiplexed data centers, with two goals: improving bandwidth utilization and reducing application completion time. CoMan first designs a novel abstraction of virtual link groups (VLGs) to establish a shared bandwidth resource pool. Based on this pool, CoMan implements a three-level bandwidth allocation model, which enables elastic bandwidth sharing among computing frameworks as well as guarantees network performance for the applications. CoMan further improves the bandwidth utilization by devising a VLG dependency graph and solves an optimization problem to guide the path selection using a 32-approximation algorithm. We conduct comprehensive trace-driven simulations as well as small-scale testbed experiments to evaluate the performance of CoMan. Extensive simulation results show that CoMan improves the bandwidth utilization and speeds up the application completion time by up to 2.83× and 6.68×, respectively, compared to the ECMP + ElasticSwitch solution. Our implementation also verifies that CoMan can realistically speed up the application completion times by 2.32× on average.
Wenxin Li 0001, Deke Guo, Alex X. Liu, Keqiu Li, Heng Qi, Song Guo 0001, Ali Munir, Xiaoyi Tao
IEEE Trans. Parallel Distributed Syst.7
2017 Multipath TCP traffic diversion attacks and countermeasures
abstract
Multipath TCP (MPTCP) is an IETF standardized suite of TCP extensions that allow two endpoints to simultaneously use multiple paths between them. In this paper, we report vulnerabilities in MPTCP that arise because of cross-path interactions between MPTCP subflows. First, an attacker eavesdropping one MPTCP subflow can infer throughput of other subflows. Second, an attacker can inject forged MPTCP packets to change priorities of any MPTCP subflow. We present two attacks to exploit these vulnerabilities. In the connection hijack attack, an attacker takes full control of the MPTCP connection by suspending the subflows he has no access to. In the traffic diversion attack, an attacker diverts traffic from one path to other paths. Proposed vulnerabilities fixes, changes to MPTCP specification, provide the guarantees that MPTCP is at least as secure as TCP and the original MPTCP. We validate attacks and prevention mechanism, using MPTCP Linux implementation (v0.91), on a real-network testbed.
Ali Munir, Zhiyun Qian, Zubair Shafiq, Alex X. Liu, Franck Le
ICNP1
2017 Multi-tenant multi-objective bandwidth allocation in datacenters using stacked congestion control
abstract
In datacenter networks, flows can have different performance objectives. We use a tenant-objective division to denote all flows of a tenant that share the same objective. Bandwidth allocation in datacenters should support not only performance isolation among divisions but also objective-oriented scheduling among flows within the same division. This paper studies the Multi-Tenant Multi-Objective (MT-MO) bandwidth allocation problem. To our best knowledge, no existing practical work support performance isolation and objective scheduling simultaneously. We propose Stacked Congestion Control (SCC), a distributed host-based bandwidth allocation design, where an underlay congestion control (UCC) layer handles contention among divisions, and a private congestion control (PCC) layer for each division optimizes its performance objective. Via the tenant-objective tunnel abstraction, SCC achieves weighted bandwidth sharing for each division in a distributed and transparent way. By adding a rate-limiting send queue in the ingress of each tunnel, mechanisms between performance isolation and objective scheduling are completely decoupled. We evaluate SCC both on a small-scale testbed and with large-scale NS-2 simulations. Compared to the direct coexistence cases, SCC reduces latency by up to 40% for Latency-Sensitive flows, deadline miss ratio by up to 3.2× for Deadline-Sensitive flows, and average flow-completion-time by up to 53% for Completion-Sensitive flows.
Chen Tian 0001, Ali Munir, Alex X. Liu, Yingtong Liu, Yanzhao Li, Fan Zhang 0016, Gong Zhang 0001
INFOCOM2
2017 A Speed Hump Sensing Approach to Global Positioning in Urban Cities without Gps Signals
abstract
Outdoor localization is of great importance for driving navigation, attracting many research efforts in past decades. Prevailing GPS achieves meter-level localization accuracy under general outdoor conditions. Yet, GPS service performs poorly in urban canyons where skyscrapers blocks GPS signals and drain mobile phone battery quickly within few hours. In this work, we exploit common city facilities, i.e. speed humps, as an indicator for vehicle location. The key insight is that when the vehicle passes through the speed bump, it experiences significant fluctuations, causing larger acceleration in the vertical direction. On this basis, we design a localization scheme that utilizes the accelerator equipped on modern smart phones to track sequence of speed bumps, which is further transferred into sequence of moving directions of the vehicle, and adopt effective road mapping technology to derive real-time location. As we have concerned, it is the first attempt to exploit the spatiotemporal characteristics generated by the speed humps to recover the trajectory of the travelling route and infer the current position. Experimental results in typical outdoor environment (campus) demonstrate a comparable performance with GPS method, yet achieve lower energy consumption.
Qiuxia Chen, Dongdong Ding, Xu Wang 0018, Alex X. Liu, Ali Munir
SMARTCOMP5
2017 PASE: Synthesizing Existing Transport Strategies for Near-Optimal Data Center Transport
abstract
Several data center transport protocols have been proposed in recent years (e.g., DCTCP, PDQ, and pFabric). In this paper, we first identify the underlying strategies used by the existing data center transports, namely, in-network Prioritization (used in pFabric), Arbitration (used in PDQ), and Self-adjusting at Endpoints (PASE) (used in DCTCP). We show that these strategies are complimentary to each other, rather than substitutes, as they have different strengths and can address each other's limitations. Unfortunately, prior data center transports use only one of these strategies. As a result, they either achieve near-optimal performance or deployment friendliness (i.e., require no changes to the data plane) but not both. Based on this insight, we design a data center transport protocol called PASE, which carefully synthesizes these strategies by assigning different transport responsibilities to each strategy. The key advantage of PASE over prior art is that it achieves both near-optimal performance as well as deployment friendliness. PASE does not require any changes in network switches (hardware or software); yet, it achieves comparable, or even better, performance than the state-of-the-art protocols (such as pFabric) that require changes to network elements. Our evaluation results show that the PASE performs well for a wide range of application workloads and network settings.
Ali Munir, Ghufran Baig, Syed Mohammad Irteza, Ihsan Ayyub Qazi, Alex X. Liu, Fahad R. Dogar
IEEE/ACM Trans. Netw.1
2016 Network Scheduling Aware Task Placement in Datacenters
abstract
To improve the performance of data-intensive applications, existing datacenter schedulers optimize either the placement of tasks or the scheduling of network flows. The task scheduler strives to place tasks close to their input data (i.e., maximize data locality) to minimize network traffic, while assuming fair sharing of the network. The network scheduler strives to finish flows as quickly as possible based on their sources and destinations determined by the task scheduler, while the scheduling is based on flow properties (e.g., size, deadline, and correlation) and not bound to fair sharing. Inconsistent assumptions of the two schedulers can compromise the overall application performance. In this paper, we propose NEAT, a task scheduling framework that leverages information from the underlying network scheduler to make task placement decisions. The core of NEAT is a task completion time predictor that estimates the completion time of a task under given network condition and a given network scheduling policy. NEAT leverages the predicted task completion times to minimize the average completion time of active tasks. Evaluation using ns2 simulations and real-testbed shows that NEAT improves application performance by up to 3.7x for the suboptimal network scheduling policies and up to 30% for the optimal network scheduling policy.
Ali Munir, Ting He 0001, Ramya Raghavendra, Franck Le, Alex X. Liu
CoNEXT1
2016 OpenFunction: An extensible data plane abstraction protocol for platform-independent software-defined middleboxes
abstract
We propose OpenFunction, an extensible data plane abstraction protocol for platform-independent software-defined middleboxes. The main challenge is how to abstract packet operations, flow states and event generations with elements. The key decision of OpenFunction is: actions/states/events operations should be defined in a uniform pattern and independent from each other. We implemented a working SDM system including one OpenFunction controller and OpenFunction boxes based on Netmap, DPDK and FPGA to verify OpenFunction abstraction.
Chen Tian 0001, Alex X. Liu, Ali Munir
ICNP3
2014 Friends, not foes: synthesizing existing transport strategies for data center networks
abstract
Many data center transports have been proposed in recent times (e.g., DCTCP, PDQ, pFabric, etc). Contrary to the common perception that they are competitors (i.e., protocol A vs. protocol B), we claim that the underlying strategies used in these protocols are, in fact, complementary. Based on this insight, we design PASE, a transport framework that synthesizes existing transport strategies, namely, self-adjusting endpoints (used in TCP style protocols), innetwork prioritization (used in pFabric), and arbitration (used in PDQ). PASE is deployment friendly: it does not require any changes to the network fabric; yet, its performance is comparable to, or better than, the state-of-the-art protocols that require changes to network elements (e.g., pFabric). We evaluate PASE using simulations and testbed experiments. Our results show that PASE performs well for a wide range of application workloads and network settings.
Ali Munir, Ghufran Baig, Syed Mohammad Irteza, Ihsan Ayyub Qazi, Alex X. Liu, Fahad R. Dogar
SIGCOMM1
2013 On achieving low latency in data centers
abstract
Today's data centers face extreme challenges in providing low latency for online services such as web search, social networking, and recommendation systems. Achieving low latency is important as it impacts user experience, which in turn impacts operator revenue. However, most current congestion control protocols approximate Processor Sharing (PS), which is known to be sub-optimal for minimizing latency. In this paper, we propose Router Assisted Capacity Sharing (RACS), a data center transport protocol that minimizes flow completion times by approximating the Shortest Remaining Processing Time (SRPT) scheduling policy, which is known to be optimal, in a distributed manner. With RACS, flows are assigned weights which determine their relative priority and thus the rate assigned to them. By changing these weights, RACS can approximate a range of scheduling disciplines. Through extensive ns-2 simulations, we demonstrate that RACS outperforms TCP, DCTCP, and RCP in data center environments. In particular, it improves completion times by up to 95% over TCP, 88% over DCTCP, and 80% over RCP. Our results also show that RACS can outperform deadline-aware transport protocols for typical data center workloads.
Ali Munir, Ihsan Ayyub Qazi, Saad B. Qaisar
ICC1
2013 Minimizing flow completion times in data centers
abstract
For provisioning large-scale online applications such as web search, social networks and advertisement systems, data centers face extreme challenges in providing low latency for short flows (that result from end-user actions) and high throughput for background flows (that are needed to maintain data consistency and structure across massively distributed systems). We propose L2DCT, a practical data center transport protocol that targets a reduction in flow completion times for short flows by approximating the Least Attained Service (LAS) scheduling discipline, without requiring any changes in application software or router hardware, and without adversely affecting the long flows. L2DCT can co-exist with TCP and works by adapting flow rates to the extent of network congestion inferred via Explicit Congestion Notification (ECN) marking, a feature widely supported by the installed router base. Though L2DCT is deadline unaware, our results indicate that, for typical data center traffic patterns and deadlines and over a wide range of traffic load, its deadline miss rate is consistently smaller compared to existing deadline-driven data center transport protocols. L2DCT reduces the mean flow completion time by up to 50% over DCTCP and by up to 95% over TCP. In addition, it reduces the completion for 99th percentile flows by 37% over DCTCP. We present the design and analysis of L2DCT, evaluate its performance, and discuss an implementation built upon standard Linux protocol stack.
Ali Munir, Ihsan Ayyub Qazi, Zartash Afzal Uzmi, Aisha Mushtaq, Saad N. Ismail, M. Safdar Iqbal, Basma Khan
INFOCOM1
2012 Performance Analysis of WiMAX Best Effort and ertPS Service Classes for Video Transmission
Hassan Abid, Haroon Raja, Ali Munir, Jaweria Amjad, Aliya Mazhar, Dong-Young Lee
ICCSA (3)3
2012 A Genetic Algorithm Assisted Resource Management Scheme for Reliable Multimedia Delivery over Cognitive Networks
Ali Munir, Saad B. Qaisar, Junaid Qadir 0001
ICCSA (3)2