Ahmed Alquraan

dblp:228/0283 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
4since 2021 · last 2026
0000-0002-1445-9619ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Distributed systems · 57% Storage systems · 23% Cloud and datacenter computing · 15%
Computer networks
3 papers
Software-defined and programmable networks · 40% Datacenter networks · 30% Routing and switching · 30%

Topics — the 22 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
fault tolerance
1.942023
Partial Network Partitioning · ACM Trans. Comput. Syst. 2023
Scalable, NearZero Loss Disaster Recovery for Distributed Data Stores · Proc. VLDB Endow. 2020
Toward a Generic Fault Tolerance Technique for Partial Network Partitioning · OSDI 2020
Storage systems
key-value storage
1.832024
LoLKV: The Logless, Linearizable, RDMA-based Key-Value Storage System · NSDI 2024
Accelerating Reads With In-Network Consistency-Aware Load Balancing · IEEE/ACM Trans. Netw. 2022
The Network-Integrated Storage System · IEEE Trans. Parallel Distributed Syst. 2020
Cloud and datacenter computing › resource management › resource pooling
resource pool management
1.012026
DROPS: Managing Serverless Resource Pools in Microsoft Azure Functions · EuroSys 2026
Cloud and datacenter computing
serverless computing
1.012026
DROPS: Managing Serverless Resource Pools in Microsoft Azure Functions · EuroSys 2026
Distributed systems › consistency models
linearizability
0.922024
LoLKV: The Logless, Linearizable, RDMA-based Key-Value Storage System · NSDI 2024
Scalable, NearZero Loss Disaster Recovery for Distributed Data Stores · Proc. VLDB Endow. 2020
Distributed systems › fault tolerance › failure models
network partitioning
0.822020
Toward a Generic Fault Tolerance Technique for Partial Network Partitioning · OSDI 2020
An Analysis of Network-Partitioning Failures in Cloud Systems · OSDI 2018
Storage systems › key-value storage
RDMA-based key-value store
0.812024
LoLKV: The Logless, Linearizable, RDMA-based Key-Value Storage System · NSDI 2024
Distributed systems
consensus
0.722022
Accelerating Reads With In-Network Consistency-Aware Load Balancing · IEEE/ACM Trans. Netw. 2022
Scalable, NearZero Loss Disaster Recovery for Distributed Data Stores · Proc. VLDB Endow. 2020
Distributed systems
distributed coordination
0.712023
Partial Network Partitioning · ACM Trans. Comput. Syst. 2023
Distributed systems › group communication
membership management
0.712023
Partial Network Partitioning · ACM Trans. Comput. Syst. 2023
Hardware reliability and fault tolerance › network fault tolerance
network partition tolerance
0.712023
Partial Network Partitioning · ACM Trans. Comput. Syst. 2023
Software-defined and programmable networks
programmable data plane
0.612022
Accelerating Reads With In-Network Consistency-Aware Load Balancing · IEEE/ACM Trans. Netw. 2022
Distributed systems › consensus
leader-based consensus
0.612022
Accelerating Reads With In-Network Consistency-Aware Load Balancing · IEEE/ACM Trans. Netw. 2022
Datacenter networks
network-storage co-design
0.412020
The Network-Integrated Storage System · IEEE Trans. Parallel Distributed Syst. 2020
Routing and switching
routing
0.412020
FLAIR: Accelerating Reads with Consistency-Aware Network Routing · NSDI 2020
Distributed systems › fault tolerance › failure recovery
disaster recovery
0.412020
Scalable, NearZero Loss Disaster Recovery for Distributed Data Stores · Proc. VLDB Endow. 2020
Distributed systems › replication › state machine replication
log replication
0.412020
Scalable, NearZero Loss Disaster Recovery for Distributed Data Stores · Proc. VLDB Endow. 2020
Distributed systems
replication
0.412020
Scalable, NearZero Loss Disaster Recovery for Distributed Data Stores · Proc. VLDB Endow. 2020
Storage systems
storage reliability
0.412020
Scalable, NearZero Loss Disaster Recovery for Distributed Data Stores · Proc. VLDB Endow. 2020
Distributed systems › peer-to-peer systems
overlay networks
0.212023
Partial Network Partitioning · ACM Trans. Comput. Syst. 2023
Storage systems › storage performance
read latency
0.112020
FLAIR: Accelerating Reads with Consistency-Aware Network Routing · NSDI 2020
Distributed systems › replication › replicated data management
replicated data stores
0.112020
FLAIR: Accelerating Reads with Consistency-Aware Network Routing · NSDI 2020

Methods — techniques the papers use, named apart from their topics

programmable switch · 1.1p4 · 1.1SLO-based pool sizing · 1.0network routing · 0.9multicast · 0.9overlay routing · 0.7synchronized clocks · 0.4pipelining · 0.4batching · 0.4asynchronous log replication · 0.4
YearPublicationVenuePosition
2026 DROPS: Managing Serverless Resource Pools in Microsoft Azure Functions
abstract
Azure Functions maintains pools of pre-warmed containers to avoid the high container-allocation latency. The size of a pool is important: a pool that is too small leads to high allocation latency, whereas a pool that is too large wastes resources and increases cost. Service providers typically oversize pools to meet service-level objectives (SLOs). Our findings indicate that the cost of maintaining pre-warmed container pools dominates the overall platform cost, motivating the need for effective pool management strategies.
Ahmed Alquraan, Abdelrahman Baba, Rafael Mendes da Silva, Sameh Elnikety, Paul Batum, Hamid Henry Safi, Seth Fine, Samer Al-Kiswany
EuroSys1
2024 LoLKV: The Logless, Linearizable, RDMA-based Key-Value Storage System
Ahmed Alquraan, Sreeharsha Udayashankar, Virendra J. Marathe, Bernard Wong 0001, Samer Al-Kiswany
NSDI1
2023 Partial Network Partitioning
abstract
We present an extensive study focused on partial network partitioning. Partial network partitions disrupt the communication between some but not all nodes in a cluster. First, we conduct a comprehensive study of system failures caused by this fault in 13 popular systems. Our study reveals that the studied failures are catastrophic (e.g., lead to data loss), easily manifest, and are mainly due to design flaws. Our analysis identifies vulnerabilities in core systems mechanisms including scheduling, membership management, and ZooKeeper-based configuration management. Second, we dissect the design of nine popular systems and identify four principled approaches for tolerating partial partitions. Unfortunately, our analysis shows that implemented fault tolerance techniques are inadequate for modern systems; they either patch a particular mechanism or lead to a complete cluster shutdown, even when alternative network paths exist. Finally, our findings motivate us to build Nifty, a transparent communication layer that masks partial network partitions. Nifty builds an overlay between nodes to detour packets around partial partitions. Nifty provides an approach for applications to optimize their operation during a partial partition. We demonstrate the benefit of this approach through integrating Nifty with VoltDB, HDFS, and Kafka.
Basil Alkhatib, Sreeharsha Udayashankar, Sara Qunaibi, Ahmed Alquraan, Mohammed Alfatafta, Wael Al-Manasrah, Alex Depoutovitch, Samer Al-Kiswany
ACM Trans. Comput. Syst.4
2022 Accelerating Reads With In-Network Consistency-Aware Load Balancing
abstract
We present FLAIR, a novel approach for accelerating read operations in leader-based consensus protocols. FLAIR leverages the capabilities of the new generation of programmable switches to serve reads from follower replicas without compromising consistency. The core of the new approach is a packet-processing pipeline that can track client requests and system replies, identify consistent replicas, and at line speed, forward read requests to replicas that can serve the read without sacrificing linearizability. An additional benefit of FLAIR is that it facilitates devising novel consistency-aware load balancing techniques. Following the new approach, we designed FlairKV, a key-value store atop Raft. FlairKV implements the processing pipeline using the P4 programming language. We evaluate the benefits of the proposed approach and compare it to previous approaches using a cluster with a Barefoot Tofino switch. Our evaluation indicates that, compared to state-of-the-art alternatives, the proposed approach can bring significant performance gains: up to 42% higher throughput and 35–97% lower latency for most workloads. Furthermore, our evaluation shows that our novel load balancing techniques can cope with heterogeneous load and hardware to achieve higher performance, and that FLAIR can scale to support large data sets and clusters.
Ibrahim Kettaneh, Ahmed Alquraan, Hatem Takruri, Ali José Mashtizadeh, Samer Al-Kiswany
IEEE/ACM Trans. Netw.2
2020 FLAIR: Accelerating Reads with Consistency-Aware Network Routing
Hatem Takruri, Ibrahim Kettaneh, Ahmed Alquraan, Samer Al-Kiswany
NSDI3
2020 Toward a Generic Fault Tolerance Technique for Partial Network Partitioning
Mohammed Alfatafta, Basil Alkhatib, Ahmed Alquraan, Samer Al-Kiswany
OSDI3
2020 Scalable, NearZero Loss Disaster Recovery for Distributed Data Stores
abstract
This paper presents a new Disaster Recovery (DR) system, called Slogger, that differs from prior works in two principle ways: (i) Slogger enables DR for a linearizable distributed data store, and (ii) Slogger adopts the continuous backup approach that strives to maintain a tiny lag on the backup site relative to the primary site, thereby restricting the data loss window, due to disasters, to milliseconds. These goals pose a significant set of challenges related to consistency of the backup site's state, failures, and scalability. Slogger employs a combination of asynchronous log replication, intra-data center synchronized clocks, pipelining, batching, and a novel watermark service to address these challenges. Furthermore, Slogger is designed to be deployable as an "add-on" module in an existing distributed data store with few modifications to the original code base. Our evaluation, conducted on Slogger extensions to a 32-sharded version of LogCabin, an open source key-value store, shows that Slogger maintains a very small data loss window of 14.2 milliseconds which is near the optimal value in our evaluation setup. Moreover, Slogger reduces the length of the data loss window by 50% compared to incremental snapshotting technique without having any performance penalty on the primary data store. Furthermore, our experiments demonstrate that Slogger achieves our other goals of scalability, fault tolerance, and efficient failover to the backup data store when a disaster is declared at the primary data store.
Ahmed Alquraan, Alex Kogan, Virendra J. Marathe, Samer Al-Kiswany
Proc. VLDB Endow.1
2020 The Network-Integrated Storage System
abstract
We present NICE, a key-value storage system design that leverages new software-defined network capabilities to build cluster-based network-efficient storage system. NICE presents novel techniques to co-design network routing and multicast with storage replication, consistency, and load balancing to achieve higher efficiency, performance, and scalability. We implement the NICEKV prototype. NICEKV follows the NICE approach in designing four essential network-centric storage mechanisms: request routing, replication, consistency, and load balancing. Our evaluation shows that the proposed approach brings significant performance gains compared with the current systems design: up to 7× put/get performance improvement, up to 2× reduction in network load, 3× to 9× load reduction on the storage nodes, and the elimination of scalability bottlenecks present in current designs.
Ibrahim Kettaneh, Ahmed Alquraan, Hatem Takruri, Suli Yang, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau, Samer Al-Kiswany
IEEE Trans. Parallel Distributed Syst.2
2018 An Analysis of Network-Partitioning Failures in Cloud Systems
Ahmed Alquraan, Hatem Takruri, Mohammed Alfatafta, Samer Al-Kiswany
OSDI1