VLDB 2026 Research / reviewers in the wild / expert
Ahmed Elmokashfi
dblp:41/3349 · also Ahmed Mustafa Elmokashfi
· DBLP profile ↗
40ranked-venue papers
9as first author
13since 2021 · last 2025
0000-0001-9964-214XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 33 · 8 first-author · 10 since 2021Systems, architecture and hardware · 2Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ADDER: Service-Specific Adaptive Data-Driven Radio Resource Control for Cellular-IoTabstractEnergy-saving methods like Discontinuous Reception (DRX) and Power Save Mode (PSM) are commonly used in Internet of Thing (IoT) applications, allowing for sleep and awake cycle adjustments to save energy. However, understanding and configuring these parameters on devices, especially actuator-type devices, is challenging for IoT service providers. Unlike sensor types, these devices must complete their sleep cycle before responding to infrequent downlink commands, making efficient parameter selection and traffic prediction vital for energy efficiency and command reception. To address this, we present ADDER, a network-side solution leveraging a context-aware traffic predictor. This predictor forecasts downlink arrival probabilities, guiding a deep deterministic policy gradient (DDPG) policymaker that selects energy-saving parameters based on thresholds defined by the IoT service providers. ADDER, leverages contextual information like day of week, hour, weather, holidays, and events, shifting the focus from individual device histories (often erratic) to analyzing broader service traffic patterns. This data-driven strategy enables ADDER to adjust energy-saving settings for the best balance between energy efficiency and latency, customizing to the unique requirements of each service and removing the burden of configuring complex network settings. We observed that ADDER meets latency needs while achieving a $\mathbf{5 . 9 \%}$ reduction in energy consumption for services requiring rapid responses. For applications prioritizing energy conservation, such as irrigation systems and city lighting, ADDER achieves a significant $32.7 \%$ reduction in energy consumption with a slight increase $(\mathbf{9 \%}$) in messages might not meet the strictest latency requirements. To evaluate the consequences of prediction inaccuracies from our predictor, we utilized a real-world shared mobility dataset provided by Austin’s Transportation Department for a case study. Yingjing Wu, Ahmed Elmokashfi, Foivos Michelinakis, Jacobus E. van der Merwe, Shandian Zhe |
WoWMoM | 2 |
| 2024 | Tracking submarine cables in the wildabstractDuring the last ten years, thousands of kilometers of submarine cables have been rolled out to connect regions around the globe and improve intercontinental connectivity. However, while it is relatively easy to get information about the frequent roll-outs of these cables, it is challenging to translate these developments into network information to facilitate networking research. For example, announcements for new submarine cables typically mention landing points and not router IP addresses. With this network information, it is easier to assess the impact of a new submarine cable on end-to-end delays in the connecting regions. In this paper, we investigate the necessary and sufficient conditions to translate public announcements for submarine cables to network information that enables networking research on this topic. We also develop and evaluate a methodology to automatically extract IP-level information for deployed submarine cables and assess their impact on end-to-end performance. Ioana Livadariu, Ahmed Elmokashfi, Georgios Smaragdakis |
Comput. Networks | 2 |
| 2024 | Bottleneck Identification in Cloudified Mobile Networks Based on Distributed TelemetryabstractCloudified mobile networks are expected to deliver a multitude of services with reduced capital and operating expenses. A characteristic example is 5G networks serving several slices in parallel. Such mobile networks, therefore, need to ensure that the SLAs of customised end-to-end sliced services are met. This requires monitoring the resource usage and characteristics of data flows at the virtualised network core, as well as tracking the performance of the radio interfaces and UEs. A centralised monitoring architecture can not scale to support millions of UEs though. This paper, proposes a 2-stage distributed telemetry framework in which UEs act as early warning sensors. After UEs flag an anomaly, a ML model is activated, at network controller, to attribute the cause of the anomaly. The framework achieves 85% F1-score in detecting anomalies caused by different bottlenecks, and an overall 89% F1-score in attributing these bottlenecks. This accuracy of our distributed framework is similar to that of a centralised monitoring system, but with no overhead of transmitting UE-based telemetry data to the centralised controller. The study also finds that passive in-band network telemetry has the potential to replace active monitoring and can further reduce the overhead of a network monitoring system. Mah-Rukh Fida, Azza H. Ahmed, Thomas Dreibholz, Andrés F. Ocampo, Ahmed Elmokashfi, Foivos Michelinakis |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Modeling Variation in Mobile Download Speed in Presence of Missing SamplesabstractA stably fast mobile broadband connectivity is key to customer retention. Mobile networks, however, suffer unpredictability in performance. Analyzing variability in network speed is, therefore, challenging since it tends to exhibit patterns at several time scales. Additionally, frequently monitoring it over time, is costly. In this paper, we analyze speed measurements from 78 stationary probes, spread across Norway. Monitoring was performed thrice per day across the year, to assess performance of the two largest network operators. Despite being unique, the dataset involves a non-trivial extent of missing data. This study investigates the effect of missing data on the extracted performance patterns. We capture patterns with tensor factorizations, that show that missing data at random has a minimal effect on the identified patterns, and that depending upon the determinism of an operator's performance, the acceptable size and structure of missing data varies. Our analysis shows that, for a probe, the difference in speed variation between real and imputed speed values can be around 7% for up to 40% missing data. We also identify that congestion, routine maintenance and sub-optimal network configuration cause high speed variability. These findings can help operators improving their offerings and deciding on optimal performance monitoring frequency. Mah-Rukh Fida, Marie Roald, Evrim Acar, Ahmed Elmokashfi |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | On Large-Scale IP Service Disruptions DependenciesabstractLarge part of the human activities has become digitalized as many of the daily activities rely on the Internet. However, this critical infrastructure is subject to disruption that can be caused by malicious actors or misconfiguration. Network operators are expected to implement community driven best-practices that contribute to the Internet's resilience. In practice many operators fail to implement them due to the high complexity of these tasks. Consequently, large-scale disruption of IP services can impact numerous customers through a cascading effect. In this work, we assess how three major Internet disruption impact their connectivity. We found that these disruptions cascade from the affected operators towards large part of their customers. However, customers that follow best practices have a short recovery time and rely on rerouted paths through a myriad of topologically close backups links. Alfred Arouna, Ioana Livadariu, Azan Latif Khanyari, Ahmed Elmokashfi |
CNSM | 4 |
| 2023 | PRINCIPIA: Opportunistic CPU and CPU-shares Allocation for Containerized Virtualization in Mobile Edge ComputingabstractLeveraging virtualization technology, Mobile Edge Computing (MEC) deploys multiple services with different execution time requirements running as isolated processes. For instance, both real-time (RT) and non-RT applications may be (are) running on the same infrastructure using containerized virtualization. Nevertheless, sharing resources (e.g., CPU) with collocated workloads could impact the RT performance of RT applications. This paper presents PRINCIPIA, a dynamic CPU and CPU-shares allocation mechanism that opportunistically enables non-RT applications to run on underutilized CPUs while providing RT guarantees to RT applications. By monitoring MEC’s system metrics like processor’s CPU utilization and container’s CPU usage, PRINCIPIA dynamically allocates both CPU and CPU-shares to containers running non-RT applications aiming at opportunistically exploiting underutilized CPUs by containers running RT applications. We evaluate PRINCIPIA on a small-scale MEC server which uses containerized virtualization along with Linux RT Kernel to deploy both RT and non-RT applications. Our findings show that PRINCIPIA mitigates the impact on the RT performance of RT applications providing bounded processing latency in comparison with the default host Kernel scheduler. Andrés F. Ocampo, Mah-Rukh Fida, Juan Felipe Botero, Ahmed Elmokashfi, Haakon Bryhni |
NOMS | 4 |
| 2023 | AI Anomaly Detection for Cloudified Mobile Core ArchitecturesabstractIT systems monitoring is a crucial process for managing and orchestrating network resources, allowing network providers to rapidly detect and react to most impediment causing network degradation. However, the high growth in size and complexity of current operational networks (2022) demands new solutions to process huge amounts of data (including alarms) reliably and swiftly. Further, as the network becomes progressively more virtualized, the hosting of NFV on cloud environments adds a magnitude of possible bottlenecks outside the control of the service owners. In this paper, we propose two deep learning anomaly detection solutions that leverage service exposure and apply it to automate the detection of service degradation and root cause discovery in a cloudified mobile network that is orchestrated by ETSI OSM. A testbed is built to validate these AI models. The testbed collects monitoring data from the OSM monitoring module, which is then exposed to the external AI anomaly detection modules, tuned to identify the anomalies and the network services causing them. The deep learning solutions are tested using various artificially induced bottlenecks. The AI solutions are shown to correctly detect anomalies and identify the network components involved in the bottlenecks, with certain limitations in a particular type of bottlenecks. A discussion of the right monitoring tools to identify concrete bottlenecks is provided. Foivos Michelinakis, Joan S. Pujol Roig, Sara Malacarne, Min Xie 0006, Thomas Dreibholz, Sayantini Majumdar, Wint Yi Poe, Georgios Patounas, Carmen Guerrero, Ahmed Elmokashfi, Vasileios Theodorou |
IEEE Trans. Netw. Serv. Manag. | 10 |
| 2023 | Opportunistic CPU Sharing in Mobile Edge Computing Deploying the Cloud-RANabstractLeveraging virtualization technology, Cloud-RAN deploys multiple virtual Base Band Units (vBBUs) along with collocated applications on the same Mobile Edge Computing (MEC) server. However, the performance of real-time (RT) applications such as the vBBU could potentially be impacted by sharing computing resources with collocated workloads. To address this challenge, this paper presents a dynamic CPU sharing mechanism, specifically designed for containerized virtualization in MEC servers, that hosts both RT and non-RT general-purpose applications. Initially, the CPU sharing problem in MEC servers is formulated as a Mixed-Integer Programming (MIP). Then, we present an algorithmic solution that breaks down the MIP into simpler subproblems that are then solved using efficient, constant factor heuristics. We assessed the performance of this mechanism against instances of a commercial solver. Further, via a small-scale testbed, we assessed various CPU sharing mechanisms and their effectiveness in reducing the impact of CPU sharing on RT application processing performance. Our findings indicate that our CPU sharing mechanism reduces the worst-case execution time by more than 150% compared to the default host RT-Kernel approach. This evidence is strengthened when evaluating this mechanism within Cloud-RAN, in which vBBUs share resources with collocated applications on a MEC server. Using our CPU sharing approach, the vBBU’s scheduling latency decreases by up to 21% in comparison with the host RT-Kernel. Andrés F. Ocampo, Mah-Rukh Fida, Juan Felipe Botero, Ahmed Elmokashfi, Haakon Bryhni |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2022 | RCAD: Real-time Collaborative Anomaly Detection System for Mobile Broadband NetworksabstractThe rapid increase in mobile data traffic and the number of connected devices and applications in networks is putting a significant pressure on the current network management approaches that heavily rely on human operators. Consequently, an automated network management system that can efficiently predict and detect anomalies is needed. In this paper, we propose, RCAD, a novel distributed architecture for detecting anomalies in network data forwarding latency in an unsupervised fashion. RCAD employs the hierarchical temporal memory (HTM) algorithm for the online detection of anomalies. It also involves a collaborative distributed learning module that facilitates knowledge sharing across the system. We implement and evaluate RCAD on real world measurements from a commercial mobile network. RCAD achieves over 0.7 F-1 score significantly outperforming current state-of-the-art methods. Azza H. Ahmed, Michael Riegler 0001, Steven Alexander Hicks, Ahmed Elmokashfi |
KDD | 4 |
| 2022 | Deep reinforcement learning-based control framework for radio access networksabstractNetwork performance optimization represents one of the major challenges for mobile network operators, especially with the increasingly use cases that have diverse performance expectations. In this work we propose a novel control framework that maximizes radio resources utilization and minimizes performance degradation in the most challenging part of cellular architecture that is the radio access network (RAN). Based on deep reinforcement learning, we devise two control schemes: centralized and distributed, respectively. Using extensive discrete event simulations, we confirm that our proposed control framework succeeds in optimizing radio resources utilization while minimizing service level agreement (SLA) violations in a multi-slice RAN. Azza H. Ahmed, Ahmed Elmokashfi |
MobiCom | 2 |
| 2022 | ICRAN: Intelligent Control for Self-Driving RAN Based on Deep Reinforcement LearningabstractMobile networks are increasingly expected to support use cases with diverse performance expectations at a very high level of reliability. These expectations imply the need for approaches that timely detect and correct performance problems. However, current approaches often focus on optimizing a single performance metric. Here, we aim to address this gap by proposing a novel control framework that maximizes radio resources utilization and minimizes performance degradation in the most challenging part of cellular architecture that is the radio access network (RAN). We devise a method called Intelligent Control for Self-driving RAN (ICRAN) which involves two deep reinforcement learning based approaches that control the RAN in a centralized and a distributed way, respectively. ICRAN defines a dual-objective optimization goals that are achieved through a set of diverse control actions. Using extensive discrete event simulations, we confirm that ICRAN succeeds in achieving its design goals, showing a greater edge over competing approaches. We believe that ICRAN is implementable and can serve as an important point on the way to realizing self-driving mobile networks. Azza H. Ahmed, Ahmed Elmokashfi |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2022 | Measuring and Localising Congestion in Mobile Broadband NetworksabstractMobile broadband networks, although increasingly popular, suffer large fluctuations in performance. Download speeds can drop by 50% or more during peak hours. Hence, understanding and dissecting the causes of these fluctuations is central to improving current and future networks. In this paper, we propose a congestion detection and localisation method, Q-TSLP, that combines and extends the two state-of-the-art congestion detection tools: Q-Probe and TSLP. Q-Probe monitors patterns in packet arrivals, while TSLP tracks shifts in RTT to detect bottleneck at different segments of an end-to-end path. QProbe can attribute congestion, at a very coarse level, to either radio or non-radio related. TSLP on the other hand cannot pinpoint radio related congestion. Q-TSLP provides a per-hop congestion attribution thus addressing these limitations. To this end, we build two small scale LTE testbeds and experiment with a series of congestion scenarios. These controlled experiments show that apart from correct congestion localisation to finer granularity, the detection accuracy improves significantly with Q-TSLP, up to 100% in some cases. We then run a three-month long measurement campaign of congestion over two commercial operators in Norway. Overall, we run 17 million tests from a large number of geographically distributed probes. We find that both operators suffer congestion at different parts of the network. Our findings indicate that apart from mobile radio access, a non-trivial fraction of cases is related to congested mobile operator and Internet paths beyond the mobile network core. These findings hint that operators may need significant infrastructure upgrades to cope with potential 5G traffic volumes. Mah-Rukh Fida, Andrés F. Ocampo, Ahmed Elmokashfi |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2021 | Dissecting Energy Consumption of NB-IoT Devices Empiricallyabstract3GPP has recently introduced NB-IoT, a new mobile communication standard offering a robust and energy-efficient connectivity option to the rapidly expanding market of the Internet-of-Things (IoT) devices. To unleash its full potential, end devices are expected to work in a plug-and-play fashion, with zero or minimal configuration of parameters, still exhibiting excellent energy efficiency. We performed the most comprehensive set of empirical measurements with commercial IoT devices and different operators to date, quantifying the impact of several parameters to energy consumption. Our findings prove that parameters' settings do impact energy consumption, so proper configuration is necessary. We shed light on this aspect by first illustrating how the nominal standard operational modes map into real current consumption patterns of NB-IoT devices. Furthermore, we investigated which device-reported metadata metrics better reflected performance and implemented an algorithm to automatically identify device state in the current time-series logs. We worked with two major western European operators to provide a measurement-driven analysis of energy consumption and network performance of two popular NB-IoT boards under different parameter configurations. We observed that energy consumption is mostly affected by the paging interval in connected state, set by the base station. However, not all operators correctly implement such settings. Furthermore, under the default configuration, energy consumption in not strongly affected by packet size nor by signal quality, unless it is extremely bad. Our observations indicate that simple modifications to the default parameters' settings can yield great energy savings. Foivos Michelinakis, Anas Saeed Al-Selwi, Martina Capuzzo, Andrea Zanella, Kashif Mahmood, Ahmed Elmokashfi |
IEEE Internet Things J. | 6 |
| 2020 | Evaluating the Cloud-RAN architecture: functional splitting and switched Ethernet XhaulabstractThe Cloud-RAN architecture is a key enabler to building future mobile networks in a flexible and cost-efficient way. For instance, switched Ethernet is a prime candidate for mobile transport networks (Xhaul), due to its flexibility, ubiquity, and cost-effectiveness. Understanding its performance under different network configurations would allow concluding about its appeal for Cloud-RAN. On the other hand, evaluating resource sharing mechanisms is relevant to put in place best solutions to host multiple virtual Base Band Units (vBBUs) into the same compute infrastructure. This paper assesses the feasibility of using a switched Ethernet Xhaul, by instantiating two vBBUs using different functional splits. Moreover, this paper evaluates two mechanisms for sharing network interface cards (NIC) in a general purpose server (GPS) hosting vBBUs. Our results point to a marginal performance degradation caused by the switched Ethernet Xhaul and the NIC sharing mechanisms. Such deviations could be seen from the increase in average and maximum Jitter and RTT results. Andrés F. Ocampo, Mah-Rukh Fida, Ahmed Elmokashfi, Haakon Bryhni |
CNSM | 3 |
| 2020 | Performance bottlenecks identification in cloudified mobile networksabstractThe recent trend towards cloudifying mobile networks brings more flexibility and shortens deployment times. However, it results in an architecture spanning several independent layers from the bare metal to the service level thus complicating troubleshooting and service assurance. In this work, we experimentally explore whether we can accurately and efficiently identify bottlenecks across the different locations of the network and layers of the cloudified architecture. Our findings confirm the complexity of this task and lead us to promising solutions through the use of Machine Learning. Georgios Patounas, Xenofon Foukas, Ahmed Elmokashfi, Mahesh K. Marina |
MobiCom | 3 |
| 2020 | An agent-based model of IPv6 adoption
Ioana Livadariu, Ahmed Elmokashfi, Amogh Dhamdhere |
Networking | 2 |
| 2020 | On the usability of transport protocols other than TCP: A home gateway and internet path traversal study
Runa Barik, Michael Welzl, Gorry Fairhurst, Ahmed Elmokashfi, Thomas Dreibholz, Stein Gjessing |
Comput. Networks | 4 |
| 2020 | Characterization and Identification of Cloudified Mobile Network Performance BottlenecksabstractThis study is a first attempt to experimentally explore the range of performance bottlenecks that 5G mobile networks can experience. To this end, we leverage a wide range of measurements obtained with a prototype testbed that captures the key aspects of a cloudified mobile network. We investigate the relevance of the metrics and a number of approaches to accurately and efficiently identify bottlenecks across the different locations of the network and layers of the system architecture. Our findings validate the complexity of this task in the multi-layered architecture and highlight the need for novel monitoring approaches that intelligently fuse metrics across network layers and functions. In particular, we find that distributed analytics performs reasonably well both in terms of bottleneck identification accuracy and incurred computational and communication overhead. Georgios Patounas, Xenofon Foukas, Ahmed Elmokashfi, Mahesh K. Marina |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2019 | Multiway Reliability Analysis of Mobile Broadband NetworksabstractUnderstanding and characterizing the reliability of a mobile broadband network is a challenging task due to the presence of a multitude of root causes that operate at different temporal and spatial scales. This, in turn, limits the use of classical statistical methods for characterizing the mobile network's reliability. We propose leveraging tensor factorizations, a well-established data mining method, to address this challenge. We represent a year-long time series of outages, from two mobile operators as multi-way arrays, and demonstrate how tensor factorizations help in extracting the outage patterns at various time-scales, making it easy to locate possible root causes. Unlike traditional methods of time series analysis, tensor factorizations provide a compact and interpretable picture of outages. Mah-Rukh Fida, Evrim Acar, Ahmed Elmokashfi |
Internet Measurement Conference | 3 |
| 2019 | Tracking the deployment of IPv6: Topology, routing and performance
Siyuan Jia, Matthew J. Luckie, Bradley Huffaker, Ahmed Elmokashfi, Emile Aben, K. C. Claffy, Amogh Dhamdhere |
Comput. Networks | 4 |
| 2019 | On the utility of unregulated IP DiffServ Code Point (DSCP) usage by end systemsabstractDiffServ was designed to implement service provider quality of service (QoS) policies, where routers change and react upon the DiffServ Code Point (DSCP) in the IP header. However, nowadays, applications are beginning to directly set the DSCP themselves, in the hope that this will yield a more appropriate service for their respective video, audio and data streams. WebRTC is a prime example of such an application. We present measurements, for both IPv4 and IPv6, of what happens to DSCP values along Internet paths after an end system has set them without any prior agreement between a customer and a service provider. We find that the DSCP is often changed or zeroed along the path, but detrimental effects from using the DSCP are extremely rare; moreover, DSCP values sometimes remain intact (potentially having an effect on traffic) for several AS hops. This positive result motivates an analysis of the potential latency impact from such DSCP usage, for which we present the first measurement results. We find that routers at approximately 3% of more than 100,000 links differentiate between the WebRTC DSCP values (EF, AF42 and CS1) and consistently reduce delay in comparison with probes carrying a zero value (CS0) under congestion. In contrast, routers at around 2% of these links increase the delay by a comparable amount under congestion, uniformly for EF, AF42 and CS1. Runa Barik, Michael Welzl, Ahmed Elmokashfi, Thomas Dreibholz, Safiqul Islam, Stein Gjessing |
Perform. Evaluation | 3 |
| 2018 | WiFi and Multiple Interfaces: Adequate for Virtual Reality?abstractIn this paper, we investigate whether IEEE 802.11ac WiFi can support VR applications. To this end we conduct a controlled study of WiFi performance in an indoor setting. Our measurements reveal that WiFi transmissions suffer from high latency and jitter, which makes WiFi systems inadequate for VR applications. For example, the round-trip delay can be as high as 228ms, and more than 24.2% packets experience jitter higher than 1ms. To locate the root cause of the high latency and jitter, we dissect the network stack layer by layer and find that the main culprit is the wireless channel transmission time. To reduce the channel transmission time, we propose using multiple network interfaces running on non-overlapping channels. By using only two interfaces, we (1) reduce the median round-trip delay by 28.6% and jitters of higher than 1ms by 11.5% compared to the best single interface in UDP transmissions, and (2) reduce the median round-trip delay by 38.9% in TCP transmissions. We believe that this paper sheds some light on whether we can make today's WiFi systems VR-ready by using multiple interfaces. Huanle Zhang, Ahmed Elmokashfi, Prasant Mohapatra |
ICPADS | 2 |
| 2018 | Inferring Carrier-Grade NAT Deployment in the WildabstractGiven the increasing scarcity of IPv4 addresses, network operators are resorting to measures to expand their address pool or prolong the life of existing addresses. One such approach is Carrier-Grade NAT (CGN), where many end-users in a network share a single public IPv4 address. There is limited data about the prevalence of CGN, despite the implications on performance, security, and ultimately, the adoption of IPv6. In this work, we present passive measurement-based techniques for detecting CGN deployments across the entire Internet, without the requirement of access to machines behind a CGN. Specifically, we identify patterns in how client IP addresses are observed at M-Lab servers and at the UCSD network telescope to infer whether those clients are behind a CGN. We apply our methods on data collected from 2014 to 2016. We find that CGN deployment is increasing rapidly. Overall, we infer that 4.1K autonomous systems are deploying CGN, 6 times the number inferred by the most recent studies. Ioana Livadariu, Karyn Benson, Ahmed Elmokashfi, Amogh Dhamdhere, Alberto Dainotti |
INFOCOM | 3 |
| 2017 | Adding the Next Nine: An Investigation of Mobile Broadband Networks AvailabilityabstractThe near ubiquitous availability and success of mobile broadband networks has motivated verticals that range from public safety communication to intelligent transportation systems and beyond to consider choosing them as the communication mean of choice. Several of these verticals, however, expect high availability of multiple nines. This paper leverages end-to-end measurements to investigate the potential of current mobile broadband networks to support these expectations. We conduct a large-scale measurement study of network availability in four networks in Norway. This study is based on three years of measurements from hundreds of stationary measurement nodes and several months of measurements from four mobile nodes. We find that the mobile network centralized architecture and infrastructure sharing between operators are responsible for a non-trivial fraction of network failures. Most episodes of degraded availability, however, are uncorrelated. We also find that using two networks simultaneously can result in more than five nines of availability for stationary nodes and three nines of availability for mobile nodes. Our findings point to potential avenues for enhancing the availability of future mobile networks. Ahmed Elmokashfi, Dziugas Baltrunas |
MobiCom | 1 |
| 2017 | On IPv4 transfer markets: Analyzing reported transfers and inferring transfers in the wild
Ioana Livadariu, Ahmed Elmokashfi, Amogh Dhamdhere |
Comput. Commun. | 2 |
| 2016 | Characterizing IPv6 control and data plane stabilityabstractEnd-to-end IPv6 performance is a factor that can influence IPv6 adoption. The stability of IPv6 - both in the control and data plane - is an important determinant of end-to-end performance, as it influences packet loss, network latency, and hence application performance. In this paper we compare stability and performance measurements from the control and data plane in IPv6 and IPv4. To study control plane stability, we use BGP feeds from five dual-stacked vantage points to measure routing dynamics towards IPv4 and IPv6 destinations. To study data plane stability, we probe dual-stacked webservers in 629 target ASes to determine the availability, RTT performance and RTT stability of paths toward these targets. In both control and data plane experiments IPv6 exhibited less stability than IPv4. In the control plane, most routing dynamics were generated by a small fraction of pathological unstable prefixes. In the data-plane, episodes of unavailability were longer on IPv6 than on IPv4. We found evidence of correlated performance degradation over IPv4 and IPv6 caused by shared infrastructure. Ioana Livadariu, Ahmed Elmokashfi, Amogh Dhamdhere |
INFOCOM | 2 |
| 2016 | The good, the bad and the implications of profiling mobile broadband coverage
Andra Lutu, Yuba Raj Siwakoti, Özgü Alay, Dziugas Baltrunas, Ahmed Elmokashfi |
Comput. Networks | 5 |
| 2015 | Dissecting packet loss in mobile broadband networks from the edgeabstractThis paper demonstrates that end-to-end active measurements can give invaluable insights into the nature and characteristics of packet loss in cellular networks. We conduct a large-scale measurement study of packet loss in four UMTS networks. The study is based on active measurements from hundreds of measurement nodes over a period of one year. We find that a significant fraction of loss occurs during pathological and normal Radio Resource Control (RRC) state transitions. The remaining loss exhibits pronounced diurnal patterns and shows a relatively strong correlation between geographically diverse measurement nodes. Our results indicate that the causes of a significant part of the remaining loss lie beyond the radio access network. Dziugas Baltrunas, Ahmed Elmokashfi, Amund Kvalbein |
INFOCOM | 2 |
| 2014 | Measuring the Reliability of Mobile Broadband NetworksabstractMobile broadband networks play an increasingly important role in society, and there is a strong need for independent assessments of their robustness and performance. A promising source of such information is active end-to-end measurements. It is, however, a challenging task to go from individual measurements to an assessment of network reliability, which is a complex notion encompassing many stability and performance related metrics. This paper presents a framework for measuring the user-experienced reliability in mobile broadband networks. We argue that reliability must be assessed at several levels, from the availability of the network connection to the stability of application performance. Based on the proposed framework, we conduct a large-scale measurement study of reliability in 5 mobile broadband networks. The study builds on active measurements from hundreds of measurement nodes over a period of 10 months. The results show that the reliability of mobile broadband networks is lower than one could hope: more than 20% of connections from stationary nodes are unavailable more than 10 minutes per day. There is, however, a significant potential for improving robustness if a device can connect simultaneously to several networks. We find that in most cases, our devices can achieve 99.999% ("five nines") connection availability by combining two operators. We further show how both radio conditions and network configuration play important roles in determining reliability, and how external measurements can reveal weaknesses and incidents that are not always captured by the operators' existing monitoring tools. Dziugas Baltrunas, Ahmed Elmokashfi, Amund Kvalbein |
Internet Measurement Conference | 2 |
| 2014 | The Nornet Edge platform for mobile broadband measurementsabstractWe present Nornet Edge (NNE), a dedicated infrastructure for measurements and experimentation in mobile broadband networks. NNE is unprecedented in size, consisting of more than 400 measurement nodes geographically distributed all over Norway. Each measurement node is a Linux-based embedded computer, and is connected to multiple mobile broadband providers. In addition, NNE includes an extensive backend system for deploying and managing experiments and collecting data. NNE makes it possible to run long-term measurement experiments to assess and compare quality and performance across different network operators on a national scale. Particular focus is put on allowing experiments to run in parallel on multiple network connections, and on collecting rich context information related to the experiments. In this paper we give a detailed presentation of NNE, and describe three different measurement experiments that illustrate how the infrastructure can be used. We also provide a roadmap for further development of NNE. Amund Kvalbein, Dziugas Baltrunas, Kristian Evensen, Ahmed Elmokashfi, Simone Ferlin |
Comput. Networks | 5 |
| 2013 | Geography matters: building an efficient transport network for a better video conferencing experienceabstractSome network applications have requirements that exceed the service levels offered by the best-effort Internet. Several network-layer Quality of Service architectures with extended service levels have been designed, but the massive scale and distributed nature of the Internet have prohibited their wide deployment. It now seems clear that the special needs of demanding applications must be met through other approaches. This paper describes how incrementally-deployed innovative solutions at the network level can contribute to an improved service for a particular type of applications, namely high-quality, wide-area video conferencing. We have built a global IP network, aiming to give users a better video conferencing experience primarily through packet loss reduction in transport networks. The key concept in our approach is a well-provisioned network-layer overlay, combined with geography-based "cold potato" BGP routing. Through an extensive set of experiments we show how our design choices impact routing and data plane behavior in the network, and demonstrate that we are able to significantly reduce packet loss compared to wide-area transport through global transit providers. Ahmed Elmokashfi, Eugene Myakotnykh, Jan Marius Evang, Amund Kvalbein, Tarik Cicic |
CoNEXT | 1 |
| 2013 | A first look at IPv4 transfer marketsabstractIn February 2011 the Internet Assigned Numbers Authority (IANA) exhausted its free pool of IPv4 addresses, and the regional registries (RIRs) have started to run out of IPv4 addresses as well. As RIRs have started rationing allocations, IPv4 transfer markets have emerged as a new mechanism to acquire IPv4 addresses. Barring a few high-profile exceptions, IPv4 transfers have largely flown under the radar. In this work, we use the lists of transfers published by three RIRs to characterise the transfer market - the types of players involved, the sizes and characteristics of transferred address blocks, and the visibility of transferred address blocks in the routing table before and after the transfer. Next, we take first steps toward detecting address transfers using BGP data from the Routeviews and RIPE repositories from 2004-2013. We identify reasons why legitimate changes in prefix origin could be mistakenly inferred to be transfers, and implement a series of 10 filters that remove 86% of candidate transfers. Our results indicate that BGP-based detection of transfers is prone to false positives due to significant noise in BGP data, while some transfers remain undetectable as they involve non-BGP speakers. We describe some additional data sources and analysis techniques that may help reveal an opaque market for IPv4 address block transfers. Ioana Livadariu, Ahmed Elmokashfi, Amogh Dhamdhere, K. C. Claffy |
CoNEXT | 2 |
| 2012 | Measuring the deployment of IPv6: topology, routing and performanceabstractWe use historical BGP data and recent active measurements to analyze trends in the growth, structure, dynamics and performance of the evolving IPv6 Internet, and compare them to the evolution of IPv4. We find that the IPv6 network is maturing, albeit slowly. While most core Internet transit providers have deployed IPv6, edge networks are lagging. Early IPv6 network deployment was stronger in Europe and the Asia-Pacific region, than in North America. Current IPv6 network deployment still shows the same pattern. The IPv6 topology is characterized by a single dominant player -- Hurricane Electric -- which appears in a large fraction of IPv6 AS paths, and is more dominant in IPv6 than the most dominant player in IPv4. Routing dynamics in the IPv6 topology are largely similar to those in IPv4, and churn in both networks grows at the same rate as the underlying topologies. Our measurements suggest that performance over IPv6 paths is comparable to that over IPv4 paths if the AS-level paths are the same, but can be much worse than IPv4 if the AS-level paths differ. Amogh Dhamdhere, Matthew J. Luckie, Bradley Huffaker, K. C. Claffy, Ahmed Elmokashfi, Emile Aben |
Internet Measurement Conference | 5 |
| 2012 | Characterizing Delays in Norwegian 3G Networks
Ahmed Elmokashfi, Amund Kvalbein, Kristian Evensen |
PAM | 1 |
| 2012 | BGP Churn Evolution: A Perspective From the CoreabstractThe scalability limitations of BGP have been a major concern lately. An important aspect of this issue is the rate of routing updates (churn) that BGP routers must process. This paper presents an analysis of the evolution of churn in four networks at the backbone of the Internet over a period of seven years and eight months, using BGP update traces from the RouteViews project. The churn rate varies widely over time and between networks. Instead of descriptive “black-box” statistical analysis, we take an exploratory data analysis approach attempting to understand the reasons behind major observed characteristics of the churn time series. We find that duplicate announcements are a major churn contributor, responsible for most large spikes. Remaining spikes are mostly caused by routing incidents that affect a large number of prefixes simultaneously. More long-term intense periods of churn, on the other hand, are caused by misconfigurations or other special events at or close to the monitored autonomous system (AS). After filtering pathologies and effects that are not related to the long-term evolution of churn, we analyze the remaining “baseline” churn and find that it is increasing at a rate that is similar to the growth of the number of ASs. Ahmed Elmokashfi, Amund Kvalbein, Constantinos Dovrolis |
IEEE/ACM Trans. Netw. | 1 |
| 2011 | On Update Rate-Limiting in BGPabstractIn order to reduce the number of BGP updates that routers need to process, it is common to rate-limit such updates using a timer that specifies the minimum time between two consecutive updates for a given destination prefix. Rate-limiting plays an important role in determining the number of routing updates that are generated after a routing event, and the time it takes before the network converges to a new stable state. Still, there are few guidelines for how rate-limiting timers should be configured in order to achieve the desired convergence properties. This work takes a first step in this direction, by exploring how different rate-limiting implementations and configurations affect the resulting churn level in a live BGP session. Measurements are performed on multiple parallel BGP sessions to a stub AS, configured with and without rate-limiting timers. We find that the daily rate of updates is reduced by two thirds when configuring the timer to the default value recommended by BGP standards. We further investigate different rate-limiting implementations and configurations using the measured BGP update patterns on emulated BGP sessions, and find that increasing the rate-limiting timer gives a logarithmic decrease in churn. Finally, using BGP update traces from RouteViews, we present the first empirical model that quantifies the impact of rate-limiting in terms of churn reduction given the observed arrival pattern of BGP updates. Ahmed Elmokashfi, Amund Kvalbein, Tarik Cicic |
ICC | 1 |
| 2010 | BGP Churn Evolution: a Perspective from the CoreabstractThe scalability limitations of BGP have been a major concern in the networking community lately. An important issue in this respect is the rate of routing updates (churn) that BGP routers must process. This paper presents an analysis of the evolution of churn in four networks in the backbone of the Internet over the last six years, using update traces from the Routeviews project. The churn rate varies widely over time and between networks, and cannot be understood through "black-box'' statistical analysis. Instead we take a different approach with a focus on investigating the underlying reasons for BGP churn evolution. Through our analysis we are able to identify and isolate the main reasons behind many of the anomalies in the churn time series. We find that duplicate announcements is a major churn contributor, and responsible for most large spikes in the churn time series. Other intense periods of churn are caused by misconfigurations or other special events in or close to the monitored AS, and hence limiting these is an important mean to limit churn. We then analyze the remaining "baseline'' churn, and find that it is increasing with a rate much slower than the increase in the routing table size. Ahmed Elmokashfi, Amund Kvalbein, Constantinos Dovrolis |
INFOCOM | 1 |
| 2010 | On the Scalability of BGP: The Role of Topology GrowthabstractThe scalability of BGP routing is a major concern for the Internet community. Scalability is an issue in two different aspects: increasing routing table size, and increasing rate of BGP updates. In this paper, we focus on the latter. Our objective is to characterize the churn increase experienced by ASes in different levels of the Internet hierarchy as the network grows. We look at several "what-if" growth scenarios that are either plausible directions in the evolution of the Internet or educational corner cases, and investigate their scalability implications and interaction with different failure types. Our findings explain the dramatically different impact of multihoming and peering on BGP scalability, highlight negative and positive effects of multihoming on churn and reachability, and identify which topological growth scenarios will lead to faster churn increase for different failure types. Ahmed Elmokashfi, Amund Kvalbein, Constantinos Dovrolis |
IEEE J. Sel. Areas Commun. | 1 |
| 2008 | On the scalability of BGP: the roles of topology growth and update rate-limitingabstractThe scalability of BGP routing is a major concern for the Internet community. Scalability is an issue in two different aspects: increasing routing table size, and increasing rate of BGP updates. In this paper, we focus on the latter. Our objective is to characterize the churn increase experienced by ASes in different levels of the Internet hierarchy as the network grows. We look at several what-if growth scenarios that are either plausible directions in the evolution of the Internet or educational corner cases, and investigate their scalability implications. In addition, we examine the effect of the BGP update rate-limiting timer (MRAI), considering both major variations with which it has been deployed. Our findings explain the dramatically different impact of multi-homing and peering on BGP scalability, identify which topological growth scenarios will lead to faster churn increase, and emphasize the importance of not rate-limiting explicit withdrawals (despite what RFC-4271 recently required). Ahmed Elmokashfi, Amund Kvalbein, Constantinos Dovrolis |
CoNEXT | 1 |
| 2007 | NetForecast: A Delay Prediction Scheme for Provider Controlled NetworksabstractOver the last years, the Internet has evolved towards becoming the dominant platform for deploying real time and multimedia services. This evolution has had as a consequence that the selection of an appropriate server, proxy or super node with reference to some specific QoS parameter becomes of paramount importance. We consider in our paper the specific case of delay estimation. An investigation of existing approaches for the estimation and prediction of network delay is provided. Based on that, we further suggest NetForecast as a way to overcome limitations of existing prediction methods. NetForecast is an algorithm for delay prediction in provider controlled networks. The algorithm is based on a combination of landmark-based distance estimation, clustering and a triangulation principle. The paper reports on preliminary performance of NetForecast, as provided by a simulation study. Our results show the feasibility of the suggested method. Ahmed Elmokashfi, Michael Kleis, Adrian Popescu 0002 |
GLOBECOM | 1 |