Engin Arslan

dblp:85/9469 · DBLP profile ↗
← Back
33ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-2355-057XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 4 first-author · 7 since 2021Computer networks · 5 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
YearPublicationVenuePosition
2026 SmartFLow: A Communication-Efficient SDN Framework for Cross-Silo Federated Learning
abstract
Cross-silo Federated Learning (FL) enables multiple institutions to collaboratively train machine learning models while preserving data privacy. In such settings, clients repeatedly exchange model weights with a central server, making the overall training time highly sensitive to network performance. However, conventional routing methods often fail to prevent congestion, leading to increased communication latency and prolonged training. Software-Defined Networking (SDN), which provides centralized and programmable control over network resources, offers a promising way to address this limitation. To this end, we propose SmartFLow, an SDN-based framework designed to enhance communication efficiency in cross-silo FL. SmartFLow dynamically adjusts routing paths in response to changing network conditions, thereby reducing congestion and improving synchronization efficiency. Experimental results show that SmartFLow decreases parameter synchronization time by up to 47% compared to shortest-path routing and 41% compared to capacity-aware routing. Furthermore, it achieves these gains with minimal computational overhead and scales effectively to networks of up to 50 clients, demonstrating its practicality for real-world FL deployments.
Osama Abu Hamdan, Hao Che, Engin Arslan, Md. Arifuzzaman
CCNC3
2026 FLEET: A Federated Learning Emulation and Evaluation Testbed for Holistic Research
abstract
Federated Learning (FL) presents a robust paradigm for privacy-preserving, decentralized machine learning. However, a significant gap persists between the theoretical design of FL algorithms and their practical performance, largely because existing evaluation tools often fail to model realistic operational conditions. Many testbeds oversimplify the critical dynamics among algorithmic efficiency, client-level heterogeneity, and continuously evolving network infrastructure. To address this challenge, we introduce the Federated Learning Emulation and Evaluation Testbed (FLEET). This comprehensive platform provides a scalable and configurable environment by integrating a versatile, framework-agnostic learning component with a high-fidelity network emulator. FLEET supports diverse machine learning frameworks, customizable real-world network topologies, and dynamic background traffic generation. The testbed collects holistic metrics that correlate algorithmic outcomes with detailed network statistics. By unifying the entire experiment configuration, FLEET enables researchers to systematically investigate how network constraints, such as limited bandwidth, high latency, and packet loss, affect the convergence and efficiency of FL algorithms. This work provides the research community with a robust tool to bridge the gap between algorithmic theory and real-world network conditions, promoting the holistic and reproducible evaluation of federated learning systems.
Osama Abu Hamdan, Hao Che, Engin Arslan, Md. Arifuzzaman
CCNC3
2026 QuRA: Reinforcement Learning Based Routing for Quantum Networks
abstract
Quantum routing deals with identifying a set of quantum repeaters to use to create entanglement between distant endpoints. Previous approaches proposed shortest-path and linear programming methods to find a solution to this problem. While the shortest path approach results in suboptimal performance, linear programming takes too long to find a solution as the network size and constraints increase. In this paper, we apply Deep Q-Reinforcement Learning (DQRL) to optimize routing in quantum networks both in terms of execution time and performance. The proposed Quantum Routing Algorithm (QuRA) first chooses which request to schedule among all requests. It then determines which route to take for the selected request. Since the number of all possible routes can be very high, we developed a hop-by-hop decision making model to lower the complexity while still attaining high performance. Experiments show that the proposed QuRA outperforms existing solutions by up to 90% in success rate and up to 79% in execution time. These results highlight QuRA as a scalable and effective solution for intelligent routing in quantum networks.
Tasdiqul Islam, Engin Arslan, Md. Arifuzzaman
CCNC2
2025 Byte vSwitch: A High-Performance Virtual Switch for Cloud Networking
abstract
Virtual switch is a fundamental component of cloud computing as it provides core networking functionalities for VMs and containers. Open vSwitch (OVS) is widely adopted in cloud environments due to its open-source nature, programmability, and rich set of features. At ByteDance, we initially adopted OVS in our public cloud, but as our cloud business grew, its generic design along with its complex code-base quickly became obstacles to improvements. Hence, we developed Byte vSwitch (BVS), a high-performance virtual switch that was specifically designed to address the performance, scalability, serviceability, and operational efficiency needs of our cloud services. More specifically, BVS adopts a simple architecture with an optimized hash table to maximize forwarding performance. In addition, we introduced several optimizations to improve BVS scalability, operability, and serviceability in cloud environments. Our evaluations show that BVS achieves up to 3.3× higher PPS and 25% lower latency compared to OVS. BVS has been deployed at scale across all regions of the ByteDance public cloud for over four years, and this paper presents our experience in designing, deploying, and operating BVS in production.
Xin Wang 0261, Deguo Li, Lidong Jiang, Shubo Wen, Daxiang Kang, Engin Arslan, Peng He 0003, Xinyu Qian, Jianwen Pi, Xiaoning Ding, Hao Luo 0013
EuroSys7
2023 UNR-IDD: Intrusion Detection Dataset using Network Port Statistics
abstract
Multiple datasets have been proposed to create Machine Learning (ML)-based Network Intrusion Detection Systems (NIDS). However, many of these datasets suffer from sub-optimal performance and inadequate tail class representation. In this paper, we propose the University of Nevada - Reno Intrusion Detection Dataset (UNR-IDD), which utilizes network port statistics for fine-grained analysis of intrusions. Evaluation results show that UNR-IDD is better than existing NIDS datasets with an Fμ, score of 94% and a minimum F-score of 86%. This is mainly because of sufficient and equal representation of various anomaly types in the UNR-IDD dataset.
Tapadhir Das, Osama Abu Hamdan, Raj Mani Shukla, Shamik Sengupta, Engin Arslan
CCNC5
2023 I/O Burst Prediction for HPC Clusters Using Darshan Logs
abstract
Understanding cluster-wide I/O patterns of large-scale HPC clusters is essential to minimize the occurrence and impact of I/O interference. Yet, most previous work in this area focused on monitoring and predicting task and node-level I/O burst events. This paper analyzes Darshan reports from three supercomputers to extract system-level read and write I/O rates in five minutes intervals. We observe significant (over 100×) fluctuations in read and write I/O rates in all three clusters. We then train machine learning models to estimate the occurrence of system-level I/O bursts 5–120 minutes ahead. Evaluation results show that we can predict I/O bursts with more than 90% accuracy (F-1 score) five minutes ahead and more than 87% accuracy two hours ahead. We also show that the ML models attain more than 70% accuracy when estimating the degree of the I/O burst. We believe that high-accuracy predictions of I/O bursts can be used in multiple ways, such as postponing delay-tolerant I/O operations (e.g., checkpointing), pausing nonessential applications (e.g., file system scrubbers), and devising I/O-aware job scheduling methods. To validate this claim, we simulated a burst-aware job scheduler that can postpone the start time of applications to avoid I/O bursts. We show that the burst-aware job scheduling can lead to an up to 5× decrease in application runtime.
Ehsan Saeedizade, Roya Taheri, Engin Arslan
e-Science3
2023 Demystifying the Performance of Data Transfers in High-Performance Research Networks
abstract
High-speed research networks are built to meet the ever-increasing needs of data-intensive distributed workflows. However, data transfers in these networks often fail to attain the promised transfer rates for several reasons, including I/O and network interference, server misconfigurations, and network anomalies. Although understanding the root causes of performance issues is critical to mitigating them and increasing the utilization of expensive network infrastructures, there is currently no available mechanism to monitor data transfers in these networks. In this paper, we present a scalable, end-to-end monitoring framework to gather and store key performance metrics for file transfers to shed light on the performance of transfers. The evaluation results show that the proposed framework can monitor up to 400 transfers per host and more than 40, 000 transfers in total while collecting performance statistics at one-second precision. We also introduce a heuristic method to automatically process the gathered performance metrics and identify the root causes of performance anomalies with an F-score of 87–98%.
Ehsan Saeedizade, Bing Zhang 0018, Engin Arslan
e-Science3
2023 Use Only What You Need: Judicious Parallelism For File Transfers in High Performance Networks
abstract
Parallelism is key to efficiently utilizing high-speed research networks when transferring large volumes of data. However, the monolithic design of existing transfer applications requires the same level of parallelism to be used for read, write, and network operations for file transfers. This, in turn, overburdens system resources since setting the parallelism level for the slowest component results in unnecessarily high parallelism for other components. Using more than necessary parallelism lead to increased overhead on system resources and unfair resource allocation among competing transfers. In this paper, we introduce modular file transfer architecture, Marlin, to separate I/O and network operations for file transfers so that parallelism can be independently adjusted for each component. Marlin adopts online gradient descent algorithm to swiftly search the solution space and find the optimal level of parallelism for read, transfer, and write operations. Experimental results collected under various network settings show that Marlin can identify and use a minimum parallelism level for each component, improving fairness among competing transfers and CPU utilization. Finally, separating network transfers from write operations allows Marlin to outperform the state-of-the-art solutions by more than 2x when transferring small datasets.
Md. Arifuzzaman, Engin Arslan
ICS2
2023 A Heuristic Approach for Scalable Quantum Repeater Deployment Modeling
abstract
Long-distance quantum communication presents a significant challenge due to the degrading fidelity of transmitted qubits. Quantum repeaters are proposed to overcome this issue through Bell measurements on entangled qubit pairs. However, despite its necessity to enable wide-area quantum internet, the deployment cost of quantum repeaters can be prohibitively expensive; thus, developing a quantum repeater deployment model that can strike a balance between cost and effectiveness is essential. In this work, we present novel heuristic models to quickly determine a minimum number of quantum repeaters to deploy in large-scale networks to provide end-to-end connectivity between all end hosts. The results show that, compared to the integer linear programming approach, the heuristic methods can find near-optimal solutions while reducing the execution time from days to seconds when evaluated against several synthetic and real-world networks such as SURFnet and ESnet. As reliability is critical for any network, we demonstrate that the heuristic method can estimate deployment models to endure up to two link/node failures.
Tasdiqul Islam, Engin Arslan
LCN2
2023 Falcon: Fair and Efficient Online File Transfer Optimization
abstract
Research networks provide high-speed wide-area network connectivity between research and education institutions to facilitate large-scale data transfers. However, scalability issues of legacy transfer applications such as scp and FTP hinder the effective utilization of these networks. Although researchers extended the legacy transfer applications to increase their performance by exploiting I/O and network parallelism, these solutions necessitate users to fine-tune parallelism level, a task that is challenging even for experienced users due to the dynamic nature of networks. In this article, we propose an online optimization algorithm,Falcon, to tune the degree of parallelism for file transfers to maximize transfer throughput while keeping system overhead at a minimum. As research networks are shared infrastructures, we introduce a game theory-inspired novel utility function to evaluate the performance of various parallelism levels such that competing transfers are guaranteed to converge to a fair and stable solution. We assessed the performance ofFalconin isolated and production high-speed networks and found that it can discover optimal transfer parallelism in as little as 20 seconds and outperform the state-of-the-art solutions by more than$2\times$. Moreover,Falconis guaranteed to converge to Nash Equilibrium when multiple transfers compete for the same resources with the help of its game theory-inspired utility function. Finally, we demonstrate thatFalconcan also be used as a central transfer scheduler to speed up convergence time, increase stability, and enforce system/user-level resource limitations in shared networks.
Md. Arifuzzaman, Brian Bockelman, Jim Basney, Engin Arslan
IEEE Trans. Parallel Distributed Syst.4
2022 Be SMART, Save I/O: A Probabilistic Approach to Avoid Uncorrectable Errors in Storage Systems
abstract
Silent data corruption poses a significant risk to the integrity of data in storage systems. Although error correction codes (ECC) can recover the majority of such errors, a non-negligible portion of them escape ECC, referred as uncorrectable errors (UEs). Despite being rare in nature, increasing scale of storage systems and fast-growing I/O rates decreased the mean time between UEs from months to hours. Yet, unlike disk failures, UEs are hard to predict with high precision, making it difficult to adopt proactive measures. In this paper, we introduce a probabilistic approach to deploy UE mitigation strategies that can capture significant portion of UE while keeping the system overhead at a tolerable range. To achieve this, we first estimate the probability of I/O operations to be exposed to UEs and find a minimum subset of disks for which employing UE avoidance strategies can lead to significant decrease in UE exposure. We demonstrate through extensive simulations that when the proposed probabilistic model is used to implement write verification strategy to detect and recover from UEs, more than 50% of all write-triggered UEs can be avoided with 1% read overhead, and more than 70% of UEs can be mitigated with less than 3.5% read overhead. We further measure the impact of incurred read overhead on write performance in production Lustre and GPFS file systems and validate our findings that more than half of UEs can be avoided while degrading write I/O throughout by less than 0.9%.
Md. Arifuzzaman, Masudul Hasan Masud Bhuiyan, Mehmet Gümüs, Engin Arslan
CLUSTER4
2022 Bandwidth and Congestion Aware Routing for Wide-Area Hybrid Networks
abstract
An increasing number of science projects rely on sensors to gather data from remote areas. For example, a wildfire detection project deploys video cameras to various high-risk areas to identify wildfires as quickly as possible. These projects typically depend on hybrid networks that are composed of a combination of wireless and wired links to stream data from sensors to datacenters. The management of such hybrid networks is a burdensome task as both link capacities and traffic rates of flows are dynamic and unpredictable.In this paper, we introduce a BAndwidth- and Congestion Aware Routing (BACAR) for hybrid networks using Software Defined Networks. We use active delay measurements to identify congested links and update their capacity to measure traffic rate to detect any link capacity degradation events. We test BACAR in Mininet and show that it offers a robust solution to detect band-width fluctuations that can happen in long-distance wireless links and find alternative routes for rate-sensitive flows to improve overall network utilization and enhance flow performance.
Osama Abu Hamdan, Scotty Strachan, Engin Arslan
LANMAN3
2022 Reliable Wide-Area Data Transfers for Streaming Workflows
abstract
Many large science projects rely on remote clusters for (near) real-time data processing, thus they demand reliable wide-area data transfer performance for smooth end-to-end workflow executions. However, data transfers are often exposed to performance variations due to the changing network (e.g., background traffic) and dataset (e.g., average file size) conditions, necessitating adaptive solutions to meet stringent performance requirements of delay-sensitive streaming workflows. In this article, we proposeFStream++to provide reliable transfer performance for large streaming science applications by dynamically adjusting transfer settings to adapt to changing transfer conditions.FStream++combines three optimization methods asdynamic tuning,online profiling, andhistorical analysisto swiftly and accurately discover optimal transfer settings that can meet workflow requirements. Dynamic tuning uses a heuristic model to predict the values of transfer parameters based on dataset characteristics and network settings. Since heuristic models fall short to incorporate many important factors such as I/O throughput and resource interference, we complement it with online profiling to execute a real-time search for a subset of transfer settings. Finally, historical analysis takes advantage of the long-running nature of streaming workflows by storing and analyzing previous performance observations to shorten the execution time of online profiling. We evaluate the performance ofFStream++by transferring several synthetic and real-world workloads in high-performance production networks and show that it offers up to$3.6x$performance improvement over legacy transfer applications and up to 24% over our previous workFStream.
Hemanta Sapkota, Engin Arslan
IEEE Trans. Parallel Distributed Syst.2
2021 Towards Generalizable Network Anomaly Detection Models
abstract
Finding the root causes of network performance anomalies is critical to satisfy the quality of service requirements. In this paper, we introduce machine learning (ML) models to process TCP socket statistics to pinpoint underlying reasons of performance issues such as packet loss and jitter. More importantly, we introduce a novel feature engineering method to transform network-dependent metrics (e.g., total packet count and round trip time) in training datasets into network-independent forms to be able to transfer the models to new network settings without requiring to retrain them. Experimental results in various network settings show that the proposed feature engineering approach improves the performance of the models in previously unseen network settings from around 60% to nearly 90%. We believe ability to transfer ML models across networks will pave the way for wide adoption of ML solutions in production networks where collecting labeled data is not possible.
Md. Arifuzzaman, Shafkat Islam, Engin Arslan
LCN3
2021 Online optimization of file transfers in high-speed networks
abstract
File transfers in high-speed networks require network and I/O parallelism to reach high speeds, however, creating arbitrarily large numbers of I/O and network threads overwhelms system resources and causes fairness issues. In this paper, we introduce Falcon that combines a novel utility function with state-of-the-art online optimization algorithms to discover the degree of I/O and network parallelism for file transfer that can maximize the throughput while keeping system overhead low and ensuring fairness among competing transfers. Our extensive evaluations in several dedicated and production high-speed networks show that Falcon can find near-optimal solution in as little as 20 seconds and outperforms existing transfer application by 2--6×. Moreover, unlike other file transfer optimization solutions that fail to ensure fair resource allocation between competing transfers, Falcon is guaranteed to converge to Nash Equilibrium when multiple Falcon agents compete for the network resources with the help of its game theory-inspired utility function.
Md. Arifuzzaman, Engin Arslan
SC2
2021 Avoiding data loss and corruption for file transfers with Fast Integrity Verification
Ahmed Alhussen, Engin Arslan
J. Parallel Distributed Comput.2
2020 RIVAChain: Blockchain-based Integrity Verification for File Transfers
abstract
File transfer integrity verification is used to detect silent data corruption by calculating and comparing the checksum of files using secure hash functions, such as SHA-256. However, it incurs significant performance overhead due to I/O and compute-intensive checksum calculation process. In this paper, we present blockchain-based ledger architecture to store the checksum of frequently accessed scientific datasets to minimize the overhead of integrity verification. In the proposed architecture, the checksum of files is calculated and pushed to a private blockchain when they are first created such that future transfers will not require data source to recalculate checksum. As scientific datasets are typically generated in one location (e.g., observatory) and streamed to many geographically distributed locations to enable collaboration, eliminating checksum calculation for data sources will save a significant amount of resource consumption. Moreover, we find that blockchain-based integrity verification reduces transfer time by up to 50% when data source is the bottleneck in the integrity verification process. Finally, we show that private blockchains can scale to thousands of transactions per second thus they are better fit for scientific applications in which data generation rate can easily outpace the transaction confirmation rates of public blockchains.
Ahmed Alhussen, Engin Arslan
IEEE BigData2
2020 Streaming File Transfer Optimization for Distributed Science Workflows
abstract
Driven by the advancements in computing and sensing technology, scientific applications started to generate a huge volume of data which needs to be streamed to highperformance computing clusters timely for real-time (or near-real time) processing, necessitating reliable network performance to operate seamlessly. However, existing data transfer applications are predominantly designed for batch workloads in a way that transfer configurations cannot be altered once they are set. This, in turn, severely limits streaming applications from adapting to changing dataset and network conditions therefore meeting stringent performance requirements. In this paper, we propose FStream to offer performance guarantees to time-sensitive streaming applications by dynamically adjusting transfer settings when system conditions deviate from initial assumptions to sustain high network performance throughout the runtime. We evaluate the performance of FStream by transferring several synthetic and real-world workloads in high-performance production networks and show that it offers up to 9x performance improvement over state-of-the-art data transfer solutions.
Davut Ucar, Engin Arslan
CLUSTER2
2020 Latency Comparison of Cloud Datacenters and Edge Servers
abstract
Edge computing has become a recent approach to bring computing resources closer to the end-user. While offline processing and aggregate data reside in the cloud, edge computing is promoted for latency-critical and bandwidth-hungry tasks. In this direction, it is crucial to quantify the expected latency reduction when edge servers are preferred over cloud locations. In this paper, we performed an extensive measurement to assess the latency characteristics of end-users with respect to the edge servers and cloud data centers. We also evaluated the impact of capacity limitations of edge servers on the latency under various user workloads. We measured latency from 8,456 end-users to 6,341 Akamai edge servers and 69 cloud locations. Measurements of latencies show that while 58% of end-users can reach a nearby edge server in less than 10 ms, only 29% of end-users obtain a similar latency from a nearby cloud location. Additionally, we observe that the latency distribution of end-users to edge servers follows a power-law distribution, which emphasizes the need for non-uniform server deployment and load balancing by an edge provider.
Batyr Charyyev, Engin Arslan, Mehmet Hadi Gunes
GLOBECOM2
2020 RIVA: Robust Integrity Verification Algorithm for High-Speed File Transfers
abstract
End-to-end integrity verification is designed to protect file transfers against silent data corruption by comparing checksum of files at source and destination end points using cryptographic hash functions such as MD5 and SHA1. However, existing implementations of end-to-end integrity verification for file transfers fall short to detect undetected disk errors that causes inconsistency between disk and cache memory. In this article, we propose Robust Integrity Verification Algorithm (RIVA) to strengthen the integrity of file transfers by forcing checksum computation tasks to read files directly from disk. RIVA achieves this by invalidating memory mappings of file pages after their transfer such that when the file is read again for checksum calculation, it will be fetched from disk and silent disk errors will be captured. We design and conduct extensive fault resilience experiments to evaluate the robustness of integrity verification algorithms against undetected disk write errors. The results indicate that while the state-of-the-art integrity verification algorithms fail to detect the injected errors for almost all file sizes, RIVA captures all of them with the help of cache invalidation. We further run statistical analysis to assess the probability of missing silent disk errors and find that RIVA reduces the likelihood by 10 to 15 orders of magnitude compared to the existing approaches. Finally, enforcing disk read in integrity verification introduces an inevitable overhead in exchange of increased robustness against silent disk errors, but RIVA keeps its overhead below 15 percent in most cases by running transfer, cache invalidation, and checksum computation processes concurrently for different portions of the same file.
Batyr Charyyev, Engin Arslan
IEEE Trans. Parallel Distributed Syst.2
2019 Towards Securing Data Transfers Against Silent Data Corruption
abstract
Scientific applications generate large volumes of data that often needs to be moved between geographically distributed sites for collaboration or backup which has led to a significant increase in data transfer rates. As an increasing number of scientific applications are becoming sensitive to silent data corruption, end-to-end integrity verification has been proposed. It minimizes the likelihood of silent data corruption by comparing checksum of files at the source and the destination using secure hash algorithms such as MD5 and SHA1. In this paper, we investigate the robustness of existing end-to-end integrity verification approaches against silent data corruption and propose a Robust Integrity Verification Algorithm (RIVA) to enhance data integrity. Extensive experiments show that unlike existing solutions, RIVA is able to detect silent disk corruptions by invalidating file contents in page cache and reading them directly from disk. Since RIVA clears page cache and reads file contents directly from the disk, it incurs delay to execution time. However, by running transfer, cache invalidation, and checksum operations concurrently, RIVA is able to keep its overhead below 15% in most cases compared to the state-of-the-art solutions in exchange of increasing the robustness to silent data corruption. We also implemented dynamic transfer and checksum parallelism to overcome performance bottlenecks and observed more than 5x increase in RIVA's speed.
Batyr Charyyev, Ahmed Alhussen, Hemanta Sapkota, Eric Pouyoul, Mehmet Hadi Gunes, Engin Arslan
CCGRID6
2019 Pooling Approach for Task Allocation in the Blockchain Based Decentralized Storage Network
abstract
Blockchain technology has provided a solid system to develop incentivization algorithms using the smart contract. Blockchain applies the distributed ledger to store transaction histories, and the information is stored across a network of computers instead of on a single server. This facilitates the development of a new set of applications such as distributed file storage systems where users can rent out their storage in return for a premium. The distributed file storage systems provide more privacy and security compared to the centralized storage models as there is no need to have a trusted party. New schemes have been developed for distributed file storage systems on top of the blockchain platform, however, the problem of task/service allocation in these models have not been studied before. In this paper, we study the task/service allocation in the distributed file storage systems considering the challenge of computation cost. First, we formalize the problem of task/service allocation in a decentralized storage network, and then we discuss different approaches to allocate storage tasks to storage servers in an efficient manner. Moreover, we study the benefits of the cooperation (a.k.a pooling) in the storage and retrieval markets of distributed storage networks. The evaluation results show the benefit of our proposed pooling based approach in storage and retrieval markets.
Iman Vakilinia, Shahin Vakilinia, Shahriar Badsha, Engin Arslan, Shamik Sengupta
CNSM4
2018 A Low-Overhead Integrity Verification for Big Data Transfers
abstract
The amount of data generated by scientific and commercial applications is growing at an ever-increasing pace. This data is often moved between geographically distributed sites for various purposes such as collaboration and backup which has led to significant increase in data transfer rates. Surge in data transfer rates when combined with proliferation of scientific applications that cannot tolerate data corruption triggered enhanced integrity verification techniques to be developed. End- to-end integrity verification minimizes the likelihood of silent data corruption by comparing checksum of files at source and destination servers using secure hash algorithms such as MD5 and SHA1. However, it imposes significant performance penalty due to overhead of checksum computation. In this paper, we propose Fast Integrity VERification (FIVER) algorithm which overlaps checksum computation and data transfer operations of files to minimize the cost of integrity verification. Extensive experiments show that FIVER is able to bring down the cost from 60% by the state-of-the-art solutions to below 10% by concurrently executing transfer and checksum operations and enabling file I/O share between them. We also implemented FIVER-Hybrid to mimic disk access patterns of sequential integrity verification approach to capture possible data corruption that may occur during file write operations which FIVER may miss. Results show that FIVER-Hybrid is able to reduce execution time by 20% compared to sequential approach without compromising the reliability of integrity verification.
Engin Arslan, Ahmed Alhussen
IEEE BigData1
2018 PCC Vivace: Online-Learning Congestion Control
Mo Dong, Tong Meng, Doron Zarchy, Engin Arslan, Yossi Gilad, Brighten Godfrey, Michael Schapira
NSDI4
2018 Big data transfer optimization through adaptive parameter tuning
Engin Arslan, Bahadir A. Pehlivan, Tevfik Kosar
J. Parallel Distributed Comput.1
2018 High-Speed Transfer Optimization Based on Historical Analysis and Real-Time Tuning
abstract
Data-intensive scientific and commercial applications increasingly require frequent movement of large datasets from one site to the other(s). Despite growing network capacities, these data movements rarely achieve the promised data transfer rates of the underlying physical network due to poorly tuned data transfer protocols. Accurately and efficiently tuning the data transfer protocol parameters in a dynamically changing network environment is a major challenge and remains as an open research problem. In this paper, we present a novel dynamic parameter tuning algorithm based on historical data analysis and real-time background traffic probing, dubbed HARP. Most of the previous work in this area are solely based on real-time network probing or static parameter tuning, which either result in an excessive sampling overhead or fail to accurately predict the optimal transfer parameters. Combining historical data analysis with real-time sampling lets HARP tune the application-layer data transfer parameters accurately and efficiently to achieve close-to-optimal end-to-end data transfer throughput with very low overhead. Instead of one-time parameter estimation, HARP uses a feedback loop to adjust the parameter values to changing network conditions in real-time. Our experimental analyses over a variety of network settings show that HARP outperforms existing solutions by up to 50 percent in terms of the achieved data transfer throughput.
Engin Arslan, Kemal Guner, Tevfik Kosar
IEEE Trans. Parallel Distributed Syst.1
2016 HARP: predictive transfer optimization based on historical analysis and real-time probing
abstract
Increasingly data-intensive scientific and commercial applications require frequent movement of large datasets from one site to the other. Despite the growing capacity of the networking capacity, these data movements rarely achieve the promised data transfer rates of the underlying physical network due to the poorly tuned data transfer protocols. Accurately and efficiently tuning the data transfer protocol parameters in a dynamically changing network environment is a big challenge and still an open research problem. In this paper, we present predictive end-to-end data transfer optimization algorithms based on historical data analysis and real-time background traffic probing, dubbed HARP. Most of the existing work in this area is solely based on real time network probing, which either cause too much sampling overhead or fail to accurately predict the correct transfer parameters. Combining historical data analysis with real time sampling enables our algorithms to tune the application level data transfer parameters accurately and efficiently to achieve close-to-optimal end-to-end data transfer throughput with very low overhead. Our experimental analysis over a variety of network settings shows that HARP outperforms existing solutions by up to 50% in terms of the achieved throughput.
Engin Arslan, Kemal Guner, Tevfik Kosar
SC1
2016 Application-Level Optimization of Big Data Transfers through Pipelining, Parallelism and Concurrency
abstract
In end-to-end data transfers, there are several factors affecting the data transfer throughput, such as the network characteristics (e.g., network bandwidth, round-trip-time, background traffic); end-system characteristics (e.g., NIC capacity, number of CPU cores and their clock rate, number of disk drives and their I/O rate); and the dataset characteristics (e.g., average file size, dataset size, file size distribution). Optimization of big data transfers over inter-cloud and intra-cloud networks is a challenging task that requires joint-consideration of all of these parameters. This optimization task becomes even more challenging when transferring datasets comprised of heterogeneous file sizes (i.e., large files and small files mixed). Previous work in this area only focuses on the end-system and network characteristics however does not provide models regarding the dataset characteristics. In this study, we analyze the effects of the three most important transfer parameters that are used to enhance data transfer throughput: pipelining,parallelism and concurrency. We provide models and guidelines to set the best values for these parameters and present two different transfer optimization algorithms that use the models developed. The tests conducted over high-speed networking and cloud testbeds show that our algorithms outperform the most popular data transfer tools like Globus Online and UDT in majority of the cases.
Esma Yildirim, Engin Arslan, JangYoung Kim, Tevfik Kosar
IEEE Trans. Cloud Comput.2
2015 Energy-aware data transfer algorithms
abstract
The amount of data moved over the Internet per year has already exceeded the Exabyte scale and soon will hit the Zettabyte range. To support this massive amount of data movement across the globe, the networking infrastructure as well as the source and destination nodes consume immense amount of electric power, with an estimated cost measured in billions of dollars. Although considerable amount of research has been done on power management techniques for the networking infrastructure, there has not been much prior work focusing on energy-aware data transfer algorithms for minimizing the power consumed at the end-systems. We introduce novel data transfer algorithms which aim to achieve high data transfer throughput while keeping the energy consumption during the transfers at the minimal levels. Our experimental results show that our energy-aware data transfer algorithms can achieve up to 50% energy savings with the same or higher level of data transfer throughput.
Ismail Alan, Engin Arslan, Tevfik Kosar
SC2
2015 Training network administrators in a game-like environment
Engin Arslan, Murat Yuksel, Mehmet Hadi Gunes
J. Netw. Comput. Appl.1
2014 Energy-Aware Data Transfer Tuning
abstract
The annual electricity consumed by data transfers in the U.S. is estimated to be 20 Terawatt hours, which translates to around 4 billion U.S. Dollars per year. There has been considerable amount of prior work looking at power management and energy efficiency in hardware and software systems, and more recently in power-aware networking. Despite the growing body of research in power management techniques for the networking infrastructure, there has been no prior work (to the best of our knowledge), focusing on saving energy at the end systems(sender and receiver nodes) during the data transfer. We argue that although network-only approaches are part of the solution, the end-system power management is a key in optimizing energy efficiency of the data transfers, which has been long ignored. In this paper, we analyze various factors that will affect the power consumption in end-to-end data transfers, such as the level of parallelism, concurrency and pipelining. Our results show that significant amount of energy savings can be achieved at the end-systems during data transfer with no or minimal performance penalty.
Ismail Alan, Engin Arslan, Tevfik Kosar
CCGRID2
2013 Dynamic Protocol Tuning Algorithms for High Performance Data Transfers
Engin Arslan, Brandon Ross, Tevfik Kosar
Euro-Par1
2011 Network management game
abstract
Network management and automated configuration of large-scale networks is one of the crucial issues for Internet Service Providers (ISPs). Since wrong configurations might lead to an enormous amount of customer traffic to be lost, highly experienced network administrators are typically the ones who are trusted for the management and configuration of a running ISP network. We frame the management and experimentation of a network as a “game” for training network administrators without having to risk the network operation. The interactive environment treats the trainee network administrators as players of a game and tests them with various network failures or dynamics. To prototype the concept of “network management as a game”, we modified NS-2 to establish an interactive simulation engine and connected the modified engine to a graphical user interface for traffic animation and interactivity with the player. We present initial results from our game applied to a small set of players.
Engin Arslan, Murat Yuksel, Mehmet Hadi Gunes
LANMAN1