VLDB 2026 Research / reviewers in the wild / expert
Tevfik Kosar
dblp:29/3036
· DBLP profile ↗
59ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0002-5600-6706ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 5 first-author · 5 since 2021Computer networks · 8 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorArtificial intelligence and machine learning · 3Software engineering, systems software and programming languages · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RapidGNN: Communication-Efficient Distributed Training on Large-Scale Graph Neural Networks
Arefin Niam, Tevfik Kosar, Md. S. Q. Zulkar Nine |
CCGrid | 2 |
| 2026 | How to Evaluate Distributed Coordination Systems?-A Survey and AnalysisabstractCoordination services and protocols are critical components of distributed systems and are essential for providing consistency, fault tolerance, and scalability. However, due to the lack of standard benchmarking and evaluation tools for distributed coordination services, coordination service developers/researchers either use a NoSQL standard benchmark and omit evaluating consistency, distribution, and fault tolerance; or create their own ad-hoc microbenchmarks and skip comparability with other services. In this study, we analyze and compare the evaluation mechanisms for known and widely used consensus algorithms, distributed coordination services, and distributed applications built on top of these services. We identify the most important requirements of distributed coordination service benchmarking, such as the metrics and parameters for the evaluation of the performance, scalability, availability, and consistency of these systems. Finally, we discuss why the existing benchmarks fail to address the complex requirements of distributed coordination system evaluation. Bekir O. Turkkan, Elvis Rodrigues, Tevfik Kosar, Aleksey Charapko, Ailidani Ailijiang, Murat Demirbas |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | Carbon-Aware Temporal Data Transfer Scheduling Across Cloud DatacentersabstractInter-datacenter communication is a significant part of cloud operations and produces a substantial amount of carbon emissions for cloud data centers, where the environmental impact has already been a pressing issue. In this paper, we present a novel carbon-aware temporal data transfer scheduling framework, called LinTS, which promises to significantly reduce the carbon emission of data transfers between cloud data centers. LinTS produces a competitive transfer schedule and makes scaling decisions, outperforming common heuristic algorithms. LinTS can lower carbon emissions during inter-datacenter transfers by up to 66 % compared to the worst case and up to 15 % compared to other solutions while preserving all deadline constraints. Elvis Rodrigues, Jacob Goldverg, Tevfik Kosar |
CLOUD | 3 |
| 2025 | Message from the Congress Program ChairsabstractWe are delighted to welcome all participants to the 2025 IEEE World Congress on Services (IEEE SERVICES 2025), which is taking place in the beautiful city of Helsinki, Finland. To support the services community of researchers and practitioners around the world, IEEE SERVICES 2025 is held in hybrid mode. Claudio A. Ardagna, Qiang He 0001, Tevfik Kosar |
SSE | 3 |
| 2025 | FlowTracer: A Tool for Uncovering Network Path Usage Imbalance in AI Training Clusters
Hasibul Jamil, Abdul Alim, Laurent Schares, Pavlos Maniotis, Liran Schour, Ali Sydney, Abdullah Kayi, Tevfik Kosar, Bengi Karaçali |
ICC | 8 |
| 2025 | Message from the Congress Program ChairsabstractWe are delighted to welcome all participants to the 2025 IEEE World Congress on Services (IEEE SERVICES 2025), which is taking place in the beautiful city of Helsinki, Finland. To support the services community of researchers and practitioners around the world, IEEE SERVICES 2025 is held in hybrid mode. Claudio A. Ardagna, Qiang He 0001, Tevfik Kosar |
ICWS | 3 |
| 2024 | Towards Sustainable Cloud Software Systems through Energy-Aware Code Smell RefactoringabstractSoftware applications and workloads, especially within the domains of Cloud computing and large-scale AI model training, exert considerable demand on computing resources, thus contributing significantly to the overall energy footprint of the IT industry. In this paper, we present an in-depth analysis of certain software coding practices that can play a substantial role in increasing the application's overall energy consumption, primarily stemming from the suboptimal utilization of computing resources. Our study encompasses a thorough investigation of 16 distinct code smells and other coding malpractices across 31 real-world open-source applications written in Java and Python. Through our research, we provide compelling evidence that vari-ous common refactoring techniques, typically employed to rectify specific code smells, can unintentionally escalate the application's energy consumption. We illustrate that a discerning and strategic approach to code smell refactoring can yield substantial energy savings. For selective refactorings, this yields a reduction of up to 13.1 % of energy consumption and 5.1 % of carbon emissions per workload on average. These findings underscore the potential of selective and intelligent refactoring to substantially increase energy efficiency of Cloud software systems. Asif Imran, Tevfik Kosar, Jaroslaw Zola, Muhammed Fatih Bulut |
CLOUD | 2 |
| 2024 | 2024 IEEE International Conference on Cloud Computing, Message from the ChairsabstractWe are delighted to welcome all participants to the 2024 IEEE International Conference on Cloud Computing (CLOUD 2024), which is taking place in the beautiful city of Shenzhen, China, from July 7th to 13th. Tevfik Kosar, Krishnan Venkateswaran, Shangguang Wang, Seetharami Seelam, Santonu Sarkar, Xuanzhe Liu |
CLOUD | 1 |
| 2024 | GreenABR+: Generalized Energy-Aware Adaptive Bitrate StreamingabstractAdaptive bitrate (ABR) algorithms play a critical role in video streaming by making optimal bitrate decisions in dynamically changing network conditions to provide a high quality of experience (QoE) for users. However, most existing ABRs suffer from limitations such as predefined rules and incorrect assumptions about streaming parameters. They often prioritize higher bitrates and ignore the corresponding energy footprint, resulting in increased energy consumption, especially for mobile device users. Additionally, most ABR algorithms do not consider perceived quality, leading to suboptimal user experience. This article proposes a novel ABR scheme called GreenABR+, which utilizes deep reinforcement learning to optimize energy consumption during video streaming while maintaining high user QoE. Unlike existing rule-based ABR algorithms, GreenABR+ makes no assumptions about video settings or the streaming environment. GreenABR+ model works on different video representation sets and can adapt to dynamically changing conditions in a wide range of network scenarios. Our experiments demonstrate that GreenABR+ outperforms state-of-the-art ABR algorithms by saving up to 57% in streaming energy consumption and 57% in data consumption while providing up to 25% more perceptual QoE due to up to 87% less rebuffering time and near-zero capacity violations. The generalization and dynamic adaptability make GreenABR+ a flexible solution for energy-efficient ABR optimization. Bekir O. Turkkan, Adithya Raman, Tevfik Kosar, Changyou Chen, Muhammed Fatih Bulut, Jaroslaw Zola, Daby M. Sow |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Message from the ChairsabstractWe are delighted to extend a warm welcome to all participants of the 2023 IEEE International Conference on Cloud Computing (CLOUD 2023) sponsored by the IEEE Technical Committee on Services Computing! Tevfik Kosar, Manish Parashar, Judy Fox, Christoph Hagleithner |
CLOUD | 1 |
| 2023 | CloudScent: A Model for Code Smell Analysis in Open-Source CloudabstractThe low cost and rapid provisioning capabilities have made open-source cloud a desirable platform to launch industrial applications. However, as open-source cloud moves towards maturity, it still suffers from quality issues like code smells. Although, a great emphasis has been provided on the economic benefits of deploying open-source cloud, low importance has been provided to improve the quality of the source code of the cloud itself to ensure its maintainability in the industrial scenario. Code refactoring has been associated with improving the maintenance and understanding of software code by removing code smells. However, analyzing what smells are more prevalent in cloud environment and designing a tool to define and detect those smells require further attention. In this paper, we propose a model called CloudScent which is an open source mechanism to detect smells in open-source cloud. We test our experiments in a real-life cloud environment using OpenStack. Results show that CloudScent is capable of accurately detecting 8 code smells in cloud. This will permit cloud service providers with advanced knowledge about the smells prevalent in open-source cloud platform, thus allowing for timely code refactoring and improving code quality of the cloud platforms. Raj Narendra Shah, Sameer Ahmed Mohamed, Asif Imran, Tevfik Kosar |
CloudCom | 4 |
| 2023 | Learning to Maximize Network Bandwidth Utilization with Deep Reinforcement LearningabstractEfficiently transferring data over long-distance, high-speed networks requires optimal utilization of available network bandwidth. One effective method to achieve this is through the use of parallel TCP streams. This approach allows applications to leverage network parallelism, thereby enhancing transfer throughput. However, determining the ideal number of parallel TCP streams can be challenging due to non-deterministic background traffic sharing the network, as well as non-stationary and partially observable network signals. We present a novel learning-based approach that utilizes deep reinforcement learning (DRL) to determine the optimal number of parallel TCP streams. Our DRL-based algorithm is designed to intelligently utilize available network bandwidth while adapting to different network conditions. Unlike rule-based heuristics, which lack generalization in unknown network scenarios, our DRL-based solution can dynamically adjust the parallel TCP stream numbers to optimize network bandwidth utilization without causing network congestion and ensuring fairness among competing transfers. We conducted extensive experiments to evaluate our DRL-based algorithm's performance and compared it with several state-of-the-art online optimization algorithms. The results demonstrate that our algorithm can identify nearly optimal solutions 40 % faster while achieving up to 15 % higher throughput. Further-more, we show that our solution can prevent network congestion and distribute the available network resources fairly among competing transfers, unlike a discriminatory algorithm. Hasibul Jamil, Elvis Rodrigues, Jacob Goldverg, Tevfik Kosar |
GLOBECOM | 4 |
| 2023 | GreenNFV: Energy-Efficient Network Function Virtualization with Service Level Agreement ConstraintsabstractNetwork Function Virtualization (NFV) platforms consume significant energy, introducing high operational costs in edge and data centers. This paper presents a novel framework called GreenNFV that optimizes resource usage for network function chains using deep reinforcement learning. GreenNFV optimizes resource parameters such as CPU sharing ratio, CPU frequency scaling, last-level cache (LLC) allocation, DMA buffer size, and packet batch size. GreenNFV learns the resource scheduling model from the benchmark experiments and takes Service Level Agreements (SLAs) into account to optimize resource usage models based on the different throughput and energy consumption requirements. Our evaluation shows that GreenNFV models achieve high transfer throughput and low energy consumption while satisfying various SLA constraints. Specifically, GreenNFV with Throughput SLA can achieve 4.4× higher throughput and 1.5× better energy efficiency over the baseline settings, whereas GreenNFV with Energy SLA can achieve 3× higher throughput while reducing energy consumption by 50%. Md. S. Q. Zulkar Nine, Tevfik Kosar, Muhammed Fatih Bulut, Jinho Hwang |
SC | 2 |
| 2022 | Energy-Efficient Data Transfer Optimization via Decision-Tree Based Uncertainty ReductionabstractThe increase and rapid growth of data produced by scientific instruments, the Internet of Things (IoT), and social media is causing data transfer performance and resource consumption to garner much attention in the research community. The network infrastructure and end systems that enable this extensive data movement use a substantial amount of electricity, measured in terawatt-hours per year. Managing energy consumption within the core networking infrastructure is an active research area, but there is a limited amount of work on reducing power consumption at the end systems during active data transfers. This paper presents a novel two-phase dynamic throughput and energy optimization model that utilizes an offline decision-search-tree based clustering technique to encapsulate and categorize historical data transfer log information and an online search optimization algorithm to find the best application and kernel layer parameter combination to maximize the achieved data transfer throughput while minimizing the energy consumption. Our model also incorporates an ensemble method to reduce aleatoric uncertainty in finding optimal application and kernel layer parameters during the offline analysis phase. The experimental evaluation results show that our decision-tree based model outperforms the state-of-the-art solutions in this area by achieving 117% higher throughput on average and also consuming 19% less energy at the end systems during active data transfers. Hasibul Jamil, Lavone Rodolph, Jacob Goldverg, Tevfik Kosar |
ICCCN | 4 |
| 2022 | GreenABR: energy-aware adaptive bitrate streaming with deep reinforcement learningabstractAdaptive bitrate (ABR) algorithms aim to make optimal bitrate decisions in dynamically changing network conditions to ensure a high quality of experience (QoE) for the users during video streaming. However, most of the existing ABRs share the limitations of predefined rules and incorrect assumptions about streaming parameters. They also come short to consider the perceived quality in their QoE model, target higher bitrates regardless, and ignore the corresponding energy consumption. This joint approach results in additional energy consumption and becomes a burden, especially for mobile device users. This paper proposes GreenABR, a new deep reinforcement learning-based ABR scheme that optimizes the energy consumption during video streaming without sacrificing the user QoE. GreenABR employs a standard perceived quality metric, VMAF, and real power measurements collected through a streaming application. GreenABR's deep reinforcement learning model makes no assumptions about the streaming environment and learns how to adapt to the dynamically changing conditions in a wide range of real network scenarios. GreenABR outperforms the existing state-of-the-art ABR algorithms by saving up to 57% in streaming energy consumption and 60% in data consumption while achieving up to 22% more perceptual QoE due to up to 84% less rebuffering time and near-zero capacity violations. Bekir O. Turkkan, Adithya Raman, Tevfik Kosar, Changyou Chen, Muhammed Fatih Bulut, Jaroslaw Zola, Daby M. Sow |
MMSys | 4 |
| 2022 | SMURF: Efficient and Scalable Metadata Access for Distributed ApplicationsabstractIn parallel with big data processing and analysis dominating the usage of distributed and Cloud infrastructures, the demand for distributed metadata access and transfer has increased. The volume of data generated by many application domains exceeds petabytes, while the corresponding metadata amounts to terabytes or even more. This article proposes a novel solution for efficient and scalable metadata access for distributed applications across wide-area networks, dubbed SMURF. Our solution combines novel pipelining and concurrent transfer mechanisms with reliability, provides distributed continuum caching and semantic locality-aware prefetching strategies to sidestep fetching latency, and achieves scalable and high-performance metadata fetch/prefetch services in the Cloud. We incorporate the phenomenon of semantic locality awareness for increased prefetch prediction rate using real-life application I/O traces from Yahoo! Hadoop audit logs and propose a novel prefetch predictor. By effectively caching and prefetching metadata based on the access patterns, our continuum caching and prefetching mechanism significantly improves the local cache hit rate and reduces the average fetching latency. We replay approximately 20 Million metadata access operations from real audit traces, where SMURF achieves 90% accuracy during prefetch prediction and reduced the average fetch latency by 50% compared to the state-of-the-art mechanisms. Bing Zhang 0018, Tevfik Kosar |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | Energy-saving Cross-layer Optimization of Big Data Transfer Based on Historical Log AnalysisabstractWith the proliferation of data movement across the Internet, global data traffic per year has already exceeded the Zettabyte scale. The network infrastructure and end-systems facilitating the vast data movement consume an extensive amount of electricity, measured in terawatt-hours per year. This massive energy footprint costs the world economy billions of dollars partially due to energy consumed at the network end-systems. Although extensive research has been done on managing power consumption within the core networking infrastructure, there is little research on reducing the power consumption at the end-systems during active data transfers. This paper presents a novel cross-layer optimization framework, called Cross-LayerHLA, to minimize energy consumption at the end-systems by applying machine learning techniques to historical transfer logs and extracting the hidden relationships between different parameters affecting both the performance and resource utilization. It utilizes offline analysis to improve online learning and dynamic tuning of application-level and kernel-level parameters with minimal overhead. This approach minimizes end-system energy consumption and maximizes data transfer throughput. Our experimental results show that Cross-LayerHLA outperforms other state-of-the-art solutions in this area. Lavone Rodolph, Md. S. Q. Zulkar Nine, Luigi Di Tacchio, Tevfik Kosar |
ICC | 4 |
| 2021 | A Two-Phase Dynamic Throughput Optimization Model for Big Data TransfersabstractThe amount of data transferred over dedicated and non-dedicated network links has been increasing much faster than the increase in the network capacity. On the other hand, the current data transfer solutions fail to guarantee even the promised achievable transfer throughput. In this article, we propose a novel two-phase dynamic throughput optimization model based on mathematical modeling with offline knowledge discovery/analysis and adaptive online decision making. In the offline analysis, we mine historical transfer logs to perform knowledge discovery about the transfer characteristics. The online phase uses the discovered knowledge from the offline analysis along with the real-time investigation of the network condition to optimize the protocol parameters. As the real-time investigation is expensive and provides partial knowledge about the current network status, our model uses historical knowledge about the network and data characteristics to reduce the real-time investigation overhead while ensuring near-optimal throughput for each transfer. Our novel approach is tested over different networks with different datasets, and it has outperformed its closest competitor by 1.7x and the default case by 5x. It also achieved up to 93 percent accuracy compared to the optimal achievable throughput possible on those networks. Md. S. Q. Zulkar Nine, Tevfik Kosar |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | The Impact of Auto-Refactoring Code Smells on the Resource Utilization of Cloud Software
Asif Imran, Tevfik Kosar |
SEKE | 2 |
| 2020 | WPaxos: Wide Area Network Flexible ConsensusabstractWPaxos is a multileader Paxos protocol that provides low-latency and high-throughput consensus across wide-area network (WAN) deployments. WPaxos uses multileaders, and partitions the object-space among these multileaders. Unlike statically partitioned multiple Paxos deployments, WPaxos is able to adapt to the changing access locality through object stealing. Multiple concurrent leaders coinciding in different zones steal ownership of objects from each other using phase-1 of Paxos, and then use phase-2 to commit update-requests on these objects locally until they are stolen by other leaders. To achieve fast phase-2 commits, WPaxos adopts the flexible quorums idea in a novel manner, and appoints phase-2 acceptors to be close to their respective leaders. We implemented WPaxos and evaluated it over WAN deployments across 5 AWS regions. The dynamic partitioning of the objectspace and emphasis on zone-local commits allow WPaxos to significantly outperform both partitioned Paxos deployments and leaderless Paxos approaches. Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas, Tevfik Kosar |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | Cross-Layer Optimization of Big Data Transfer Throughput and Energy ConsumptionabstractWith the emergence of data deluge, the energy footprint of global data movement has surpassed 100 terawatt hours, costing more than 20 billion US dollars to the world economy. During an active data transfer, depending on the number of hops between the source and destination, the networking infrastructure consumes between 10% - 75% of the total energy, and the rest is consumed by the end systems. Even though there has been extensive research on reducing the power consumption at the networking infrastructure, the work focusing on saving energy at the end systems has been limited to the tuning of a few application-level parameters. In this paper, we introduce a novel cross-layer optimization framework which jointly considers application-level and kernel-level parameters to minimize the energy consumption without sacrificing from the transfer throughput. We present three different algorithms which can dynamically tune the CPU frequency level, number of active CPU cores, number of active transfer threads, number of parallel TCP streams, and the level of transfer command pipelining to achieve different user-set goals. Experimental results show that our proposed algorithms outperform the state-of-the-art solutions, achieving up to 80% higher throughput while consuming 48% less energy. Luigi Di Tacchio, Md. S. Q. Zulkar Nine, Tevfik Kosar, Muhammed Fatih Bulut, Jinho Hwang |
CLOUD | 3 |
| 2019 | Energy-aware data throughput optimization for next generation internet
Tevfik Kosar, Ismail Alan, Muhammed Fatih Bulut |
Inf. Sci. | 1 |
| 2018 | GreenDataFlow: Minimizing the Energy Footprint of Global Data MovementabstractThe global data movement over Internet has an estimated energy footprint of 100 terawatt hours per year, costing the world economy billions of dollars. The networking infrastructure together with source and destination nodes involved in the data transfer contribute to overall energy consumption. Although considerable amount of research has rendered power management techniques for the networking infrastructure, there has not been much prior work focusing on energy-aware data transfer solutions for minimizing the power consumed at the end-systems. In this paper, we introduce a novel application-layer solution based on historical analysis and real-time tuning called GreenDataFlow, which aims to achieve high data transfer throughput while keeping the energy consumption at the minimal levels. GreenDataFlow supports service level agreements (SLAs) which give the service providers and the consumers the ability to fine tune their goals and priorities in this optimization process. Our experimental results show that GreenDataFlow outperforms the closest competing state-of-the art solution in this area 50% for energy saving and 2.5× for the achieved end-to-end performance. Md. S. Q. Zulkar Nine, Luigi Di Tacchio, Asif Imran, Tevfik Kosar, Muhammed Fatih Bulut, Jinho Hwang |
IEEE BigData | 4 |
| 2018 | OneDataShare - A Vision for Cloud-hosted Data Transfer Scheduling and Optimization as a ServiceabstractFast, reliable, and efficient data transmission across wide-area networks is a predominant bottleneck for data-intensive cloud applications. This paper introduces OneDataShare, which is designed to eliminate the issues plaguing effective cloud-based data transfers of varying file sizes and across incompatible transfer end-points. The vision of OneDataShare is to achieve high-speed data communication, interoperability between multiple transfer protocols, and accurate estimation of delivery time for advance planning, thereby maximizing user-profit through improved and faster data analysis for business intelligence. The paper elaborates on the desirable features of OneDataShare as a cloud-hosted data transfer scheduling and optimization service, and how it is aligned with the vision of harnessing the power of the cloud and distributed computing. Experimental evaluation and comparison with existing real-life file transfer services show that the transfer throughout achieved by OneDataShare is 6.5 times greater. Asif Imran, Md. S. Q. Zulkar Nine, Kemal Guner, Tevfik Kosar |
CLOSER | 4 |
| 2018 | Energy-Efficient Mobile Network I/OabstractBy year 2020, the number of smartphone users globally will reach 3 Billion and the mobile data traffic (cellular + WiFi) will exceed PC Internet traffic the first time. As the number of smartphone users and the amount of data transferred per smartphone grow exponentially, limited battery power is becoming an increasingly critical problem for mobile devices which heavily depend on network I/O. Despite the growing body of research in power management techniques for mobile devices at the hardware and networking layers, there has been little work focusing on saving energy at the application layer for mobile network I/O. In this paper, we show that significant energy savings can be achieved with application-layer solutions at the mobile systems during data transfer with no performance penalty. In many cases, performance increase and energy savings can be achieved simultaneously. Kemal Guner, Tevfik Kosar |
GLOBECOM | 2 |
| 2018 | Big data transfer optimization through adaptive parameter tuning
Engin Arslan, Bahadir A. Pehlivan, Tevfik Kosar |
J. Parallel Distributed Comput. | 3 |
| 2018 | High-Speed Transfer Optimization Based on Historical Analysis and Real-Time TuningabstractData-intensive scientific and commercial applications increasingly require frequent movement of large datasets from one site to the other(s). Despite growing network capacities, these data movements rarely achieve the promised data transfer rates of the underlying physical network due to poorly tuned data transfer protocols. Accurately and efficiently tuning the data transfer protocol parameters in a dynamically changing network environment is a major challenge and remains as an open research problem. In this paper, we present a novel dynamic parameter tuning algorithm based on historical data analysis and real-time background traffic probing, dubbed HARP. Most of the previous work in this area are solely based on real-time network probing or static parameter tuning, which either result in an excessive sampling overhead or fail to accurately predict the optimal transfer parameters. Combining historical data analysis with real-time sampling lets HARP tune the application-layer data transfer parameters accurately and efficiently to achieve close-to-optimal end-to-end data transfer throughput with very low overhead. Instead of one-time parameter estimation, HARP uses a feedback loop to adjust the parameter values to changing network conditions in real-time. Our experimental analyses over a variety of network settings show that HARP outperforms existing solutions by up to 50 percent in terms of the achieved data transfer throughput. Engin Arslan, Kemal Guner, Tevfik Kosar |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | Big data transfer optimization based on offline knowledge discovery and adaptive samplingabstractThe amount of data moved over dedicated and non-dedicated network links increases much faster than the increase in the network capacity, but the current solutions fail to guarantee even the promised achievable transfer throughputs. In this paper, we propose a novel dynamic throughput optimization model based on mathematical modeling with offline knowledge discovery/analysis and adaptive online decision making. In offline analysis, we mine historical transfer logs to perform knowledge discovery about the transfer characteristics. Online phase uses the discovered knowledge from the offline analysis along with real-time investigation of the network condition to optimize the protocol parameters. As real-time investigation is expensive and provides partial knowledge about the current network status, our model uses historical knowledge about the network and data to reduce the real-time investigation overhead while ensuring near optimal throughput for each transfer. Our novel approach is tested over different networks with different datasets and outperformed its closest competitor by 1.7× and the default case by 5×. It also achieved up to 93% accuracy compared with the optimal achievable throughput possible on those networks. Md. S. Q. Zulkar Nine, Kemal Guner, Ziyun Huang 0001, Xiangyu Wang 0017, Jinhui Xu 0001, Tevfik Kosar |
IEEE BigData | 6 |
| 2017 | Efficient Distributed Coordination at WAN-ScaleabstractTraditional coordination services for distributed applications do not scale well over wide-area networks (WAN): centralized coordination fails to scale with respect to the increasing distances in the WAN, and distributed coordination fails to scale with respect to the number of nodes involved. We argue that it is possible to achieve scalability over WAN using a hierarchical coordination architecture and a smart token migration mechanism, and lay down the foundation of a novel design for a flexible-consistent coordination framework, called WanKeeper. We implemented WanKeeper based on the ZooKeeper API and deployed it over WAN as a proof of concept. Our experimental results based on the Yahoo! Cloud Serving Benchmark (YCSB), Apache BookKeeper replicated log service, and the Shared Cloud-backed File System (SCFS) show that WanKeeper provides multiple folds improvement in write/update performance in WAN compared to ZooKeeper, while keeping the same read performance. Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas, Bekir O. Turkkan, Tevfik Kosar |
ICDCS | 5 |
| 2017 | Poster: Application-Layer Optimization of Performance vs Energy in Mobile Network I/OabstractThe number of smartphone users globally has already exceeded 2 Billion, and this number is expected to reach 3 Billion by 2020 [2]. It is also estimated that smartphone mobile data traffic (cellular + WiFi) will reach 370 Exabytes per year by that time, exceeding PC Internet traffic the first time in the history [4]. Kemal Guner, Tevfik Kosar |
MobiSys | 2 |
| 2016 | HARP: predictive transfer optimization based on historical analysis and real-time probingabstractIncreasingly data-intensive scientific and commercial applications require frequent movement of large datasets from one site to the other. Despite the growing capacity of the networking capacity, these data movements rarely achieve the promised data transfer rates of the underlying physical network due to the poorly tuned data transfer protocols. Accurately and efficiently tuning the data transfer protocol parameters in a dynamically changing network environment is a big challenge and still an open research problem. In this paper, we present predictive end-to-end data transfer optimization algorithms based on historical data analysis and real-time background traffic probing, dubbed HARP. Most of the existing work in this area is solely based on real time network probing, which either cause too much sampling overhead or fail to accurately predict the correct transfer parameters. Combining historical data analysis with real time sampling enables our algorithms to tune the application level data transfer parameters accurately and efficiently to achieve close-to-optimal end-to-end data transfer throughput with very low overhead. Our experimental analysis over a variety of network settings shows that HARP outperforms existing solutions by up to 50% in terms of the achieved throughput. Engin Arslan, Kemal Guner, Tevfik Kosar |
SC | 3 |
| 2016 | Application-Level Optimization of Big Data Transfers through Pipelining, Parallelism and ConcurrencyabstractIn end-to-end data transfers, there are several factors affecting the data transfer throughput, such as the network characteristics (e.g., network bandwidth, round-trip-time, background traffic); end-system characteristics (e.g., NIC capacity, number of CPU cores and their clock rate, number of disk drives and their I/O rate); and the dataset characteristics (e.g., average file size, dataset size, file size distribution). Optimization of big data transfers over inter-cloud and intra-cloud networks is a challenging task that requires joint-consideration of all of these parameters. This optimization task becomes even more challenging when transferring datasets comprised of heterogeneous file sizes (i.e., large files and small files mixed). Previous work in this area only focuses on the end-system and network characteristics however does not provide models regarding the dataset characteristics. In this study, we analyze the effects of the three most important transfer parameters that are used to enhance data transfer throughput: pipelining,parallelism and concurrency. We provide models and guidelines to set the best values for these parameters and present two different transfer optimization algorithms that use the models developed. The tests conducted over high-speed networking and cloud testbeds show that our algorithms outperform the most popular data transfer tools like Globus Online and UDT in majority of the cases. Esma Yildirim, Engin Arslan, JangYoung Kim, Tevfik Kosar |
IEEE Trans. Cloud Comput. | 4 |
| 2015 | Energy-aware data transfer algorithmsabstractThe amount of data moved over the Internet per year has already exceeded the Exabyte scale and soon will hit the Zettabyte range. To support this massive amount of data movement across the globe, the networking infrastructure as well as the source and destination nodes consume immense amount of electric power, with an estimated cost measured in billions of dollars. Although considerable amount of research has been done on power management techniques for the networking infrastructure, there has not been much prior work focusing on energy-aware data transfer algorithms for minimizing the power consumed at the end-systems. We introduce novel data transfer algorithms which aim to achieve high data transfer throughput while keeping the energy consumption during the transfers at the minimal levels. Our experimental results show that our energy-aware data transfer algorithms can achieve up to 50% energy savings with the same or higher level of data transfer throughput. Ismail Alan, Engin Arslan, Tevfik Kosar |
SC | 3 |
| 2014 | Energy-Aware Data Transfer TuningabstractThe annual electricity consumed by data transfers in the U.S. is estimated to be 20 Terawatt hours, which translates to around 4 billion U.S. Dollars per year. There has been considerable amount of prior work looking at power management and energy efficiency in hardware and software systems, and more recently in power-aware networking. Despite the growing body of research in power management techniques for the networking infrastructure, there has been no prior work (to the best of our knowledge), focusing on saving energy at the end systems(sender and receiver nodes) during the data transfer. We argue that although network-only approaches are part of the solution, the end-system power management is a key in optimizing energy efficiency of the data transfers, which has been long ignored. In this paper, we analyze various factors that will affect the power consumption in end-to-end data transfers, such as the level of parallelism, concurrency and pipelining. Our results show that significant amount of energy savings can be achieved at the end-systems during data transfer with no or minimal performance penalty. Ismail Alan, Engin Arslan, Tevfik Kosar |
CCGRID | 3 |
| 2013 | Dynamic Protocol Tuning Algorithms for High Performance Data Transfers
Engin Arslan, Brandon Ross, Tevfik Kosar |
Euro-Par | 3 |
| 2013 | Modeling throughput sampling size for a cloud-hosted data scheduling and optimization service
Esma Yildirim, JangYoung Kim, Tevfik Kosar |
Future Gener. Comput. Syst. | 3 |
| 2012 | Guest Editors' Introduction: Special Issue on Data-Intensive Computing in the Clouds
Tevfik Kosar, Ioan Raicu |
J. Grid Comput. | 1 |
| 2012 | End-to-End Data-Flow Parallelism for Throughput Optimization in High-Speed Networks
Esma Yildirim, Tevfik Kosar |
J. Grid Comput. | 2 |
| 2011 | Prediction of Optimal Parallelism Level in Wide Area Data TransfersabstractWide area data transfer may be a major bottleneck for the end-to-end performance of distributed applications. A practical way of increasing the wide area throughput at the application layer is using multiple parallel streams. Although increased number of parallel streams may yield much better performance than using a single stream, overwhelming the network by opening too many streams may have an inverse effect. The congestion created by excess number of streams may cause a drop down in the throughput achieved. Hence, it is important to decide on the optimal number of streams without congesting the network. Predicting this "optimum” number is not straightforward, since it depends on many parameters specific to each individual transfer. Generic models that try to predict this number either rely too much on historical information or fail to achieve accurate predictions. In this paper, we present a set of new models which aim to approximate the optimal number with least history information and lowest prediction overhead. An algorithm is introduced to select the best combination of historic information to do the prediction for evaluation purposes as well as optimizing prediction by reducing error rate. We measure the feasibility and accuracy of the proposed prediction models by comparing to actual GridFTP data transfer by using little historical information and have seen that we could predict the throughput of parallel streams accurately and find a very close approximation of the optimal stream number. Esma Yildirim, Dengpan Yin, Tevfik Kosar |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2011 | A Data Throughput Prediction and Optimization Service for Widely Distributed Many-Task ComputingabstractIn this paper, we present the design and implementation of an application-layer data throughput prediction and optimization service for many-task computing in widely distributed environments. This service uses multiple parallel TCP streams to improve the end-to-end throughput of data transfers. A novel mathematical model is developed to determine the number of parallel streams, required to achieve the best network performance. This model can predict the optimal number of parallel streams with as few as three prediction points. We implement this new service in the Stork Data Scheduler, where the prediction points can be obtained using Iperf and GridFTP samplings. Our results show that the prediction cost plus the optimized transfer time is much less than the nonoptimized transfer time in most cases. As a result, Stork data transfer jobs with optimization service can be completed much earlier, compared to nonoptimized data transfer jobs. Dengpan Yin, Esma Yildirim, Sivakumar Kulasekaran, Brandon Ross, Tevfik Kosar |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2010 | Design, implementation and use of a simulation data archive for coastal scienceabstractWith many researchers now having easy access to supercomputers, coastal scientists are able to develop and run simulations that model the physical and ecological processes in ocean or nearshore areas in a distributed and collaborative environment. However, the increase in capacity of computational resources does not lead directly to a rapid improvement of the simulations themselves. Instead, it brings a new challenge that motivates scientists to fully utilize the huge amount of simulation data created in supercomputers, thus fostering advanced scientific research. Driven by the urgent need for in-depth investigations in Louisiana coastal areas, especially during hurricane seasons, a data center, which provides research communities with scientific data resources on demand, is imperative. In this paper, we present the design, implementation and use of such a simulation data archive for coastal science. The simulation data archive is capable of providing interfaces based on the requirements of user groups and its application incorporates multiple use cases. The enabling technology, as well as the challenges in the development of this simulation data archive, are also described in this paper. Harsha Bhagawaty, Lei Jiang 0010, Sreekanth Pothanis, Gabrielle Allen, Nathan Brener, Tevfik Kosar |
HPDC | 6 |
| 2010 | Toward a Reliable Distributed Data Management SystemabstractModern collaborative science has placed increasing burden on data management infrastructure to handle the increasingly large data archives generated. Beside functionality, reliability and availability are also key factors in delivering a data management system that can efficiently and effectively meet the challenges posed and compounded by the unbounded increase in the size of data archive generated by scientific applications. In this paper, we present our work on increasing and improving reliability and availability in the data management system we designed for the PetaShare project, we also discuss our work on benchmarking the performance and scalability of metadata management system in PetaShare project. Ismail Akturk, Xinqi Wang, Tevfik Kosar |
ISPDC | 3 |
| 2009 | Early Error Detection and Classification in Data Transfer SchedulingabstractData transfer in distributed environment is prone to frequent failures resulting from back-end system level problems, like connectivity failure which is technically untraceable by users.Error messages are not logged efficiently, and sometimes are not relevant/useful from users' point-of-view.Our study explores the possibility of an efficient error detection and reporting system for such environments.Besides, early error detection and error classification have great importance in organizing data placement jobs.It is necessary to have well defined error detection and error reporting methods to increase the usability and serviceability of existing data transfer protocols and data management systems. Prior knowledge about the environment and awareness of the actual reason behind a failure would enable data placement scheduler to make better and accurate decisions.We investigate the applicability of proposed early error detection and error classification techniques to improve arrangement of data placement jobs and to enhance decision making of data placement schedulers. Mehmet Balman, Tevfik Kosar |
CISIS | 2 |
| 2009 | Dynamic Adaptation of Parallelism Level in Data Transfer SchedulingabstractWe discuss dynamic parameter tuning in wide-area data transfers for efficient utilization of available network capacity and optimized end-to-end application performance.Impacts of parallel TCP streams as well as concurrent data transfer jobs running simultaneously have been studied.We present an adaptive approach for tuning parallelism level of data placement jobs in distributed environments.The adaptive data scheduling includes dynamically setting parameters of data placement jobs.The proposed methodology operates without depending on any external profiles to adapt to changing network conditions. Mehmet Balman, Tevfik Kosar |
CISIS | 2 |
| 2009 | Cross-domain metadata management in data intensive distributed computing environmentabstractAs the the size of scientific datasets grows, it becomes imperative that cross-domain metadata management system needs to be developed to facilitate interdisciplinary scientific research. Three key issues need to be addressed: the development of a cross-domain metadata schema; the implementation of a metadata management system based on this schema; the integration of the metadata system into existing infrastructure with reasonable performance and scalability. In this paper, we give an overview of the research we have done as part of the PetaShare project to address the above mentioned problems. Xinqi Wang, Ismail Akturk, Tevfik Kosar |
CLUSTER | 3 |
| 2009 | An innovative application execution toolkit for multicluster gridsabstractMulticluster grids provide one promising solution to satisfying growing computation demands of compute-intensive applications by collaborating various networked clusters. However, it is challenging to seamlessly integrate all participating clusters in different domains into a virtual computation platform. In order to take full advantages of multicluster grids capability, computer scientists need to deal with how to collaborate practically and efficiently participating autonomic systems to execute Grid-enabled applications. We make efforts on grid resource management and implement a toolkit called Pelecanus to improve the overall performance of application execution in multicluster grids environment. The Pelecanus takes advantages of the DA-TC (Dynamic Assignment with Task Containers) execution model to improve resource interoperability and enhance application execution and monitoring. Experiments show that it can significantly reduce turnaround time and increase resource utilization for certain applications with large number of sequential jobs. Zhifeng Yun, Zhou Lei 0001, Gabrielle Allen, Daniel S. Katz, Tevfik Kosar, Shantenu Jha, J. Ramanujam |
CLUSTER | 5 |
| 2009 | Design and Implementation of Metadata System in PetaShare
Xinqi Wang, Tevfik Kosar |
SSDBM | 2 |
| 2009 | A new paradigm: Data-aware scheduling in grid computing
Tevfik Kosar, Mehmet Balman |
Future Gener. Comput. Syst. | 1 |
| 2008 | Semantic Enabled Metadata Framework for Data GridsabstractWe designed a semantic enabled metadata framework using ontology for multi-disciplinary and multi-institutional large scale scientific data sets in a Data Grid setting. There are two main issues we intend to address: data integration for semantically and physically heterogeneous distributed knowledge stores, and semantic reasoning for data verification and inference in such a setting. This framework enables data interoperability between otherwise semantically incompatible data sources, cross-domain query capabilities and multi-source knowledge extraction. In this paper, we present the basic system architecture for this framework, as well as an initial implementation. Dayong Huang, Xinqi Wang, Gabrielle Allen, Tevfik Kosar |
CISIS | 4 |
| 2006 | What makes workflows work in an opportunistic environment?abstractAbstract In this paper, we examine the issues of workflow mapping and execution in opportunistic environments such as the Grid. As applications become ever more complex, the process of choosing the appropriate resources and successfully executing the application components becomes ever more difficult. This may include extension or reduction of the initial workflow mapping as necessary for the actual execution. In this paper, we focus on the interplay between a workflow‐mapping component that plans the high‐level resource assignments and the workflow executor that oversees the component execution. We concentrate particularly on issues of data management and we draw from the experiences with mapping and execution systems: Pegasus, DAGMan and Stork. Copyright © 2005 John Wiley & Sons, Ltd. Ewa Deelman, Tevfik Kosar, Carl Kesselman, Miron Livny |
Concurr. Comput. Pract. Exp. | 2 |
| 2006 | Building reliable and efficient data transfer and processing pipelinesabstractAbstract Scientific distributed applications have an increasing need to process and move large amounts of data across wide area networks. Existing systems either closely couple computation and data movement, or they require substantial human involvement during the end‐to‐end process. We propose a framework that enables scientists to build reliable and efficient data transfer and processing pipelines. Our framework provides a universal interface to different data transfer protocols and storage systems. It has sophisticated flow control and recovers automatically from network, storage system, software and hardware failures. We successfully used data pipelines to replicate and process three terabytes of DPOSS astronomy image dataset and several terabytes of WCER educational video dataset. In both cases, the entire process was performed without any human intervention and the data pipeline recovered automatically from various failures. Copyright © 2005 John Wiley & Sons, Ltd. Tevfik Kosar, George Kola, Miron Livny |
Concurr. Comput. Pract. Exp. | 1 |
| 2005 | Faults in Large Distributed Systems and What We Can Do About Them
George Kola, Tevfik Kosar, Miron Livny |
Euro-Par | 2 |
| 2005 | A framework for reliable and efficient data placement in distributed computing systems
Tevfik Kosar, Miron Livny |
J. Parallel Distributed Comput. | 1 |
| 2004 | A client-centric grid knowledgebaseabstractGrid computing brings with it additional complexities and unexpected failures. Just keeping track of our jobs traversing different grid resources before completion can at times become tricky. We introduce a client-centric grid knowledgebase that keeps track of the job performance and failure characteristics on different grid resources as observed by the client. We present the design and implementation of our prototype grid knowledgebase and evaluate its effectiveness on two real life grid data processing pipelines: NCSA image processing pipeline and WCER video processing pipeline. It enabled us to easily extract useful job and resource information and interpret them to make better scheduling decisions. Using it, we were able to understand failures better and were able to devise innovative methods to automatically avoid and recover from failures and dynamically adapt to grid environment improving fault-tolerance and performance. George Kola, Tevfik Kosar, Miron Livny |
CLUSTER | 2 |
| 2004 | Profiling Grid Data Transfer Protocols and Servers
George Kola, Tevfik Kosar, Miron Livny |
Euro-Par | 2 |
| 2004 | Stork: Making Data Placement a First Class Citizen in the GridabstractTodays scientific applications have huge data requirements which continue to increase drastically every year. These data are generally accessed by many users from all across the the globe. This implies a major necessity to move huge amounts of data around wide area networks to complete the computation cycle, which brings with it the problem of efficient and reliable data placement. The current approach to solve this problem of data placement is either doing it manually, or employing simple scripts which do not have any automation or fault tolerance capabilities. Our goal is to make data placement activities first class citizens in the Grid just like the computational jobs. They will be queued, scheduled, monitored, managed, and even check-pointed. More importantly, it will be made sure that they complete successfully and without any human interaction. We also believe that data placement jobs should be treated differently from computational jobs, since they may have different semantics and different characteristics. For this purpose, we have developed Stork, a scheduler for data placement activities in the grid. Tevfik Kosar, Miron Livny |
ICDCS | 1 |
| 2004 | A fully automated fault-tolerant system for distributed video processing and off-site replicationabstractDifferent fields including biomedical-engineering, educational research and geology have an increasing need to process large amounts of video and make them electronically available at different locations. So far, this has been a failure-prone tedious operation with an operator needed to babysit the processing and off-site replication of processed video. In this work, we developed a fault-tolerant system that handles large scale processing and replication of digital video in a fully automated manner. The system is highly resilient and handles a variety of hardware, software and network failures making it possible to process videos using commodity clusters or grid resources. Finally, we discuss how the system is being used in educational research to process several hundred terabytes of video. George Kola, Tevfik Kosar, Miron Livny |
NOSSDAV | 2 |
| 2003 | Managing eBusiness on Demand SLA Contracts in Business Terms Using the Cross-SLA Execution Manager SAMabstractIt is imperative for a competitive e-business service provider to be positioned to manage the execution of its service level agreement (SLA) contracts in business terms (e.g., minimizing financial penalties for service-level violations, maximizing service-level measurement based customer satisfaction metrics). This paper briefly describes the design rationale of an integrated set of business-oriented service level management (SLM) technologies under development in the SAM project at IBM TJ Watson Research Center. The e-business SLA execution manager SAM, (1) enables the provider to deploy an effective means of capturing and managing contractual SLA data as well as provider-facing non-contractual SLM data; (2) assists service personnel to prioritize the processing of action-demanding quality management alerts as per the provider's SLM objectives; and (3) automates the prioritization and execution management Of approved SLM processes on behalf of the provider, including assigning SLM tasks to service personnel. Melissa J. Buco, Rong Chang 0001, Laura Z. Luan, Christopher Ward, Joel L. Wolf, Philip S. Yu, Tevfik Kosar, Syed Umair Ahmed Shah |
ISADS | 7 |
| 2003 | A Framework for Self-optimizing, Fault-tolerant, High Performance Bulk Data Transfers in a Heterogeneous Grid EnvironmentabstractThe drastic increase in the data requirements of scientific applications combined with an increasing trend towards collaborative research has resulted in the need to transfer large amounts of data among the participating sites. The general approach to transferring such large amounts of data has been to either dump data to tapes and mail them or employ scripts with an operator at each site to babysit the transfers to deal with failures. We introduce a framework which automates the whole process of data movement between different sites. The framework does not require any human intervention and it can recover automatically from various kinds of storage system, network, and software failures, guaranteeing completion of the transfers. The framework has sophisticated monitoring and tuning capability that increases the performance of the data transfers on the fly. The framework also generates on-the-fly visualization of the transfers making identification of problems and bottlenecks in the system simple. Tevfik Kosar, George Kola, Miron Livny |
ISPDC | 1 |