Eric Pouyoul

dblp:41/601 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 87% Distributed systems · 13%
Computer networks
1 paper
Transport protocols and congestion control · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Transport protocols and congestion control
transport protocols
0.112012
Protocols for wide-area data-intensive applications: design and performance issues · SC 2012
High-performance computing
data-intensive computing
0.112012
Protocols for wide-area data-intensive applications: design and performance issues · SC 2012
High-performance computing › data transfer
wide-area data transfer
0.112012
Protocols for wide-area data-intensive applications: design and performance issues · SC 2012
Distributed systems › middleware
communication middleware
0.012012
Protocols for wide-area data-intensive applications: design and performance issues · SC 2012

Methods — techniques the papers use, named apart from their topics

task synchronization · 0.3flow control · 0.3connection management · 0.3
YearPublicationVenuePosition
2019 Analysis and Prediction of Data Transfer Throughput for Data-Intensive Workloads
abstract
Scientific workflows are increasingly transferring large amounts of data between high performance computing (HPC) systems. Even though these HPC systems are connected via high-speed dedicated networks and use dedicated data transfer nodes (DTNs), it is still difficult to predict the data transfer throughput because of variations in data transfer protocols, host configurations, performance of file systems, and overlapping workloads. In order to provide reliable performance prediction for better resource management and job scheduling, we need models for predicting data transfer throughput under real-world conditions. In this paper, we explore different machine learning approaches for building data-driven models to improve performance and prediction of large-scale data transfer throughput. In addition to the variables already collected by the network monitoring system, we also develop heuristics to derive additional metrics for improving the prediction accuracy. We use the prediction results to identify the importance of different network parameters in predicting the throughput for large-scale data transfers. Through extensive tests, we identify key network parameters, discover interesting variations among different HPC sites, and show that we can predict throughput with high accuracy. We also analyze our models and results to provide recommendations for improving the performance of big data transfers.
Devarshi Ghoshal, Kesheng Wu, Eric Pouyoul, Erich Strohmaier
IEEE BigData3
2019 Towards Securing Data Transfers Against Silent Data Corruption
abstract
Scientific applications generate large volumes of data that often needs to be moved between geographically distributed sites for collaboration or backup which has led to a significant increase in data transfer rates. As an increasing number of scientific applications are becoming sensitive to silent data corruption, end-to-end integrity verification has been proposed. It minimizes the likelihood of silent data corruption by comparing checksum of files at the source and the destination using secure hash algorithms such as MD5 and SHA1. In this paper, we investigate the robustness of existing end-to-end integrity verification approaches against silent data corruption and propose a Robust Integrity Verification Algorithm (RIVA) to enhance data integrity. Extensive experiments show that unlike existing solutions, RIVA is able to detect silent disk corruptions by invalidating file contents in page cache and reading them directly from disk. Since RIVA clears page cache and reads file contents directly from the disk, it incurs delay to execution time. However, by running transfer, cache invalidation, and checksum operations concurrently, RIVA is able to keep its overhead below 15% in most cases compared to the state-of-the-art solutions in exchange of increasing the robustness to silent data corruption. We also implemented dynamic transfer and checksum parallelism to overcome performance bottlenecks and observed more than 5x increase in RIVA's speed.
Batyr Charyyev, Ahmed Alhussen, Hemanta Sapkota, Eric Pouyoul, Mehmet Hadi Gunes, Engin Arslan
CCGRID4
2018 Calibers: A bandwidth calendaring paradigm for science workflows
Fatma Alali, Nathan Hanford, Eric Pouyoul, Rajkumar Kettimuthu, Mariam Kiran, Ben Mack-Crane, Brian Tierney, Yatish Kumar, Dipak Ghosal
Future Gener. Comput. Syst.3
2018 Enabling intent to configure scientific networks for high performance demands
Mariam Kiran, Eric Pouyoul, Anu Mercian, Brian Tierney, Chin Guok, Inder Monga
Future Gener. Comput. Syst.2
2018 mdtmFTP and its evaluation on ESNET SDN testbed
Liang Zhang 0009, Wenji Wu, Phil DeMar, Eric Pouyoul
Future Gener. Comput. Syst.4
2018 AmoebaNet: An SDN-enabled network service for big data science
Syed Asif Raza Shah, Wenji Wu, Qiming Lu, Liang Zhang 0009, Sajith Sasidharan, Phil DeMar, Chin Guok, John MacAuley, Eric Pouyoul, Seo-Young Noh
J. Netw. Comput. Appl.9
2016 Improving network performance on multicore systems: Impact of core affinities on high throughput flows
Nathan Hanford, Vishal Ahuja, Matthew K. Farrens, Dipak Ghosal, Mehmet Balman, Eric Pouyoul, Brian Tierney
Future Gener. Comput. Syst.6
2014 Impact of the end-system and affinities on the throughput of high-speed flows
abstract
Network throughput is scaling "up" to higher data transfer rates while processors are scaling "out" to multiple cores. As a result, network adapter "offloads" and performance "tuning" have received a good deal of attention lately. However, much of this attention is focused on the "how" and not the "why" of performance efficiency. There are two types of efficiencies that we have found particularly intriguing: First, processor core "affinity," or "binding" is fundamentally the choice of which processor core or cores handle certain tasks in a network- or I/O-heavy application running on a MIMD machine. Second, Ethernet "pause frames" slightly violate the "end-to-end" nature of TCP/IP in order to perform link-to-link flow control. The goal of our research is to delve deeper into why these tuning suggestions and this offload exist, and how they affect the end-to-end performance and efficiency of a single, large TCP flow.
Nathan Hanford, Vishal Ahuja, Matthew K. Farrens, Dipak Ghosal, Mehmet Balman, Eric Pouyoul, Brian Tierney
ANCS6
2012 Efficient data transfer protocols for big data
abstract
Data set sizes are growing exponentially, so it is important to use data movement protocols that are the most efficient available. Most data movement tools today rely on TCP over sockets, which limits flows to around 20Gbps on today's hardware. RDMA over Converged Ethernet (RoCE) is a promising new technology for high-performance network data movement with minimal CPU impact over circuit-based infrastructures. We compare the performance of TCP, UDP, UDT, and RoCE over high latency 10Gbps and 40Gbps network paths, and show that RoCE-based data transfers can fill a 40Gbps path using much less CPU than other protocols. We also show that the Linux zero-copy system calls can improve TCP performance considerably, especially on current Intel “Sandy Bridge”-based PCI Express 3.0 (Gen3) hosts.
Brian Tierney, Ezra Kissel, D. Martin Swany, Eric Pouyoul
eScience4
2012 Protocols for wide-area data-intensive applications: design and performance issues
abstract
Providing high-speed data transfer is vital to various data-intensive applications.While there have been remarkable technology advances to provide ultra-high-speed network bandwidth, existing protocols and applications may not be able to fully utilize the bare-metal bandwidth due to their inefficient design.We identify the same problem remains in the field of Remote Direct Memory Access (RDMA) networks.RDMA offloads TCP/IP protocols to hardware devices.However, its benefits have not been fully exploited due to the lack of efficient software and application protocols, in particular in wide-area networks.In this paper, we address the design choices to develop such protocols.We describe a protocol implemented as part of a communication middleware.The protocol has its flow control, connection management, and task synchronization.It maximizes the parallelism of RDMA operations.We demonstrate its performance benefit on various local and wide-area testbeds, including the DOE ANI testbed with RoCE links and InfiniBand links.
Yufei Ren, Dantong Yu, Shudong Jin, Thomas G. Robertazzi, Brian Tierney, Eric Pouyoul
SC7