EDBT 2026 Demo / reviewers in the wild / expert
Zhengchun Liu
dblp:138/8484
· DBLP profile ↗
24ranked-venue papers
7as first author
12since 2021 · last 2025
0000-0002-6647-4423ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 6 first-author · 7 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Artificial intelligence and machine learning · 1Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PBench: Workload Synthesizer with Real Statistics for Cloud Analytics BenchmarkingabstractCloud service providers commonly use standard benchmarks like TPC-H and TPC-DS to evaluate and optimize cloud data analytics systems. However, these benchmarks rely on fixed query patterns and fail to capture real execution statistics of production cloud workloads. Although some cloud database vendors have recently released real workload traces, these traces alone do not qualify as benchmarks, as they typically lack essential components (i.e., queries and databases). To overcome this limitation, this paper studies a new problem of workload synthesis with real statistics , which generates synthetic workloads that closely approximate real execution statistics, including key performance metrics and operator distributions. To address this problem, we propose PBench, a novel workload synthesizer that constructs synthetic workloads by (1) selecting and combining workload components from existing benchmarks and (2) augmenting new workload components. This paper studies the key challenges in PBench. First, we address the challenge of balancing performance metrics and operator distributions by introducing a multi-objective optimization-based component selection method. Second, to capture the temporal dynamics of real workloads, we design a timestamp assignment method that progressively reines workload timestamps. Third, to handle the disparity between the original workload and the candidate workload, we propose a component augmentation approach that leverages large language models (LLMs) to generate additional workload components while maintaining statistical idelity. Experimental results show that PBench reduces approximation error by up to 6X compared to state-of-the-art methods. Chunwei Liu, Bhuvan Urgaonkar, Zhengle Wang, Magnus Mueller, Chao Zhang 0034, Songyue Zhang, Pascal Pfeil, Dominik Horn, Zhengchun Liu, Davide Pagano, Tim Kraska, Samuel Madden 0001, Ju Fan |
Proc. VLDB Endow. | 10 |
| 2024 | MalleTrain: Deep Neural Networks Training on Unfillable Supercomputer NodesabstractFirst-come first-serve scheduling can result in substantial (up to 10%) of transiently idle nodes on supercomputers. Recognizing that such unfilled nodes are well-suited for deep neural network (DNN) training, due to the flexible nature of DNN training tasks, Liu et al. proposed that the re-scaling DNN training tasks to fit gaps in schedules be formulated as a mixed-integer linear programming (MILP) problem, and demonstrated via simulation the potential benefits of the approach. Here, we introduce MalleTrain, a system that provides the first practical implementation of this approach and that furthermore generalizes it by allowing it to be used even for DNN training applications for which model information is unknown before runtime. Key to this latter innovation is the use of a lightweight online job profiling advisor (JPA) to collect critical scalability information for DNN jobs---information that it then employs to optimize resource allocations dynamically, in real time. We describe the MalleTrain architecture and present the results of a detailed experimental evaluation on a supercomputer GPU cluster and several representative DNN training workloads, including neural architecture search and hyperparameter optimization. Our results not only confirm the practical feasibility of leveraging idle supercomputer nodes for DNN training but improve significantly on prior results, improving training throughput by up to 22.3% without requiring users to provide job scalability information. Feng Yan 0001, Lei Yang 0001, Ian T. Foster, Michael E. Papka, Zhengchun Liu, Rajkumar Kettimuthu |
ICPE | 6 |
| 2024 | Why TPC Is Not Enough: An Analysis of the Amazon Redshift FleetabstractDatabase research and development is heavily influenced by benchmarks, such as the industry-standard TPC-H and TPC-DS for analytical systems. However, these twenty-year-old benchmarks neither capture how databases are deployed nor what workloads modern cloud data warehouse systems face these days. In this paper, we summarize well-known, confirm suspected, and unearth novel discrepancies between TPC-H/DS and actual workloads using empirical data. We base our analysis on telemetrics from Amazon Redshift - one of the largest cloud data warehouse deployments. Among others, we show how write-heavy data pipelines are prominent, workloads vary over time (in both load and type), queries are repetitive, and how most properties of queries or workloads experience very long tailed distributions. We conclude that data warehouse benchmarks, just like database systems, need to become more holistic and stop focusing solely on query engine performance. Finally, we publish a dataset containing query statistics of 200 randomly selected Redshift serverless and provisioned instances (each) over a three-month period, as a basis for building more realistic benchmarks. Alexander van Renen, Dominik Horn, Pascal Pfeil, Kapil Vaidya, Wenjian Dong, Murali Narayanaswamy, Zhengchun Liu, Andreas Kipf, Tim Kraska |
Proc. VLDB Endow. | 7 |
| 2023 | FreeTrain: A Framework to Utilize Unused Supercomputer Nodes for Training Neural NetworksabstractSupercomputer scheduling policies commonly result in many transient idle nodes, a phenomenon that is only partially alleviated by backfill scheduling methods that promote small jobs to run before large jobs. Here we describe how to realize a novel use for these otherwise wasted resources, namely, deep neural network (DNN) training. This important workload is easily organized as many small fragments that can be configured dynamically to fit essentially any node × time hole in a supercomputer's schedule. We describe how the task of rescaling suitable DNN training tasks to fit dynamically changing holes can be formulated as a deterministic mixed integer linear programming (MILP)-based resource allocation algorithm, and show that this MILP problem can be solved efficiently at run time. We show further how this MILP problem can be adapted to optimize for administrator- or user-defined metrics. We validate our method with supercomputer scheduler logs and different DNN training scenarios, and demonstrate efficiencies of up to 93% compared with running the same training tasks on dedicated nodes. Our method thus enables substantial supercomputer resources to be allocated to DNN training with no impact on other applications. Zhengchun Liu, Rajkumar Kettimuthu, Michael E. Papka, Ian T. Foster |
CCGrid | 1 |
| 2023 | Tomo2Mesh: Fast Porosity Mapping and Visualization for Synchrotron TomographyabstractApplications of X-ray computed tomography (CT) in porosity characterization of engineering materials often involve an extensive data analysis workflow. This workflow includes CT reconstruction of raw projection data, binarization, labeling, and mesh extraction. Mapping porosity in larger samples presents a significant computational challenge, as it requires processing gigabytes of raw data to extract porosity information, which becomes a critical bottleneck in the analysis. In this study, we present algorithms and an implementation of an end-to-end porosity mapping framework. Our framework processes raw projection data obtained from a synchrotron CT instrument, generating a porosity map and a visualization in the form of a triangular face mesh. To achieve this objective, we introduce a novel subset reconstruction scheme for X-ray CT, combining filtered back-projection and a convolutional neural network. This scheme allows us to reconstruct subsets of a tomography object with arbitrary shapes. Building upon this scheme, we have developed a fast and efficient framework for porosity mapping. Initially, our framework detects potential voids by performing a coarse reconstruction on down-sampled projections. Subsequently, we enhance the shape of these voids by reconstructing selected subsets from the original raw data, providing higher detail. To evaluate the performance of our framework, we measured the processing time from raw data to a triangular face mesh across multiple visualization scenarios. Our experiments were conducted on a single high-performance workstation equipped with a GPU. The results demonstrate that our framework enables the visualization of local porosity within an 8-gigavoxel CT volume (12 gigabytes raw data) in just 1 to 2 minutes. Moreover, for a larger 64-gigavoxel CT volume (100 gigabytes of raw data), the visualization can be generated within 3 to 7 minutes, showcasing the efficiency of our approach. Aniket Tekawade, Viktor V. Nikitin, Yashas Satapathy, Zhengchun Liu, Peter Kenesei, Weijian Zheng, Francesco De Carlo, Ian T. Foster, Rajkumar Kettimuthu |
e-Science | 4 |
| 2023 | RAPIDS: Reconciling Availability, Accuracy, and Performance in Managing Geo-Distributed Scientific DataabstractIn modern science, big data plays an increasingly important role. Many scientific applications, such as running simulations on supercomputers or conducting experiments on advanced instruments, produce huge amount of data at unprecedented speed. Analyzing and understanding such big data is the key for scientists to make scientific breakthroughs. However, data might become unavailable for scientists to access when outages or maintenance of the storage system occur, which severely hinders scientific discovery. To improve the data availability, data duplication and erasure coding (EC) are often used. But as the scientific data gets larger, using these two methods can cause considerable storage and network overhead. Lipeng Wan 0001, Jieyang Chen, Xin Liang 0001, Ana Gainaru, Qian Gong, Qing Liu 0002, Ben Whitney, Joy Arulraj, Zhengchun Liu, Ian T. Foster, Scott Klasky |
HPDC | 9 |
| 2023 | Moving small files in a networked environment
Chao Jin 0001, David Abramson 0001, Jake Carroll, Zhengchun Liu, Rajkumar Kettimuthu |
Future Gener. Comput. Syst. | 4 |
| 2022 | fairDMS: Rapid Model Training by Data and Model ReuseabstractExtracting actionable information rapidly from data produced by instruments such as the Linac Coherent Light Source (LCLS-II) and Advanced Photon Source Upgrade (APS-U) is becoming ever more challenging due to high (up to TB/s) data rates. Conventional physics-based information retrieval methods are hard-pressed to detect interesting events fast enough to enable timely focusing on a rare event or correction of an error. Machine learning (ML) methods that learn cheap surrogate classifiers present a promising alternative, but can fail catastrophically when changes in instrument or sample result in degradation in ML performance. To overcome such difficulties, we present a new data storage and ML model training architecture designed to organize large volumes of data and models so that when model degradation is detected, prior models and/or data can be queried rapidly and a more suitable model retrieved and fine-tuned for new conditions. We show that our approach can achieve up to 100x data labelling speedup compared to the current state-of-the-art, 200x improvement in training speed, and 92x speedup in-terms of end-to-end model updating time. Hemant Sharma, Rajkumar Kettimuthu, Peter Kenesei, Dennis Trujillo, Antonino Miceli, Ian T. Foster, Ryan N. Coffee, Jana Thayer, Zhengchun Liu |
CLUSTER | 10 |
| 2022 | What does Inter-Cluster Job Submission and Execution Behavior Reveal to Us?abstractModern High Performing Computing (HPC) facil-ities have multiple computing clusters that serve different pur-poses. These include large-scale computing clusters and smaller data visualization and analysis clusters, which are meant to shift the load of data analytics jobs from the large-scale systems. We perform the first in-depth characterization of cross-cluster behavior of users and jobs and provide an analysis of three inter-related systems at the Argonne Leadership Computing Facility (ALCF). Our analysis reveals interesting trends related to the resource utilization and predictability of user and job behavior across different clusters. Tirthak Patel, Devesh Tiwari, Rajkumar Kettimuthu, William E. Allcock, Paul M. Rich, Zhengchun Liu |
CLUSTER | 6 |
| 2022 | SciStream: Architecture and Toolkit for Data Streaming between Federated Science InstrumentsabstractModern scientific instruments, such as detectors at synchrotron light sources, generate data at such high rates that online processing is needed for data reduction, feature detection, experiment steering, and other purposes. The same high data rates also demand memory-to-memory streaming from instrument to remote computer, because local computational capacity is limited and data transmissions that engage the file system introduce unacceptable latencies. But efficient and secure memory-to-memory data streaming is challenging to realize in practice, because of a lack of direct external network connectivity for scientific instruments and because of authentication and security requirements. In response, we propose here SciStream, a middlebox-based architecture with control protocols to enable efficient and secure memory-to-memory data streaming between producers and consumers that lack direct network connectivity. We describe the protocols that SciStream uses to establish authenticated and transparent connections between producers and consumers, and we discuss the experiments that we have conducted to evaluate alternative implementation approaches for key SciStream components. Experiments on the Chameleon cloud show that SciStream improves the throughput of a streaming pipeline by an order of magnitude compared with state-of-the-art data transfer methods and adds only ~4μsec latency compared with an ideal scenario in which producers and consumers have direct external connectivity. Joaquin Chung 0001, Wojciech Zacherek, A. J. Wisniewski, Zhengchun Liu, Tekin Bicer, Rajkumar Kettimuthu, Ian T. Foster |
HPDC | 4 |
| 2021 | 3d Autoencoders For Feature Extraction In X-Ray TomographyabstractReal-time steering of time-resolved or in-situ X-ray tomography requires capturing changes in morphological descriptors in a sample (e.g., porosity, particle size, and crack width) during continuous data acquisition. Segmentation of 2D or 3D images followed by quantitative measurement is the conventional method for tracking changes in these descriptors with respect to a previous time-step or a 3D search in a volume. However, image segmentation is expensive. As a faster and unsupervised alternative, we propose a feature-extraction approach using a convolutional autoencoder, where the latent space of the encoder responds to relative changes in morphology without prior knowledge of the morphological descriptors. To test this approach in parametric experiments, we used a digital twin for micro-CT to generate realistic datasets of porous materials. We show that for some configurations of the autoencoder, its n-dimensional latent space encodes a sample's local porosity metrics while disregarding contrast information (relative proportion of absorption and phase contrast) determined by the imaging modality and not the sample. Through dimensionality reduction, the vector's response is visualized in 2D space to find clusters of data with similar porosity metrics. Our approach can extract features from grayscale tomographic data more than 4x faster than a segmentation + pore analysis workflow. Aniket Tekawade, Zhengchun Liu, Peter Kenesei, Tekin Bicer, Francesco De Carlo, Rajkumar Kettimuthu, Ian T. Foster |
ICIP | 2 |
| 2021 | End-to-end online performance data capture and analysis for scientific workflows
George Papadimitriou 0002, Cong Wang 0014, Karan Vahi, Rafael Ferreira da Silva, Anirban Mandal, Zhengchun Liu, Rajiv Mayani, Mats Rynge, Mariam Kiran, Vickie E. Lynch, Rajkumar Kettimuthu, Ewa Deelman, Jeffrey S. Vetter, Ian T. Foster |
Future Gener. Comput. Syst. | 6 |
| 2020 | Characterization and identification of HPC applications at leadership computing facilityabstractHigh Performance Computing (HPC) is an important method for scientific discovery via large-scale simulation, data analysis, or artificial intelligence. Leadership-class supercomputers are expensive, but essential to run large HPC applications. The Petascale era of supercomputers began in 2008, with the first machines achieving performance in excess of one petaflops, and with the advent of new supercomputers in 2021 (e.g., Aurora, Frontier), the Exascale era will soon begin. However, the high theoretical computing capability (i.e., peak FLOPS) of a machine is not the only meaningful target when designing a supercomputer, as the resources demand of applications varies. A deep understanding of the characterization of applications that run on a leadership supercomputer is one of the most important ways for planning its design, development and operation. Zhengchun Liu, Ryan Lewis, Rajkumar Kettimuthu, Kevin Harms, Philip H. Carns, Nageswara S. V. Rao, Ian T. Foster, Michael E. Papka |
ICS | 1 |
| 2020 | Job characteristics on large-scale systems: long-term analysis, quantification, and implicationsabstractHPC workload analysis and resource consumption characteristics are the key to driving better operation practices, system procurement decisions, and designing effective resource management techniques. Unfortunately, the HPC community does not have easy accessibility to long-term introspective work-load analysis and characterization for production-scale HPC systems. This study bridges this gap by providing detailed long-term quantification, characterization, and analysis of job characteristics on two supercomputers: Intrepid and Mira. This study is one of the largest of its kind - covering trends and characteristics for over three billion compute hours, 750 thousand jobs, and spanning a decade. We confirm several long-held conventional wisdom, and identify many previously undiscovered trends and its implications. We also introduce a learning based technique to predict the resource requirement of future jobs with high accuracy, using features available prior to the job submission and without requiring any application-specific tracing or application-intrusive instrumentation. Tirthak Patel, Zhengchun Liu, Rajkumar Kettimuthu, Paul M. Rich, William E. Allcock, Devesh Tiwari |
SC | 2 |
| 2019 | Spatiotemporal Real-Time Anomaly Detection for Supercomputing SystemsabstractThe demands of increasingly large scientific application workflows lead to the need for more powerful supercomputers. As the scale of supercomputing systems have grown, the prediction of fault tolerance has become an increasingly critical area of study, since the prediction of system failures can improve performance by saving checkpoints in advance. We propose a real-time failure detection algorithm that adopts an event-based prediction model. The prediction model is a convolutional neural network that utilizes both traditional event attributes and additional spatio-temporal features. We present a case study using our proposed method with six years of reliability, availability, and serviceability event logs recorded by Mira, a Blue Gene/Q supercomputer at Argonne National Laboratory. In the case study, we have shown that our failure prediction model is not limited to predict the occurrence of failures in general. It is capable of accurately detecting specific types of critical failures such as coolant and power problems within reasonable lead time ranges. Our case study shows that the proposed method can achieve a F1score of 0.56 for general failures, 0.97 for coolant failures, and 0.86 for power failures. Qiao Kang, Ankit Agrawal 0001, Alok N. Choudhary, Alex Sim, Kesheng Wu, Rajkumar Kettimuthu, Pete Beckman, Zhengchun Liu, Wei-keng Liao |
IEEE BigData | 8 |
| 2019 | Data Transfer between Scientific Facilities - Bottleneck Analysis, Insights and OptimizationsabstractWide area file transfers play an important role in many science applications. File transfer tools typically deliver the highest performance for datasets with a small number of large files, but many science datasets consist of many small files. Thus it is important to understand the factors that contribute to the decrease in wide area data transfer performance for datasets with many small files. To this end, we (i) benchmark the performance of subsystems involved in end-to-end file transfer between two HPC facilities for a many-file dataset that is representative of production science transfers; (ii) characterize the per-file overhead introduced by different subsystems; (iii) identify potential dependencies and bottlenecks; (iv) study the effectiveness of transferring many files concurrently as a means of reducing per-file overheads; and (v) prototype a prefetching mechanism as an alternative of concurrency to reduce the per-file overhead on source storage system. We show that both concurrency and prefetching can help reduce the per-file overhead significantly. A reasonable level of concurrency combined with prefetching can bring the per-file overhead down to a negligible level. Yuanlai Liu, Zhengchun Liu, Rajkumar Kettimuthu, Nageswara S. V. Rao, Zizhong Chen, Ian T. Foster |
CCGRID | 2 |
| 2019 | Toward an Elastic Data Transfer InfrastructureabstractData transfer over wide area networks is an integral part of many science workflows that must, for example, move data from scientific facilities to remote resources for analysis, sharing, and storage. Yet despite continued enhancements in data transfer infrastructure (DTI), our previous analyses of approximately 40 billion GridFTP command logs collected over four years from the Globus transfer service show that data transfer nodes (DTNs) are idle (i.e., are performing no transfers) 94.3% of the time. On the other hand, we have also observed periods in which CPU resource scarcity negatively impacts DTN throughput. Motivated by the opportunity to optimize DTI performance, we present here an elastic DTI architecture in which the pool of nodes allocated to DTN activities expands and shrinks over time, based on demand. Our results show that this elastic DTI can save up to ~95% of resources compared with a typical static DTN deployment, with the median slowdown incurred remaining close to one for most of the evaluated scenarios. Joaquin Chung 0001, Zhengchun Liu, Rajkumar Kettimuthu, Ian T. Foster |
eScience | 2 |
| 2019 | Elastic Data Transfer Infrastructure (DTI) on the Chameleon CloudabstractMany science workflows are distributed in nature and rely on wide area networks (WANs) to move data between geographically distributed resources for analysis, sharing, and storage. In spite of continued enhancements in campus cyberinfrastructure, data transfer nodes (DTNs) are grossly underutilized. Our previous analysis of logs from 1,800 DTNs shows that they were completely idle for 94.3% of the time in 2017. Motivated by the opportunity to optimize DTN usage, here we present an elastic data transfer infrastructure (DTI) architecture in which the pool of nodes allocated to DTN activities expands and shrinks over time, based on demand. Our results show that this elastic DTI can save up to $\sim 95$% of resources compared with a typical static DTN deployment. Joaquin Chung 0001, Zhengchun Liu, Rajkumar Kettimuthu, Ian T. Foster |
ICNP | 2 |
| 2018 | Cross-geography scientific data transferring trends and behaviorabstractWide area data transfers play an important role in many science applications but rely on expensive infrastructure that often delivers disappointing performance in practice. In response, we present a systematic examination of a large set of data transfer log data to characterize transfer characteristics, including the nature of the datasets transferred, achieved throughput, user behavior, and resource usage. This analysis yields new insights that can help design better data transfer tools, optimize networking and edge resources used for transfers, and improve the performance and experience for end users. Our analysis shows that (i) most of the datasets as well as individual files transferred are very small; (ii) data corruption is not negligible for large data transfers; and (iii) the data transfer nodes utilization is low. Insights gained from our analysis suggest directions for further analysis. Zhengchun Liu, Rajkumar Kettimuthu, Ian T. Foster, Nageswara S. V. Rao |
HPDC | 1 |
| 2018 | A Comprehensive Study of Wide Area Data Movement at a Scientific Computing FacilityabstractWide-area data transfer is central to distributed science. Network capacity, data movement infrastructure, and tools in science environments continuously evolve to meet the requirements of distributed-science applications. Research and education (R&E) networks such as the U.S. Department of Energy's Energy Sciences network and Internet2 provide multiple 100 Gbps backbone networks. Large scientific facilities and research institutions have 100 Gbps wide-area network connectivity, and 10 Gbps wide-area network connectivity is common for a lot of R&E institutions. Many of these institutions employ Science DMZs, dedicated data transfer node(s), and high performance data movement tools to improve wide area data transfer performance. Large facilities may use 10 or more dedicated data transfer nodes to meet the needs of their users. In this work, we analyze various logs pertaining to wide area data transfers in and out of a large scientific facility to obtain insights on data transfer characteristics and behavior. We also show some of the inefficiencies in the state-of-the-art data movement tool and discuss approaches to address these inefficiencies. Zhengchun Liu, Rajkumar Kettimuthu, Ian T. Foster, Yuanlai Liu |
ICDCS | 1 |
| 2018 | Transferring a petabyte in a day
Rajkumar Kettimuthu, Zhengchun Liu, David Wheeler, Ian T. Foster, Katrin Heitmann, Franck Cappello |
Future Gener. Comput. Syst. | 2 |
| 2018 | Toward a smart data transfer node
Zhengchun Liu, Rajkumar Kettimuthu, Ian T. Foster, Pete Beckman |
Future Gener. Comput. Syst. | 1 |
| 2017 | A Mathematical Programming- and Simulation-Based Framework to Evaluate Cyberinfrastructure Design ChoicesabstractModern scientific experimental facilities such as x-ray light sources increasingly require on-demand access to large-scale computing for data analysis, for example to detect experimental errors or to select the next experiment. As the number of such facilities, the number of instruments at each facility, and the scale of computational demands all grow, the question arises as to how to meet these demands most efficiently and cost-effectively. A single computer per instrument is unlikely to be cost-effective because of low utilization and high operating costs. A single national compute facility, on the other hand, introduces a single point of failure and perhaps excessive communication costs. We introduce here methods for evaluating these and other potential design points, such as per-facility computer systems and a distributed multisite "superfacility." We use the U.S. Department of Energy light sources as a use case and build a mixed-integer programming model and a customizable superfacility simulator to enable joint optimization of design choices and associated operational decisions. The methodology and tools provide new insights into design choices for on-demand computing facilities for real-time analysis of scientific experiment data. The simulator can also be used to support facility operations, for example by simulating the impact of events such as outages. Zhengchun Liu, Rajkumar Kettimuthu, Sven Leyffer, Prashant Palkar, Ian T. Foster |
eScience | 1 |
| 2017 | Explaining Wide Area Data Transfer PerformanceabstractDisk-to-disk wide-area file transfers involve many subsystems and tunable application parameters that pose significant challenges for bottleneck detection, system optimization, and performance prediction. Performance models can be used to address these challenges but have not proved generally usable because of a need for extensive online experiments to characterize subsystems. We show here how to overcome the need for such experiments by applying machine learning methods to historical data to estimate parameters for predictive models. Starting with log data for millions of Globus transfers involving billions of files and hundreds of petabytes, we engineer features for endpoint CPU load, network interface card load, and transfer characteristics; and we use these features in both linear and nonlinear models of transfer performance, We show that the resulting models have high explanatory power. For a representative set of 30,653 transfers over 30 heavily used source-destination pairs ("edges''),totaling 2,053 TB in 46.6 million files, we obtain median absolute percentage prediction errors (MdAPE) of 7.0% and 4.6% when using distinct linear and nonlinear models per edge, respectively; when using a single nonlinear model for all edges, we obtain an MdAPE of 7.8%. Our work broadens understanding of factors that influence file transfer rate by clarifying relationships between achieved transfer rates, transfer characteristics, and competing load. Our predictions can be used for distributed workflow scheduling and optimization, and our features can also be used for optimization and explanation. Zhengchun Liu, Prasanna Balaprakash, Rajkumar Kettimuthu, Ian T. Foster |
HPDC | 1 |