EDBT 2026 Demo / reviewers in the wild / expert
Xingfu Wu
dblp:w/XingfuWu
· DBLP profile ↗
24ranked-venue papers
16as first author
9since 2021 · last 2026
0000-0001-8150-5171ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 10 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorTheory of computation · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SmartCap: Coordinated CPU-GPU Power Capping for Performance-Assurance Energy EfficiencyabstractPerformance prediction is essential for energy-efficient computing in heterogeneous computing systems that integrate CPUs and GPUs. However, traditional performance modeling methods often rely on exhaustive offline profiling, which becomes impractical due to the large setting space and the high cost of profiling large-scale applications. In this paper, we present OPEN, a framework consists of offline and online phases. The offline phase involves building a performance predictor and constructing an initial dense matrix. In the online phase, OPEN performs lightweight online profiling, and leverages the performance predictor with collaborative filtering to make performance prediction. We evaluate OPEN on multiple heterogeneous systems, including those equipped with A100 and A30 GPUs. Results show that OPEN achieves prediction accuracy up to 98.29\%. This demonstrates that OPEN effectively reduces profiling cost while maintaining high accuracy, making it practical for power-aware performance modeling in modern HPC environments. Overall, OPEN provides a lightweight solution for performance prediction under power constraints, enabling better runtime decisions in power-aware computing environments. Xingfu Wu, Valerie Taylor 0001, Michael E. Papka, Zhiling Lan |
ICS | 2 |
| 2026 | Beyond Throughput: Performance and Energy Insights of LLM Inference Across AI Accelerators
Giacomo Brunetta, Varuni Sastry 0001, Xingfu Wu, Valerie Taylor 0001, Michael E. Papka, Zhiling Lan |
IPDPS | 3 |
| 2025 | Evaluating Energy Efficiency of Ai Accelerators Using Two Mlperf BenchmarksabstractThe significantly increasing use of artificial intelligence (AI) has led to the availability of specialized AI accelerators, aiming to enhance the performance and energy efficiency of AI workloads. In this paper, we conduct an initial study to evaluate the energy requirements of four AI accelerators: Nvidia A100 GPUs, Intel Habana Gaudi Processing Units (HPUs), Graphcore Bow-Pod64 Intelligence Processing Units (IPUs), and GroqRack Language Processing Units (LPUs) using two popular MLPerf benchmarks: BERT-Large and ResNet50. We report the energy requirements for the two benchmarks to achieve a common MLPerfspecified target accuracy. The benchmarks and AI accelerators were chosen based on the following criteria: publicly available tools or libraries from the vendors to monitor power consumption, publicly available optimized models from the vendors, and access to the AI accelerators. Our experimental results indicate that for ResNet50, Intel Gaudi2 HPUs delivered the highest throughput, the lowest energy consumption for both training and inference, and the highest inference energy efficiency, while Graphcore demonstrated the highest training energy efficiency. For BERTLarge pre-training, Intel Gaudi2 outperformed both Nvidia A100 and Graphcore in terms of time, energy consumption, and training energy efficiency. However, for BERT-Large inference, Nvidia A100 achieved the shortest time and lowest energy consumption; Graphcore exhibited the highest throughput in both pre-training and inference, along with the highest inference energy efficiency. We discuss our observations and findings while exploring the associated tradeoffs. Farah Ferdaus, Xingfu Wu, Valerie Taylor 0001, Zhiling Lan, Sanjif Shanmugavelu, Venkatram Vishwanath, Michael E. Papka |
CCGrid | 2 |
| 2025 | ytopt: Autotuning Scientific Applications for Energy Efficiency at Large ScalesabstractABSTRACT As we enter the exascale computing era, efficiently utilizing power and optimizing the performance of scientific applications under power and energy constraints has become critical and challenging. We propose a low‐overhead autotuning framework to autotune performance and energy for various hybrid MPI/OpenMP scientific applications at large scales and to explore the tradeoffs between application runtime and power/energy for energy efficient application execution, then use this framework to autotune four ECP proxy applications—XSBench, AMG, SWFFT, and SW4lite. Our approach uses Bayesian optimization with a Random Forest surrogate model to effectively search parameter spaces with up to 6 million different configurations on two large‐scale HPC production systems, Theta at Argonne National Laboratory and Summit at Oak Ridge National Laboratory. The experimental results show that our autotuning framework at large scales has low overhead and achieves good scalability. Using the proposed autotuning framework to identify the best configurations, we achieve up to 91.59% performance improvement, up to 21.2% energy savings, and up to 37.84% EDP (energy delay product) improvement on up to 4096 nodes. Xingfu Wu, Prasanna Balaprakash, Michael Kruse, Jaehoon Koo, Brice Videau, Paul D. Hovland, Valerie Taylor 0001, Brad Geltz, Siddhartha Jana, Mary W. Hall |
Concurr. Comput. Pract. Exp. | 1 |
| 2023 | Transfer-learning-based Autotuning using Gaussian CopulaabstractAs diverse high-performance computing (HPC) systems are built, many opportunities arise for applications to solve larger problems than ever before. Given the significantly increased complexity of these HPC systems and application tuning, empirical performance tuning, such as autotuning, has emerged as a promising approach in recent years. Despite its effectiveness, autotuning is often a computationally expensive approach. Transfer learning (TL)-based autotuning seeks to address this issue by leveraging the data from prior tuning. Current TL methods for autotuning spend significant time modeling the relationship between parameter configurations and performance, which is ineffective for few-shot (that is, few empirical evaluations) tuning on new tasks. We introduce the first generative TL-based autotuning approach based on the Gaussian copula (GC) to model the high-performing regions of the search space from prior data and then generate high-performing configurations for new tasks. This allows a sampling-based approach that maximizes few-shot performance and provides the first probabilistic estimation of the few-shot budget for effective TL-based autotuning. We compare our generative TL approach with state-of-the-art autotuning techniques on several benchmarks. We find that the GC is capable of achieving 64.37% of peak few-shot performance in its first evaluation. Furthermore, the GC model can determine a few-shot transfer budget that yields up to 33.39× speedup, a dramatic improvement over the 20.58× speedup using prior techniques. Thomas Randall, Jaehoon Koo, Brice Videau, Michael Kruse, Xingfu Wu, Paul D. Hovland, Mary W. Hall, Rong Ge 0002, Prasanna Balaprakash |
ICS | 5 |
| 2023 | Utilizing ensemble learning for performance and power modeling and improvement of parallel cancer deep learning CANDLE benchmarksabstractAbstract Machine learning (ML) continues to grow in importance across nearly all domains in modeling to learn from data. Often a tradeoff exists between a model's ability to minimize bias and variance. In this article, we utilize ensemble learning to combine linear, nonlinear, and tree‐/rule‐based ML methods to cope with the bias‐variance tradeoff and result in more accurate models. We use the datasets collected for two parallel cancer deep learning CANDLE benchmarks, NT3 and P1B2, to build performance and power models based on hardware performance counters using single‐object and multiple‐objects ensemble learning to identify the most important counters for improvement on the Cray XC40 Theta at Argonne National Laboratory. Based on the insights from these models, we improve the performance and energy of P1B2 and NT3 by optimizing the deep learning environments TensorFlow, Keras, Horovod, and Python under the huge page size of 8 MB. Experimental results show that ensemble learning not only produces more accurate models but also provides more robust performance counter ranking. We achieve up to 61.15% performance improvement and up to 62.58% energy saving for P1B2 and up to 55.81% performance improvement and up to 52.60% energy saving for NT3 on up to 24,576 cores. Xingfu Wu, Valerie Taylor 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2023 | Performance and power modeling and prediction using MuMMI and 10 machine learning methodsabstractSummary Energy‐efficient scientific applications require insight into how high performance computing system features impact the applications' power and performance. This insight can result from the development of performance and power models. In this article, we use the modeling and prediction tool MuMMI (Multiple Metrics Modeling Infrastructure) and 10 machine learning methods to model and predict performance and power consumption and compare their prediction error rates. We use an algorithm‐based fault‐tolerant linear algebra code and a multilevel checkpointing fault‐tolerant heat distribution code to conduct our modeling and prediction study on the Cray XC40 Theta and IBM BG/Q Mira at Argonne National Laboratory and the Intel Haswell cluster Shepard at Sandia National Laboratories. Our experimental results show that the prediction error rates in performance and power using MuMMI are less than 10% for most cases. By utilizing the models for runtime, node power, CPU power, and memory power, we identify the most significant performance counters for potential application optimizations, and we predict theoretical outcomes of the optimizations. Based on two collected datasets, we analyze and compare the prediction accuracy in performance and power consumption using MuMMI and 10 machine learning methods. Xingfu Wu, Valerie Taylor 0001, Zhiling Lan |
Concurr. Comput. Pract. Exp. | 1 |
| 2022 | Autotuning PolyBench benchmarks with LLVM Clang/Polly loop optimization pragmas using Bayesian optimizationabstractAbstract We develop a ytopt autotuning framework that leverages Bayesian optimization to explore the parameter space search and compare four different supervised learning methods within Bayesian optimization and evaluate their effectiveness. We select six of the most complex PolyBench benchmarks and apply the newly developed LLVM Clang/Polly loop optimization pragmas to the benchmarks to optimize them. We then use the autotuning framework to optimize the pragma parameters to improve their performance. The experimental results show that our autotuning approach outperforms the other compiling methods to provide the smallest execution time for the benchmarks syr2k, 3mm, heat‐3d, lu, and covariance with two large datasets in 200 code evaluations for effectively searching the parameter spaces with up to 170,368 different configurations. We find that the Floyd–Warshall benchmark did not benefit from autotuning. To cope with this issue, we provide some compiler option solutions to improve the performance. Then we present loop autotuning without a user's knowledge using a simple mctree autotuning framework to further improve the performance of the Floyd–Warshall benchmark. We also extend the ytopt autotuning framework to tune a deep learning application. Xingfu Wu, Michael Kruse, Prasanna Balaprakash, Hal Finkel, Paul D. Hovland, Valerie Taylor 0001, Mary W. Hall |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | A Dynamic Power Capping Library for HPC ApplicationsabstractAs HPC systems increase in scale and capability, the cost of supplying power to these systems grows significantly. This introduces the urgent need for energy efficient computing through the management of power consumption. The PowerStack initiative defines a holistic power management framework at three levels, i.e., at the cluster level, the job level, and the node level [1] –[3]. It is expected that a cluster will be given a system-wide power budget (aka an allocated power budget), which can be intelligently managed and distributed at the job and node levels. Our work provides a method for intelligently managing node-level power during execution. Zhiling Lan, Xingfu Wu, Valerie Taylor 0001 |
CLUSTER | 3 |
| 2020 | Toward an End-to-End Auto-tuning Framework in HPC PowerStackabstractEfficiently utilizing procured power and optimizing performance of scientific applications under power and energy constraints are challenging. The HPC PowerStack defines a software stack to manage power and energy of high-performance computing systems and standardizes the interfaces between different components of the stack. This survey paper presents the findings of a working group focused on the end-to-end tuning of the PowerStack. First, we provide a background on the PowerStack layer-specific tuning efforts in terms of their high-level objectives, the constraints and optimization goals, layer-specific telemetry, and control parameters, and we list the existing software solutions that address those challenges. Second, we propose the PowerStack end-to-end auto-tuning framework, identify the opportunities in co-tuning different layers in the PowerStack, and present specific use cases and solutions. Third, we discuss the research opportunities and challenges for collective auto-tuning of two or more management layers (or domains) in the PowerStack. This paper takes the first steps in identifying and aggregating the important R&D challenges in streamlining the optimization efforts across the layers of the PowerStack. Xingfu Wu, Aniruddha Marathe, Siddhartha Jana, Ondrej Vysocky, Jophin John, Andrea Bartolini, Lubomir Riha, Michael Gerndt, Valerie Taylor 0001, Sridutt Bhalachandra |
CLUSTER | 1 |
| 2019 | Performance, Energy, and Scalability Analysis and Improvement of Parallel Cancer Deep Learning CANDLE BenchmarksabstractTraining scientific deep learning models requires the significant compute power of high-performance computing systems. In this paper, we analyze the performance characteristics of the benchmarks from the exploratory research project CANDLE (Cancer Distributed Learning Environment) with a focus on the hyperparameters epochs, batch sizes, and learning rates. We present the parallel methodology that uses the distributed deep learning framework Horovod to parallelize the CANDLE benchmarks. We then use scaling strategies for both epochs and batch size with linear learning rate scaling to investigate how they impact the execution time and accuracy as well as the power, energy, and scalability of the parallel CANDLE benchmarks under conditions of strong scaling and weak scaling on the IBM Power9 heterogeneous system Summit at Oak Ridge National Laboratory and the Cray XC40 Theta at Argonne National Laboratory. This study provides insights into how to set the proper numbers of epochs, batch sizes, and compute resources for these benchmarks to preserve the high accuracy and to reduce the execution time of the benchmarks. We identify the data-loading performance bottleneck and then improve the performance and energy for better scalability. Results with the modified benchmarks on Summit indicate up to 78.25% in performance improvement and up to 78% in energy saving under strong scaling on up to 384 GPUs, and up to 79.5% in performance improvement and up to 77.11% in energy saving under weak scaling on up to 3,072 GPUs. On Theta, we achieve up to 45.22% performance improvement and up to 41.78% in energy saving under strong scaling on up to 384 nodes. Moreover, the modification dramatically reduces the broadcast overhead. Xingfu Wu, Valerie Taylor 0001, Justin M. Wozniak, Rick L. Stevens, Thomas S. Brettin, Fangfang Xia |
ICPP | 1 |
| 2017 | An Energy Efficient Demand-Response Model for High Performance Computing SystemsabstractDemand response refers to reducing energy consumption of participating systems in response to transient surge in power demand or other emergency events. Demand response is particularly important for maintaining power grid transmission stability, as well as achieving overall energy saving. High Performance Computing (HPC) systems can be considered as ideal participants for demand-response programs, due to their massive energy demand. However, the potential loss of performance must be weighed against the possible gain in power system stability and energy reduction. In this paper, we explore the opportunity of demand response on HPC systems by proposing a new HPC job scheduling and resource provisioning model. More specifically, the proposed model applies power-bound energy-conservation job scheduling during the critical demand-response events, while maintaining the traditional performance-optimized job scheduling during the normal period. We expect such a model can attract willing participation of the HPC systems in the demand response programs, as it can improve both power stability and energy saving without significantly compromising application performance. We implement the proposed method in a simulator and compare it with the traditional scheduling approach. Using trace-driven simulation, we demonstrate that the HPC demand response is a viable approach toward power stability and energy savings with only marginal increase in the jobs' execution time. Kishwar Ahmed, Jason Liu 0001, Xingfu Wu |
MASCOTS | 3 |
| 2013 | MuMMI: Multiple Metrics Modeling InfrastructureabstractThe MuMMI (Multiple Metrics Modeling Infrastructure) project is an infrastructure that facilitates systematic measurement, modeling, and prediction of performance, power consumption and performance-power tradeoffs for parallel systems. In this paper, we present the MuMMI framework, which consists of an Instrument or, Databases and Analyzer. The MuMMI instrument or provides for automatic performance and power data collection and storage with low overhead. The MuMMI Databases store performance, power and energy consumption and hardware performance counters' data. The MuMMI Analyzer entails performance and power modeling and performance-power tradeoff and optimizations. As part of the MuMMI project, we mainly focus on discussing the design and development of a MuMMI Instrument or to provide automatic performance and power data collection and storage with low overhead on multicore systems in detail, then utilize the MuMMI Instrument or to collect performance and power data for a hybrid MPI/OpenMP earthquake application to discuss application performance-power trade-off and optimizations. Our experimental results show that we reduce up to 8.5% the application execution time and lower up to 18.35% the energy consumption by applying Dynamic Voltage and Frequency Scaling (DVFS), Dynamic Concurrency Throttling (DCT) and loop optimizations. Xingfu Wu, Charles W. Lively, Valerie Taylor 0001, Hung-Ching Chang, Chun-Yi Su, Kirk W. Cameron, Shirley Moore, Daniel Terpstra, Vincent M. Weaver |
SNPD | 1 |
| 2013 | Performance Characteristics of Hybrid MPI/OpenMP Scientific Applications on a Large-Scale Multithreaded BlueGene/Q SupercomputerabstractIn this paper, we investigate the performance characteristics of five hybrid MPI/OpenMP scientific applications (two NAS Parallel benchmarks Multi-Zone SP-MZ and BT-MZ, an earthquake simulation PEQdyna, an aerospace application PMLB and a 3D particle-in-cell application GTC) on a large-scale multithreaded Blue Gene/Q supercomputer at Argonne National laboratory, and quantify the performance gap resulting from using different number of threads per node. We use performance tools and MPI profile and trace libraries available on the supercomputer to analyze and compare the performance of these hybrid scientific applications with increasing the number OpenMP threads per node, and find that increasing the number of threads to some extent saturates or worsens performance of these hybrid applications. For the strong-scaling hybrid scientific applications such as SP-MZ, BT-MZ, PEQdyna and PLMB, using 32 threads per node results in much better application efficiency than using 64 threads per node, and as increasing the number of threads per node, the FPU (Floating Point Unit) percentage decreases, and the MPI percentage (except PMLB) and IPC (Instructions per cycle) per core (except BT-MZ) increase. For the weak-scaling hybrid scientific application such as GTC, the performance trend (relative speedup) is very similar with increasing number of threads per node no matter how many nodes (32, 128, 512) are used. Xingfu Wu, Valerie Taylor 0001 |
SNPD | 1 |
| 2013 | Performance modeling of hybrid MPI/OpenMP scientific applications on large-scale multicore supercomputers
Xingfu Wu, Valerie Taylor 0001 |
J. Comput. Syst. Sci. | 1 |
| 2012 | Performance Characteristics of Hybrid MPI/OpenMP Implementations of NAS Parallel Benchmarks SP and BT on Large-Scale Multicore ClustersabstractThe NAS Parallel Benchmarks (NPB) are well-known applications with fixed algorithms for evaluating parallel systems and tools. Multicore clusters provide a natural programming paradigm for hybrid programs, whereby OpenMP can be used with the data sharing with the multicores that comprise a node, and MPI can be used with the communication between nodes. In this paper, we use Scalar Pentadiagonal (SP) and Block Tridiagonal (BT) benchmarks of MPI NPB 3.3 as a basis for a comparative approach to implement hybrid MPI/OpenMP versions of SP and BT. In particular, we can compare the performance of the hybrid SP and BT with the MPI counterparts on large-scale multicore clusters, Intrepid (BlueGene/P) at Argonne National Laboratory and Jaguar (Cray XT4/5) at Oak Ridge National Laboratory. Our performance results indicate that the hybrid SP outperforms the MPI SP by up to 20.76%, and the hybrid BT outperforms the MPI BT by up to 8.58% on up to 10 000 cores on Intrepid and Jaguar. We also use performance tools and MPI trace libraries available on these clusters to further investigate the performance characteristics of the hybrid SP and BT. Xingfu Wu, Valerie Taylor 0001 |
Comput. J. | 1 |
| 2009 | Performance projection of HPC applications using SPEC CFP2006 benchmarksabstractPerformance projections of high performance computing (HPC) applications onto various hardware platforms are important for hardware vendors and HPC users. The projections aid hardware vendors in the design of future systems, enable them to compare the application performance across different existing and future systems, and help HPC users with system procurement and application refinements. In this paper, we present a method for projecting the node level performance of HPC applications using published data of industry standard benchmarks, the SPEC CFP2006, and hardware performance counter data from one base machine. In particular, we project performance of eight HPC applications onto four systems, utilizing processors from different vendors, using data from one base machine, the IBM p575. The projected performance of the eight applications was within 7.2% average difference with respect to measured runtimes for IBM POWER6 systems and standard deviation of 5.3%. For two Intel based systems with different micro-architecture and instruction set architecture (ISA) than the base machine, the average projection difference to measured runtimes was 10.5% with standard deviation of 8.2%. Sameh Sharkawi, Don DeSota, Raj Panda, Rajeev Indukuru, Stephen Stevens, Valerie Taylor 0001, Xingfu Wu |
IPDPS | 7 |
| 2008 | Performance Analysis of Parallel Visualization Applications and Scientific Applications on an Optical GridabstractOne major challenge for grid environments is how to efficiently utilize geographically distributed resources given large communication latency introduced by wide area networks interconnecting different grid sites. In this paper, we use optical networks to connect four clusters from three different sites: Texas A&M University, University of Illinois at Chicago, and University of California at San Diego to form an optical grid as an OptIPuter testbed, and execute parallel visualization applications and scientific applications to analyze the performance of these applications on the optical grid. Our experiments indicate that on-demand scheduling for resource allocation on different grid sites ensures to run MPI programs on the grid successfully. We explore the dedicated optical grid for different purpose of usage with on-demand scheduling to address how the optical grid impacts the performance of parallel programs. To avoid the large across-site communication latency, the ideal applications for the optical grid are embarrassingly parallel. Xingfu Wu, Valerie Taylor 0001 |
CW | 1 |
| 2006 | Performance Analysis, Modeling and Prediction of a Parallel Multiblock Lattice Boltzmann Application Using Prophesy SystemabstractRecently, the Lattice Boltzmann method is widely used in simulating fluid flows. In this paper, we present the performance analysis, modeling and prediction of a parallel multiblock Lattice Boltzmann application on up to 512 processors on three SMP clusters: two IBM SP systems at San Diego Supercomputing Center (DataStar-p655 and p690) and one IBM SP system at the DOE National Energy Research Scientific Computing Center (Seaborg) using the Prophesy system. By characterizing the performance of the Lattice Boltzmann application as the problem size and the number of processors increase, we can identify and eliminate performance bottlenecks, and predict the application performance. The experimental results indicate that the application with large problem sizes scales well across these three clusters, and performance models using the coupling method are accurate with less than 4.8% average relative prediction error. Xingfu Wu, Valerie Taylor 0001, Shane Garrick, Dazhi Yu, Jacques Richard |
CLUSTER | 1 |
| 2004 | Isocoupling: Reusing Kernel Coupling Values to Predict the Performance of Parallel ApplicationsabstractSummary form only given. Kernel coupling quantifies the interaction between adjacent and chains of kernels in an application. A kernel can be a loop, procedure or file. In our previous work, we used the kernel coupling values to identify how to combine the execution times of the individual kernels that compose the application to predict the execution time of the full application. The results of this previous work using the NAS Parallel Benchmark SP demonstrated that the use of coupling values resulted in very good predictions with average errors in the range of only 1.18% in contrast to simply summing the execution times of the kernels that resulted in average errors in the range of 20.54%. The major concern with the coupling values is the fact that values are needed for each different problem size, number of processors and machine. We explore the ability to reuse coupling values. In particular, we explore the reuse in terms of the three dimensional space consisting of the following axes: number of processors, problem size and system architecture. The experimental results indicate that when considering parallel systems, with increasing number of processors and problem sizes, we found clear transitions with the coupling values resulting in the ability to reuse values. Further, reusing coupling values is feasible on classes of systems such as clusters, distributed shared memory and other distributed memory systems. Xingfu Wu, Jonathan Geisler, Rick L. Stevens |
IPDPS | 1 |
| 2002 | Using Kernel Couplings to Predict Parallel Application PerformanceabstractPerformance models provide significant insight into the performance relationships between an application and the system used for execution. The major obstacle to developing performance models is the lack of knowledge about the performance relationships between the different functions that compose an application. This paper addresses the issue by using a coupling parameter, which quantifies the interaction between kernels, to develop performance predictions. The results, using three NAS parallel application benchmarks, indicate that the predictions using the coupling parameter were greatly improved over a traditional technique of summing the execution times of the individual kernels in an application. In one case the coupling predictor had less than 1% relative error in contrast the summation methodology that had over 20% relative error. Further, as the problem size and number of processors scale, the coupling values go through a finite number of major value changes that is dependent on the memory subsystem of the processor architecture. Valerie Taylor 0001, Xingfu Wu, Jonathan Geisler, Rick L. Stevens |
HPDC | 2 |
| 2000 | Prophesy: An Infrastructure for Analyzing and Modeling the Performance of Parallel and Distributed ApplicationsabstractEfficient execution of applications requires insight into how the system features impact the performance of the application. For distributed systems, the task of gaining this insight is complicated by the complexity of the system features. This insight generally results from significant experimental analysis and possibly the development of performance models. This paper presents the Prophesy project, an infrastructure that aids in gaining this needed insight based upon experience. The core component of Prophesy is a relational database that allows for the recording of performance data, system features and application details. Xingfu Wu, Valerie Taylor 0001, Jonathan Geisler, Zhiling Lan, Rick L. Stevens, Mark Hereld, Ivan R. Judson |
HPDC | 1 |
| 1998 | Performance models for scalable cluster computing
Xingfu Wu, Wei Li 0022 |
J. Syst. Archit. | 1 |
| 1997 | An Approach to Scalability of Parallel Matrix Multiplication Algorithms
Xingfu Wu |
COCOON | 1 |