Vincent W. Freeh

dblp:90/829 · DBLP profile ↗
← Back
37ranked-venue papers
7as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 6 first-authorSecurity and privacy · 4Applied, interdisciplinary, general and emerging computing · 4Software engineering, systems software and programming languages · 3 · 1 first-authorArtificial intelligence and machine learning · 1Computer networks · 1Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
11 papers
Energy-efficient computing · 43% High-performance computing · 17% Memory systems · 15%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 30 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing › power management
dynamic voltage and frequency scaling
0.232006
MPI and communication - Adaptive, transparent frequency and voltage scaling of communication phases in MPI programs · SC 2006
Just In Time Dynamic Voltage Scaling: Exploiting Inter-Node Slack to Save Energy in MPI Programs · SC 2005
Using multiple energy gears in MPI programs on a power-scalable cluster · PPoPP 2005
Energy-efficient computing
power management
0.122007
Bounding energy consumption in large-scale MPI programs · SC 2007
Using multiple energy gears in MPI programs on a power-scalable cluster · PPoPP 2005
Storage systems
distributed storage
0.122006
Constructing collaborative desktop storage caches for large scientific datasets · ACM Trans. Storage 2006
FreeLoader: Scavenging Desktop Storage Resources for Scientific Data · SC 2005
High-performance computing
scientific data management
0.122006
FreeLoader: Scavenging Desktop Storage Resources for Scientific Data · SC 2005
Constructing collaborative desktop storage caches for large scientific datasets · ACM Trans. Storage 2006
Energy-efficient computing › voltage scaling
dynamic voltage scaling
0.112007
Bounding energy consumption in large-scale MPI programs · SC 2007
Energy-efficient computing
energy-aware scheduling
0.112007
Bounding energy consumption in large-scale MPI programs · SC 2007
Energy-efficient computing › power-performance tradeoff
energy-delay tradeoff
0.112007
Analyzing the Energy-Time Trade-Off in High-Performance Computing Applications · IEEE Trans. Parallel Distributed Syst. 2007
High-performance computing
performance optimization
0.112007
Analyzing the Energy-Time Trade-Off in High-Performance Computing Applications · IEEE Trans. Parallel Distributed Syst. 2007
High-performance computing
performance optimization at scale
0.112007
Bounding energy consumption in large-scale MPI programs · SC 2007
Memory systems › cache management › storage caching
cooperative caching
0.112006
Constructing collaborative desktop storage caches for large scientific datasets · ACM Trans. Storage 2006
Energy-efficient computing
energy-constrained computing
0.112006
Minimizing execution time in MPI programs on an energy-constrained, power-scalable cluster · PPoPP 2006
Parallel and multicore computing › parallel programming models › message passing
MPI runtime
0.112006
MPI and communication - Adaptive, transparent frequency and voltage scaling of communication phases in MPI programs · SC 2006
Performance modeling and evaluation
scheduling optimization
0.112006
Minimizing execution time in MPI programs on an energy-constrained, power-scalable cluster · PPoPP 2006
Memory systems › cache management
storage caching
0.112006
Constructing collaborative desktop storage caches for large scientific datasets · ACM Trans. Storage 2006
Parallel and multicore computing › parallel programming models › message passing
MPI application optimization
0.112005
Using multiple energy gears in MPI programs on a power-scalable cluster · PPoPP 2005
Energy-efficient computing
power-performance tradeoff
0.112005
Using multiple energy gears in MPI programs on a power-scalable cluster · PPoPP 2005
Memory systems
memory bandwidth
0.011999
Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive Architecture · SC 1999
Memory systems
processing-in-memory
0.011999
Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive Architecture · SC 1999
Memory systems › memory interface
processor-memory interface
0.011999
Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive Architecture · SC 1999
Performance modeling and evaluation
bottleneck analysis
0.012007
Analyzing the Energy-Time Trade-Off in High-Performance Computing Applications · IEEE Trans. Parallel Distributed Syst. 2007
Mathematical optimization
linear programming
0.012007
Bounding energy consumption in large-scale MPI programs · SC 2007
Parallel and multicore computing › parallel programming models › message passing
MPI applications
0.012006
MPI and communication - Adaptive, transparent frequency and voltage scaling of communication phases in MPI programs · SC 2006
Storage systems › disk array
data striping
0.012005
FreeLoader: Scavenging Desktop Storage Resources for Scientific Data · SC 2005
Parallel and multicore computing
load balancing
0.012005
Just In Time Dynamic Voltage Scaling: Exploiting Inter-Node Slack to Save Energy in MPI Programs · SC 2005
Parallel and multicore computing › load balancing
load imbalance
0.012005
Just In Time Dynamic Voltage Scaling: Exploiting Inter-Node Slack to Save Energy in MPI Programs · SC 2005
Memory systems › shared memory
distributed shared memory
0.011996
Dynamically Controlling False Sharing in Distributed Shared Memory · HPDC 1996
Memory systems › cache coherence
false sharing
0.011996
Dynamically Controlling False Sharing in Distributed Shared Memory · HPDC 1996
Parallel and multicore computing › parallelization strategies
fine-grained parallelism
0.011994
Distributed Filaments: Efficient Fine-Grain Parallelism on a Cluster of Workstations · OSDI 1994
Hardware accelerators and domain-specific architectures
irregular application acceleration
0.011999
Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive Architecture · SC 1999
Parallel and multicore computing › parallel programming models
shared-memory parallelization
0.011996
Dynamically Controlling False Sharing in Distributed Shared Memory · HPDC 1996

Methods — techniques the papers use, named apart from their topics

dynamic voltage scaling · 0.2linear programming · 0.1data striping · 0.1runtime prediction · 0.1measurement · 0.1frequency/voltage scaling · 0.1performance prediction · 0.1performance modeling · 0.1energy-delay product optimization · 0.1dynamic frequency and voltage scaling · 0.1
YearPublicationVenuePosition
2018 Micky: A Cheaper Alternative for Selecting Cloud Instances
abstract
Most cloud computing optimizers explore and improve one workload at a time. When optimizing many workloads, the single-optimizer approach can be prohibitively expensive. Accordingly, we examine "collective optimizer" that concurrently explore and improve a set of workloads significantly reducing the measurement costs. Our large-scale empirical study shows that there is often a single cloud configuration which is surprisingly near-optimal for most workloads. Consequently, we create a collective-optimizer, MICKY, that reformulates the task of finding the near-optimal cloud configuration as a multi-armed bandit problem. MICKY efficiently balances exploration (of new cloud configurations) and exploitation (of known good cloud configuration). Our experiments show that MICKY can achieve on average 8.6 times reduction in measurement cost as compared to the state-of-the-art method while finding near-optimal solutions. Hence we propose MICKY as the basis of a practical collective optimization method for finding good cloud configurations (based on various constraints such as budget and tolerance to near-optimal configurations).
Chin-Jung Hsu, Vivek Nair, Tim Menzies, Vincent W. Freeh
IEEE CLOUD4
2018 Arrow: Low-Level Augmented Bayesian Optimization for Finding the Best Cloud VM
abstract
With the advent of big data applications, which tend to have longer execution time, choosing the right cloud VM has significant performance and economic implications. For example, in our large-scale empirical study of 107 different workloads on three popular big data systems, we found that a wrong choice can lead to a 20 times slowdown or an increase in cost by 10 times. Bayesian optimization is a technique for optimizing expensive (black-box) functions. Previous work has only used instance-level information (such as core counts and memory size) which is not sufficient to represent the search space. In this work, we discover that this may lead to the fragility problem-either incurs high search cost or finds only the sub-optimal solution. The central insight of this paper is to use low-level performance information to augment the process of Bayesian Optimization. Our novel low-level augmented Bayesian Optimization is rarely worse than current practices and often performs much better (in 46 of 107 cases). Further, it significantly reduces the search cost in nearly half of our case studies. Based on this work, we conclude that it is often insufficient to use general-purpose off-the-shelf methods for configuring cloud instances without augmenting those methods with essential systems knowledge such as CPU utilization, working memory size and I/O wait time.
Chin-Jung Hsu, Vivek Nair, Vincent W. Freeh, Tim Menzies
ICDCS3
2017 Trilogy: Data placement to improve performance and robustness of cloud computing
abstract
Infrastructure as a Service, one of the most disruptive aspects of cloud computing, enables configuring a cluster for each application for each workload. When the workload changes, a cluster will be either underutilized (wasting resources) or unable to meet demand (incurring opportunity costs). Consequently, efficient cluster resizing requires proper data replication and placement. Our work reveals that coarse-grain, workload-aware replication addresses over-utilization but cannot resolve under-utilization. With fine-grain partitioning of the dataset, data replication can reduce both under- and over-utilization. In our empirical studies, compared to a näive uniform data replication a coarse-grain workload-aware replication increases throughput by 81% on a highly-skewed workload. A fine-grain scheme further reaches 166% increase. Furthermore, a surprisingly small increase in granularity is sufficient to obtain most benefits. Evaluations also show that maximizing the number of unique partitions per node increases robustness to tolerate workload deviation while minimizing this number reduces storage footprint.
Chin-Jung Hsu, Vincent W. Freeh, Flavio Villanustre
IEEE BigData2
2016 Inside-Out: Reliable Performance Prediction for Distributed Storage Systems in the Cloud
abstract
Many storage systems are undergoing a significant shift from dedicated appliance-based model to software-defined storage (SDS) because the latter is flexible, scalable and cost-effective for modern workloads. However, it is challenging to provide a reliable guarantee of end-to-end performance in SDS due to complex software stack, time-varying workload and performance interference among tenants. Therefore, modeling and monitoring the performance of storage systems is critical for ensuring reliable QoS guarantees. Existing approaches such as performance benchmarking and analytical modeling are inadequate because they are not efficient in exploring large configuration space, and cannot support elastic operations and diverse storage services in SDS. This paper presents Inside-Out, an automatic model building tool that creates accurate performance models for distributed storage services. Inside-Out is a black-box approach. It builds high-level performance models by applying machine learning techniques to low-level system performance metrics collected from individual components of the distributed SDS system. Inside-Out uses a two-level learning method that combines two machine learning models to automatically filter irrelevant features, boost prediction accuracy and yield consistent prediction. Our in-depth evaluation shows that Inside-Out is a robust solution that enables SDS to predict end-to-end performance even in challenging conditions, e.g., changes in workload, storage configuration, available cloud resources, size of the distributed storage service, and amount of interference due to multi-tenants. Our experiments show that Inside-Out can predict end-to-end performance with 91.1% accuracy on average. Its prediction accuracy is consistent across diverse storage environments.
Chin-Jung Hsu, Rajesh Krishna Panta, Moo-Ryong Ra, Vincent W. Freeh
SRDS4
2015 Dynamically Controlling Node-Level Parallelism in Hadoop
abstract
Hadoop is a widely used large scale data processing framework. Applications run in Hadoop as containers, the concurrency of which affects completion time of an application as well as system resource usage. When there are too many concurrent containers, resource bottlenecks occur and when there too few, system resources are underutilized. The default and best practice settings underutilize resources which results in longer application completion times. In this work, we develop an approach to dynamically change the parallelism for concurrent containers to suit an application. Our approach ensures efficient utilization of resources and avoids bottlenecks for all types of MapReduce applications. Our approach improves performance of MapReduce applications by as much as 28% and 60% respectively when compared to the best practice and default settings.
Kamal Kc, Vincent W. Freeh
CLOUD2
2015 Evaluation of MapReduce in a Large Cluster
abstract
MapReduce is a widely used framework that runs large scale data processing applications. However, there are very few systematic studies of MapReduce on large clusters and thus there is a lack of reference for expected behavior or issues while running applications in a large cluster. This paper describes our findings of running applications on Pivotal's Analytics Workbench, which consists of a 540-node Hadoop cluster. Our experience sheds light on how applications behave in a large-scale cluster. This paper discusses our experiences in three areas. The first describes scaling behavior of applications as the dataset size increases. The second discusses the appropriate settings for parallelism and overlap of map and reduce tasks. The third area discusses general observations. These areas have not been reported or studied previously. Our findings show that IO-intensive applications do not scale as data size increases and MapReduce applications require different amounts of parallelism and overlap to minimize completion time. Additionally, our observations also highlight the need for appropriate memory allocation for a MapReduce component and the importance of decreasing log file size.
Kamal Kc, Chin-Jung Hsu, Vincent W. Freeh
CLOUD3
2011 Mitigating code-reuse attacks with control-flow locking
abstract
Code-reuse attacks are software exploits in which an attacker directs control flow through existing code with a malicious result. One such technique, return-oriented programming, is based on "gadgets" (short pre-existing sequences of code ending in a ret instruction) being executed in arbitrary order as a result of a stack corruption exploit. Many existing codereuse defenses have relied upon a particular attribute of the attack in question (e.g., the frequency of ret instructions in a return-oriented attack), which leads to an incomplete protection, while a smaller number of efforts in protecting all exploitable control flow transfers suffer from limited deploy-ability due to high performance overhead. In this paper, we present a novel cost-effective defense technique called control flow locking, which allows for effective enforcement of control flow integrity with a small performance overhead. Specifically, instead of immediately determining whether a control flow violation happens before the control flow transfer takes place, control flow locking lazily detects the violation after the transfer. To still restrict attackers' capability, our scheme guarantees that the deviation of the normal control flow graph will only occur at most once. Further, our scheme ensures that this deviation cannot be used to craft a malicious system call, which denies any potential gains an attacker might obtain from what is permitted in the threat model. We have developed a proof-of-concept prototype in Linux and our evaluation demonstrates desirable effectiveness and competitive performance overhead with existing techniques. In several benchmarks, our scheme is able to achieve significant gains.
Tyler K. Bletsch, Xuxian Jiang, Vincent W. Freeh
ACSAC3
2011 Jump-oriented programming: a new class of code-reuse attack
abstract
Return-oriented programming is an effective code-reuse attack in which short code sequences ending in a ret instruction are found within existing binaries and executed in arbitrary order by taking control of the stack. This allows for Turing-complete behavior in the target program without the need for injecting attack code, thus significantly negating current code injection defense efforts (e.g., W⊕X). On the other hand, its inherent characteristics, such as the reliance on the stack and the consecutive execution of return-oriented gadgets, have prompted a variety of defenses to detect or prevent it from happening.
Tyler K. Bletsch, Xuxian Jiang, Vincent W. Freeh, Zhenkai Liang
AsiaCCS3
2011 On the Expressiveness of Return-into-libc Attacks
Mark Etheridge, Tyler K. Bletsch, Xuxian Jiang, Vincent W. Freeh, Peng Ning
RAID5
2011 Adaptive, transparent CPU scaling algorithms leveraging inter-node MPI communication regions
Min Yeol Lim, Vincent W. Freeh, David K. Lowenthal
Parallel Comput.2
2009 PADD: Power Aware Domain Distribution
abstract
Modern data centers usually have computing resources sized to handle expected peak demand, but average demand is generally much lower than peak. This means that the systems in the data center usually operate at very low utilization rates. Past techniques have exploited this fact to achieve significant power savings, but they generally focus on centrally managed, throughput-oriented systems that process a single fine-grained request stream. We propose a more general solution - a technique to save power by dynamically migrating virtual machines and packing them onto fewer physical machines when possible. We call our scheme power-aware domain distribution (PADD). In this paper, we report on simulation results for PADD and demonstrate that the power and performance changes from using PADD are primarily dependent on how much buffering or reserve capacity it maintains. Our adaptive buffering scheme achieves energy savings within 7% of the idealized system that has no performance penalty. Our results also show that we can achieve an energy savings up to 70% with fewer than 1% of the requests violating their service level agreements.
Min Yeol Lim, Freeman L. Rawson III, Tyler K. Bletsch, Vincent W. Freeh
ICDCS4
2009 Adagio: making DVS practical for complex HPC applications
abstract
Power and energy are first-order design constraints in high performance computing. Current research using dynamic voltage scaling (DVS) relies on trading increased execution time for energy savings, which is unacceptable for most high performance computing applications. We present Adagio, a novel runtime system that makes DVS practical for complex, real-world scientific applications by incurring only negligible delay while achieving significant energy savings. Adagio improves and extends previous state-of-the-art algorithms by combining the lessons learned from static energy-reducing CPU scheduling with a novel runtime mechanism for slack prediction. We present results using Adagio for two real-world programs, UMT2K and ParaDiS, along with the NAS Parallel Benchmark suite. While requiring no modification to the application source code, Adagio provides total system energy savings of 8% and 20% for UMT2K and ParaDiS, respectively, with less than 1% increase in execution time.
Barry Rountree, David K. Lowenthal, Bronis R. de Supinski, Martin Schulz 0001, Vincent W. Freeh, Tyler K. Bletsch
ICS5
2009 Resource-efficient computing paradigm for computational protein modeling applications
abstract
Many computational protein modeling applications using numerical methods such as Molecular Dynamics (MD), Monte Carlo (MC), or Genetic Algorithms (GA) require a large number of energy estimations of the protein molecular system. A typical energy function describing the protein energy is a combination of a number of terms characterizing various interactions within the protein molecule as well as the protein-solvent interactions. Evaluating the energy function of a relatively large protein molecule is rather computationally costly and usually occupies the major computation time in the protein simulation process. In this paper, we present a resource-efficient computing paradigm based on ldquoconsolidationrdquo to reduce the computational time of evaluating the energy function of large protein molecule. The fundamental idea of consolidation is to increase computational density to a computer in order to increase the CPU utilizations. Consolidation will be particularly efficient when the consolidated computations have heterogeneous resource demands. In computational protein modeling applications with costly energy function evaluation, we advocate the use of ldquothread consolidation,rdquo which is to spawn concurrent threads to carry out parallel energy function terms computations. Our computational results show that 7%~11% speedup in a protein loop structure prediction program on various hardware architectures where memory-intensive and computation-intensive terms coexist in the energy function. For an MD protein simulation program where computation-intensive energy function evaluations are divided and carried out by concurrent threads, we also find slight performance improvement when the thread consolidation technique is applied.
Yaohang Li, Douglas Wardell, Vincent W. Freeh
IPDPS3
2008 Just-in-time dynamic voltage scaling: Exploiting inter-node slack to save energy in MPI programs
Vincent W. Freeh, Nandini Kappiah, David K. Lowenthal, Tyler K. Bletsch
J. Parallel Distributed Comput.1
2007 Scaling and Packing on a Chip Multiprocessor
abstract
Power management is critical in server and high-performance computing environments as well as in mobile computing. Many mechanisms have been developed over recent years to support a wide a variety of power management techniques. In particular, general purpose microprocessors now support dynamically modifying the power-performance state through voltage and frequency changes. This development spawned a very important area of this research in dynamic voltage and frequency scaling (DVFS). On the other hand, in a multiprocessor environment one can perform power management by offlining and idling processors when computational demand is low in a technique called CPU packing. This paper examines the effect of combining voltage and frequency scaling and CPU packing in a multiprocessor. Furthermore, it examines DVFS on a chip multiprocessor in which multiple processor cores are placed on a single die. This paper shows that in general one should use DVFS first, then CPU packing. Furthermore, we find that the effectiveness of CPU packing is application-dependent: commercial workloads (e.g. Apache) with periods of low utilization can reduce power by as much as 19% via packing, while the improvement to HPC workloads ranges from small to negligible.
Vincent W. Freeh, Tyler K. Bletsch, Freeman L. Rawson III
IPDPS1
2007 Determining the Minimum Energy Consumption using Dynamic Voltage and Frequency Scaling
abstract
While improving raw performance is of primary interest to most users of high-performance computers, energy consumption also is a critical concern. Some microprocessors allow voltage and frequency scaling, which enables a system to reduce CPU power and performance when the CPU is not on the critical path. When properly directed, such dynamic voltage and frequency scaling can produce significant energy savings with little performance penalty. Various DVFS scaling algorithms have been proposed. However, the benefit is application-dependent. We cannot see if they achieve the energy consumption as minimum as possible. So, it is important to establish the baseline of the DVFS scheduling for any application. This paper determines minimum energy consumption in voltage and frequency scaling systems for a given time delay. We assume we have a set of fixed points where scaling can occur. A brute-force solution is intractable even for a moderately sized set (although all programs presented in this paper can be solved with the brute-force). Our algorithm efficiently chooses the exact optimal schedule satisfying the given time constraint by estimation. Besides, our time and energy estimations from the optimal schedule have reasonable accuracy with 1.48% of differences at maximum.
Min Yeol Lim, Vincent W. Freeh
IPDPS2
2007 Bounding energy consumption in large-scale MPI programs
abstract
Power is now a first-order design constraint in large-scale parallel computing. Used carefully, dynamic voltage scaling can execute parts of a program at a slower CPU speed to achieve energy savings with a relatively small (possibly zero) time delay. However, the problem of when to change frequencies in order to optimize energy savings is NP-complete, which has led to many heuristic energy-saving algorithms. To determine how closely these algorithms approach optimal savings, we developed a system that determines a bound on the energy savings for an application. Our system uses a linear programming solver that takes as inputs the application communication trace and the cluster power characteristics and then outputs a schedule that realizes this bound. We apply our system to three scientific programs, two of which exhibit load imbalance---particle simulation and UMT2K. Results from our bounding technique show particle simulation is more amenable to energy savings than UMT2K.
Barry Rountree, David K. Lowenthal, Shelby H. Funk, Vincent W. Freeh, Bronis R. de Supinski, Martin Schulz 0001
SC4
2007 Analyzing the Energy-Time Trade-Off in High-Performance Computing Applications
abstract
Although users of high-performance computing are most interested in raw performance both energy and power consumption has become critical concerns. One approach to lowering energy and power is to use high-performance cluster nodes that have several power-performance states so that the energy-time trade-off can be dynamically adjusted. This paper analyzes the energy-time trade-off of a wide range of applications-serial and parallel-on a power-scalable cluster. We use a cluster of frequency and voltage-scalable AMD-64 nodes, each equipped with a power meter. We study the effects of memory and communication bottlenecks via direct measurement of time and energy. We also investigate metrics that can, at runtime, predict when each type of bottleneck occurs. Our results show that, for programs that have a memory or communication bottleneck, a power-scalable cluster can save significant energy with only a small time penalty. Furthermore, we find that, for some programs, it is possible to both consume less energy and execute in less time by increasing the number of nodes while reducing the frequency-voltage setting of each node
Vincent W. Freeh, David K. Lowenthal, Nandini Kappiah, Robert Springer, Barry Rountree, Mark E. Femal
IEEE Trans. Parallel Distributed Syst.1
2006 Positioning Dynamic Storage Caches for Transient Data
abstract
Simulations, experiments and observatories are generating a deluge of scientific data. Even more staggering is the ever growing application demand to process and assimilate these datasets. Application users perform a range of data operations, collaborate and share data in many novel ways. The current storage landscape is struggling to keep up with these trends in scientific data processing. Application users pay the price due to over-crowded shared filesystems, or expensive storage area networks, or not enough local storage, or high-latency archival or wide-area transfers. In order to sustain and maximize I/O bandwidth relative to increasing CPU speeds, applications must take advantage of large amounts of intermediate commodity storage, However, intermediate storage presents new challenges above and beyond the traditional distributed file system paradigm: persistent scheduling, storage/CPU coallo-cation, namespace management, lifetime management, and novel application interfaces. In this paper, we describe applications that require intermediate storage management, suggest several open research problems, and illustrate two systems - Freeloader and Tactical Storage - that attack different aspects of these problems
Sudharshan S. Vazhkudai, Douglas Thain, Xiaosong Ma, Vincent W. Freeh
CLUSTER4
2006 Coupling prefix caching and collective downloads for remote dataset access
abstract
Scientific datasets are typically archived at mass storage systems or data centers close to supercomputers/instruments. End-users of these datasets, however, usually perform parts of their workflows at their local computers. In such cases, client-side caching can offer significant gains by reducing the cost of wide-area data movement.Scientific data caches, however, traditionally cache entire data-sets, which may not be necessary. In this paper, we propose a novel combination of prefix caching and collective download. Prefix caching allows the bootstrapping of dataset downloads by caching only a prefix of the dataset, while collective download facilitates efficient parallel patching of the missing suffix from an external data source. To estimate the optimal prefix size, we further present an analytical model that considers both the initial download over-head and the downloading speed. We implemented our proposed approach in the FreeLoader distributed cache prototype. Experimental results (using multiple scientific data repositories and data transfer tools, as well as a real-world scientific dataset access trace) demonstrate that prefix caching and collective download can be implemented efficiently, our model can select an appropriate prefix size, and the cache hit rate can be improved significantly without hurting the local access rate of cached datasets.
Xiaosong Ma, Vincent W. Freeh, Sudharshan S. Vazhkudai, Tyler A. Simon, Stephen L. Scott
ICS2
2006 Minimizing execution time in MPI programs on an energy-constrained, power-scalable cluster
abstract
Recently, the high-performance computing community has realized that power is a performance-limiting factor. One reason for this is that supercomputing centers have limited power capacity and machines are starting to hit that limit. In addition, the cost of energy has become increasingly significant, and the heat produced by higher-energy components tends to reduce their reliability. One way to reduce power (and therefore energy) requirements is to use high-performance cluster nodes that are frequency- and voltage-scalable (e.g., AMD-64 processors).The problem we address in this paper is: given a target program, a power-scalable cluster, and an upper limit for energy consumption, choose a schedule (number of nodes and CPU frequency) that simultaneously (1) satisfies an external upper limit for energy consumption and (2) minimizes execution time. There are too many schedules for an exhaustive search. Therefore, we find a schedule through a novel combination of performance modeling, performance prediction, and program execution. Using our technique, we are able to find a near-optimal schedule for all of our benchmarks in just a handful of partial program executions.
Robert Springer, David K. Lowenthal, Barry Rountree, Vincent W. Freeh
PPoPP4
2006 MPI and communication - Adaptive, transparent frequency and voltage scaling of communication phases in MPI programs
abstract
Although users of high-performance computing are most interested in raw performance, both energy and power consumption have become critical concerns. Some microprocessors allow frequency and voltage scaling, which enables a system to reduce CPU performance and power when the CPU is not on the critical path. When properly directed, such dynamic frequency and voltage scaling can produce significant energy savings with little performance penalty.This paper presents an MPI runtime system that dynamically reduces CPU performance during communication phases in MPI programs. It dynamically identifies such phases and, without profiling or training, selects the CPU frequency in order to minimize energy-delay product. All analysis and subsequent frequency and voltage scaling is within MPI and so is entirely transparent to the application. This means that the large number of existing MPI programs, as well as new ones being developed, can use our system without modification. Results show that the average reduction in energy-delay product over the NAS benchmark suite is 10%---the average energy reduction is 12% while the average execution time increase is only 2.1%.
Min Yeol Lim, Vincent W. Freeh, David K. Lowenthal
SC2
2006 Constructing collaborative desktop storage caches for large scientific datasets
abstract
High-end computing is suffering a data deluge from experiments, simulations, and apparatus that creates overwhelming application dataset sizes. This has led to the proliferation of high-end mass storage systems, storage area clusters, and data centers. These storage facilities offer a large range of choices in terms of capacity and access rate, as well as strong data availability and consistency support. However, for most end-users, the “last mile” in their analysis pipeline often requires data processing and visualization at local computers, typically local desktop workstations. End-user workstations---despite having more processing power than ever before---are ill-equipped to cope with such data demands due to insufficient secondary storage space and I/O rates. Meanwhile, a large portion of desktop storage is unused.We propose the FreeLoader framework, which aggregates unused desktop storage space and I/O bandwidth into a shared cache/scratch space, for hosting large, immutable datasets and exploiting data access locality. This article presents the FreeLoader architecture, component design, and performance results based on our proof-of-concept prototype. Its architecture comprises contributing benefactor nodes, steered by a management layer, providing services such as data integrity, high performance, load balancing, and impact control. Our experiments show that FreeLoader is an appealing low-cost solution to storing massive datasets by delivering higher data access rates than traditional storage facilities, namely, local or remote shared file systems, storage systems, and Internet data repositories. In particular, we present novel data striping techniques that allow FreeLoader to efficiently aggregate a workstation's network communication bandwidth and local I/O bandwidth. In addition, the performance impact on the native workload of donor machines is small and can be effectively controlled. Further, we show that security features such as data encryptions and integrity checks can be easily added as filters for interested clients. Finally, we demonstrate how legacy applications can use the FreeLoader API to store and retrieve datasets.
Sudharshan S. Vazhkudai, Xiaosong Ma, Vincent W. Freeh, Jonathan W. Strickland, Nandan Tammineedi, Tyler A. Simon, Stephen L. Scott
ACM Trans. Storage3
2005 Using multiple energy gears in MPI programs on a power-scalable cluster
abstract
Recently, system architects have built low-power, high-performance clusters, such as Green Destiny. The idea behind these clusters is to improve the energy efficiency of nodes. However, these clusters save power at the expense of performance. Our approach is instead to use high-performance cluster nodes that are frequency- and voltage-scalable; energy can than be saved by scaling down the CPU. Our prior work has examined the costs and benefits of executing an entire application at a single reduced frequency.This paper presents a framework for executing a single application in several frequency-voltage settings. The basic idea is to first divide programs into phases and then execute a series of experiments, with each phase assigned a prescribed frequency. During each experiment, we measure energy consumption and time and then use a heuristic to choose the assignment of frequency to phase for the next experiment.Our results show that significant energy can be saved without an undue performance penalty; particularly, our heuristic finds assignments of frequency to phase that is superior to any fixed-frequency solution. Specifically, this paper shows that more than half of the NAS benchmarks exhibit a better energy-time tradeoff using multiple gears than using a single gear. For example, IS using multiple gears uses 9% less energy and executes in 1% less time than the closest single-gear solution. Compared to no frequency scaling, multiple gear IS uses 16% less energy while executing only 1% longer.
Vincent W. Freeh, David K. Lowenthal
PPoPP1
2005 Just In Time Dynamic Voltage Scaling: Exploiting Inter-Node Slack to Save Energy in MPI Programs
abstract
Recently, improving the energy efficiency of HPC machines has become important. As a result, interest in using powerscalable clusters, where frequency and voltage can be dynamically modified, has increased. On power-scalable clusters, one opportunity for saving energy with little or no loss of performance exists when the computational load is not perfectly balanced. This situation occurs frequently, as balancing load between nodes is one of the long standing problems in parallel and distributed computing. In this paper we present a system called Jitter, which reduces the frequency on nodes that are assigned less computation and therefore have slack time. This saves energy on these nodes, and the goal of Jitter is to attempt to ensure that they arrive "just in time" so that they avoid increasing overall execution time. For example, in Aztec, from the ASCI Purple suite, our algorithm uses 8% less energy while increasing execution time by only 2.6%.
Nandini Kappiah, Vincent W. Freeh, David K. Lowenthal
SC2
2005 FreeLoader: Scavenging Desktop Storage Resources for Scientific Data
abstract
High-end computing is suffering a data deluge from experiments, simulations, and apparatus that creates overwhelming application dataset sizes. End-user workstations-despite more processing power than ever before-are ill-equipped to cope with such data demands due to insufficient secondary storage space and I/O rates. Meanwhile, a large portion of desktop storage is unused. We present the FreeLoader framework, which aggregates unused desktop storage space and I/O bandwidth into a shared cache/scratch space, for hosting large, immutable datasets and exploiting data access locality. Our experiments show that FreeLoader is an appealing low-cost solution to storing massive datasets, by delivering higher data access rates than traditional storage facilities. In particular, we present novel data striping techniques that allow FreeLoader to efficiently aggregate a workstation’s network communication bandwidth and local I/O bandwidth. In addition, the performance impact on the native workload of donor machines is small and can be effectively controlled.
Sudharshan S. Vazhkudai, Xiaosong Ma, Vincent W. Freeh, Jonathan W. Strickland, Nandan Tammineedi, Stephen L. Scott
SC3
2003 Web server performance in a WAN environment
abstract
This paper analyzes web server performance under simulated WAN conditions. The workload simulates many critical network characteristics, such as network delay, bandwidth limit for client connections, and small MTU sizes to dial-up clients. A novel aspect of this study is the examination of the internal behavior of the web server at the network protocol stack and the device driver. We discovered that WAN network characteristics may significantly change the behavior of the web server compared to LAN-based simulations and may make many of optimizations of server design irrelevant in this environment. Particularly, we found that the small MTU size of the dial-up or wireless user connections can increase the processing overhead several times.
Vsevolod Panteleenko, Vincent W. Freeh
ICCCN2
2002 Streaming extensibility in the Modify-on-Access file system
H. Richard Kendall, Vincent W. Freeh, Paul W. Schermerhorn, Robert J. Minerick, Peter W. Rijks
J. Syst. Softw.2
2002 The design and implementation of the exported procedure call
abstract
Abstract This paper describes the exported procedure call, a mechanism that pushes computation out of the operating system kernel and into user space. It supports a simple, secure model for system extensions. An exported procedure call incurs overhead crossing of the user‐kernel boundary, but once in user space, it has greater security and usability and is significantly simpler than an equivalent kernel operation. This paper demonstrates the capabilities of the exported procedure call by discussing two implementations. One is the Modify‐on‐access (Mona) file system and the other is the Magi device interface. Mona and Magi use the exported procedure call in order to safely execute untrusted or complex system extensions. This paper shows that in situations where raw kernel performance is not paramount, the exported procedure call is desirable. Copyright © 2001 John Wiley & Sons, Ltd.
H. Richard Kendall, Vincent W. Freeh
Softw. Pract. Exp.2
2000 Architecture-independent parallelism for both shared- and distributed-memory machines using the Filaments package
David K. Lowenthal, Vincent W. Freeh
Parallel Comput.2
1999 Microservers: a new memory semantics for massively parallel computing
abstract
Article Microservers: a new memory semantics for massively parallel computing Share on Authors: Jay B. Brockman Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, IN Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, INView Profile , Peter M. Kogge Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, IN Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, INView Profile , Thomas L. Sterling Center for Advanced Computing Research, California Institute of Technology, Pasadena, CA Center for Advanced Computing Research, California Institute of Technology, Pasadena, CAView Profile , Vincent W. Freeh Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, IN Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, INView Profile , Shannon K. Kuntz Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, IN Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, INView Profile Authors Info & Claims ICS '99: Proceedings of the 13th international conference on SupercomputingJune 1999 Pages 454–463https://doi.org/10.1145/305138.305234Online:01 May 1999Publication History 29citation567DownloadsMetricsTotal Citations29Total Downloads567Last 12 Months9Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Jay B. Brockman, Peter M. Kogge, Thomas L. Sterling, Vincent W. Freeh, Shannon K. Kuntz
International Conference on Supercomputing4
1999 Mapping Irregular Applications to DIVA, a PIM-based Data-Intensive Architecture
abstract
Processing-in-memory (PIM) chips that integrate processor logic into memory devices offer a new opportunity for bridging the growing gap between processor and memory speeds, especially for applications with high memory-bandwidth requirements.The Data-IntensiVe Architecture (DIVA) system combines PIM memories with one or more external host processors and a PIM-to-PIM interconnect.DIVA increases memory bandwidth through two mechanisms: (1) performing selected computation in memory, reducing the quantity of data transferred across the processor-memory interface; and (2) providing communication mechanisms called parcels for moving both data and computation throughout memory, further bypassing the processor-memory bus.DIVA uniquely supports acceleration of important irregular applications, including sparse-matrix and pointer-based computations.In this paper, we focus on several aspects of DIVA designed to effectively support such computations at very high performance levels: (1) the memory model and parcel definitions; (2) the PIM-to-PIM interconnect; and, (3) requirements for the processor-to-memory interface.We demonstrate the potential of PIMbased architectures in accelerating the performance of three irregular computations, sparse conjugate gradient, a natural-join database operation and an object-oriented database query.
Mary W. Hall, Peter M. Kogge, Jefferey G. Koller, Pedro C. Diniz, Jacqueline Chame, Jeffrey T. Draper, Jeff LaCoss, John J. Granacki, Jay B. Brockman, Apoorv Srivastava, William C. Athas, Vincent W. Freeh, Joonseok Park
SC12
1998 Efficient support for fine-grain parallelism on shared-memory machines
abstract
A coarse-grain parallel program typically has one thread (task) per processor, whereas a fine-grain program has one thread for each independent unit of work. Although there are several advantages to fine-grain parallelism, conventional wisdom is that coarse-grain parallelism is more efficient. This paper illustrates the advantages of fine-grain parallelism and presents an efficient implementation for shared-memory machines. The approach has been implemented in a portable software package called Filaments, which employs a unique combination of techniques to achieve efficiency. The performance of the fine-grain programs discussed in this paper is always within 13% of a hand-coded coarse-grain program and is usually within 5%. © 1998 John Wiley & Sons, Ltd.
David K. Lowenthal, Vincent W. Freeh, Gregory R. Andrews
Concurr. Pract. Exp.2
1996 Dynamically Controlling False Sharing in Distributed Shared Memory
abstract
Distributed shared memory (DSM) alleviates the need to program message passing explicitly on a distributed-memory machine. In order to reduce memory latency, a DSM replicates copies of data. This paper examines several current approaches to controlling thrashing caused by false sharing in a DSM. Then it introduces a novel memory consistency protocol, writer-owns, which detects and eliminates false sharing at run time. In iterative computations, where the data is accessed similarly every iteration, the writer-owns protocol can have tremendous benefits because the overhead of eliminating false sharing is only incurred once. Performance results show that the writer-owns protocol is competitive with and often better than existing approaches.
Vincent W. Freeh, Gregory R. Andrews
HPDC1
1996 A Comparison of Implicit and Explicit Parallel Programming
Vincent W. Freeh
J. Parallel Distributed Comput.1
1996 Using Fine-Grain Threads and Run-Time Decision Making in Parallel Computing
David K. Lowenthal, Vincent W. Freeh, Gregory R. Andrews
J. Parallel Distributed Comput.2
1994 Distributed Filaments: Efficient Fine-Grain Parallelism on a Cluster of Workstations
Vincent W. Freeh, David K. Lowenthal, Gregory R. Andrews
OSDI1