Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

David K. Lowenthal

dblp:l/DavidKLowenthal · DBLP profile ↗
← Back
50ranked-venue papers
3as first author
2since 2021 · last 2023
0000-0002-4235-1541ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 41 · 3 first-author · 1 since 2021Computer networks · 5Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
20 papers
Cloud and datacenter computing · 30% Energy-efficient computing · 28% High-performance computing · 18%
Computer networks
3 papers
Datacenter networks · 75% Routing and switching · 9% Cellular and mobile networks · 9%

Topics — the 30 heaviest of 60, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Datacenter networks › data center network topology
fat-tree
0.512021
Jigsaw: A High-Utilization, Interference-Free Job Scheduler for Fat-Tree Clusters · HPDC 2021
Cloud and datacenter computing
job scheduling
0.512021
Jigsaw: A High-Utilization, Interference-Free Job Scheduler for Fat-Tree Clusters · HPDC 2021
Cloud and datacenter computing
resource management
0.522016
Exploiting Redundancy and Application Scalability for Cost-Effective, Time-Constrained Execution of HPC Applications on Amazon EC2 · IEEE Trans. Parallel Distributed Syst. 2016
Practical Resource Management in Power-Constrained, High Performance Computing · HPDC 2015
High-performance computing › distributed computing infrastructure
cloud HPC
0.422016
Exploiting Redundancy and Application Scalability for Cost-Effective, Time-Constrained Execution of HPC Applications on Amazon EC2 · IEEE Trans. Parallel Distributed Syst. 2016
Exploiting redundancy for cost-effective, time-constrained execution of HPC applications on amazon EC2 · HPDC 2014
Energy-efficient computing
power-constrained computing
0.422015
Finding the limits of power-constrained application performance · SC 2015
Practical Resource Management in Power-Constrained, High Performance Computing · HPDC 2015
Energy-efficient computing
power management
0.332015
Practical Resource Management in Power-Constrained, High Performance Computing · HPDC 2015
Bounding energy consumption in large-scale MPI programs · SC 2007
Using multiple energy gears in MPI programs on a power-scalable cluster · PPoPP 2005
High-performance computing
performance optimization at scale
0.322015
Finding the limits of power-constrained application performance · SC 2015
Bounding energy consumption in large-scale MPI programs · SC 2007
Cloud and datacenter computing › job scheduling › cloud scheduling
spot instance scheduling
0.212016
Exploiting Redundancy and Application Scalability for Cost-Effective, Time-Constrained Execution of HPC Applications on Amazon EC2 · IEEE Trans. Parallel Distributed Syst. 2016
Energy-efficient computing › power management
power budgeting
0.212015
Analyzing and mitigating the impact of manufacturing variability in power-constrained supercomputing · SC 2015
Cloud and datacenter computing › resource management
cloud resource management
0.212014
Exploiting redundancy for cost-effective, time-constrained execution of HPC applications on amazon EC2 · HPDC 2014
Hardware reliability and fault tolerance
redundancy
0.212014
Exploiting redundancy for cost-effective, time-constrained execution of HPC applications on amazon EC2 · HPDC 2014
Cloud and datacenter computing › utility computing › cloud pricing
spot market
0.212014
Exploiting redundancy for cost-effective, time-constrained execution of HPC applications on amazon EC2 · HPDC 2014
Energy-efficient computing › power management
dynamic voltage and frequency scaling
0.232006
MPI and communication - Adaptive, transparent frequency and voltage scaling of communication phases in MPI programs · SC 2006
Just In Time Dynamic Voltage Scaling: Exploiting Inter-Node Slack to Save Energy in MPI Programs · SC 2005
Using multiple energy gears in MPI programs on a power-scalable cluster · PPoPP 2005
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.112021
Jigsaw: A High-Utilization, Interference-Free Job Scheduler for Fat-Tree Clusters · HPDC 2021
Routing and switching
adaptive routing
0.112018
Mitigating inter-job interference using adaptive flow-aware routing · SC 2018
Cellular and mobile networks › interference management
interference mitigation
0.112018
Mitigating inter-job interference using adaptive flow-aware routing · SC 2018
Parallel and multicore computing
data distribution
0.122005
The MHETA Execution Model for Heterogeneous Clusters · SC 2005
Accurate data redistribution cost estimation in software distributed shared memory systems · PPoPP 2001
Energy-efficient computing › voltage scaling
dynamic voltage scaling
0.112007
Bounding energy consumption in large-scale MPI programs · SC 2007
Energy-efficient computing
energy-aware scheduling
0.112007
Bounding energy consumption in large-scale MPI programs · SC 2007
Energy-efficient computing › power-performance tradeoff
energy-delay tradeoff
0.112007
Analyzing the Energy-Time Trade-Off in High-Performance Computing Applications · IEEE Trans. Parallel Distributed Syst. 2007
High-performance computing
performance optimization
0.112007
Analyzing the Energy-Time Trade-Off in High-Performance Computing Applications · IEEE Trans. Parallel Distributed Syst. 2007
High-performance computing › supercomputing
exascale computing
0.112015
Practical Resource Management in Power-Constrained, High Performance Computing · HPDC 2015
Parallel and multicore computing › parallel programming models › hybrid programming models
hybrid MPI/OpenMP
0.112015
Finding the limits of power-constrained application performance · SC 2015
Parallel and multicore computing
parallel programming models
0.112015
Finding the limits of power-constrained application performance · SC 2015
Hardware reliability and fault tolerance
process variation
0.112015
Analyzing and mitigating the impact of manufacturing variability in power-constrained supercomputing · SC 2015
Internet of things and sensor networks › wireless sensor network
energy-efficient communication
0.112006
Client-Centered, Energy-Efficient Wireless Communication on IEEE 802.11b Networks · IEEE Trans. Mob. Comput. 2006
Program analysis › error detection
array bounds checking
0.112006
Implicit array bounds checking on 64-bit architectures · ACM Trans. Archit. Code Optim. 2006
Operating systems › resource management › memory management
virtual memory
0.112006
Implicit array bounds checking on 64-bit architectures · ACM Trans. Archit. Code Optim. 2006
Energy-efficient computing
energy-constrained computing
0.112006
Minimizing execution time in MPI programs on an energy-constrained, power-scalable cluster · PPoPP 2006
Parallel and multicore computing › parallel programming models › message passing
MPI runtime
0.112006
MPI and communication - Adaptive, transparent frequency and voltage scaling of communication phases in MPI programs · SC 2006

Methods — techniques the papers use, named apart from their topics

checkpointing · 0.4flow-aware routing · 0.3performance modeling · 0.3adaptive scheduling · 0.2variation-aware power budgeting · 0.2upper bound analysis · 0.2power provisioning · 0.2dynamic voltage scaling · 0.2redundancy · 0.2comparative study · 0.2linear programming · 0.1traffic shaping · 0.1connection tracking · 0.1compiler and OS infrastructure · 0.1
YearPublicationVenuePosition
2023 Evaluating the Potential of Coscheduling on High-Performance Computing Systems
Jason Hall, Arjun Lathi, David K. Lowenthal, Tapasya Patki
JSSPP3
2021 Jigsaw: A High-Utilization, Interference-Free Job Scheduler for Fat-Tree Clusters
abstract
Jobs on HPC clusters can suffer significant performance degradation due to inter-job network interference. Approaches to mitigating this interference primarily focus on reactive routing schemes. A better approach---in that it completely eliminates inter-job interference---is to implement scheduling policies that proactively enforce network isolation for every job. However, existing schedulers that allocate isolated partitions lead to lowered system utilization, which creates a barrier to adoption.
Staci A. Smith, David K. Lowenthal
HPDC2
2020 The Case of Performance Variability on Dragonfly-based Systems
abstract
Performance of a parallel code running on a large supercomputer can vary significantly from one run to another even when the executable and its input parameters are left unchanged. Such variability can occur due to perturbation of the computation and/or communication in the code. In this paper, we investigate the case of performance variability arising due to network effects on supercomputers that use a dragonfly topology - specifically, Cray XC systems equipped with the Aries interconnect. We perform post-mortem analysis of network hardware counters, profiling output, job queue logs, and placement information, all gathered from periodic representative application runs. We investigate the causes of performance variability using deviation prediction and recursive feature elimination. Additionally, using time-stepped performance data of individual applications, we train machine learning models that can forecast the execution time of future time steps.
Abhinav Bhatele, Jayaraman J. Thiagarajan, Taylor L. Groves, Rushil Anirudh, Staci A. Smith, Brandon Cook 0001, David K. Lowenthal
IPDPS7
2019 Mitigating Inter-Job Interference via Process-Level Quality-of-Service
abstract
Jobs on most high-performance computing (HPC) systems share the network with other concurrently executing jobs. This sharing creates contention that can severely degrade performance. We investigate the use of Quality of Service (QoS) mechanisms to reduce the negative impacts of network contention. Our results show that careful use of QoS reduces the impact of contention for specific jobs, resulting in up to a 27% performance improvement. In some cases the impact of contention is completely eliminated. These improvements are achieved with limited negative impact to other jobs; any job that experiences performance loss typically degrades less than 5%, often much less. Our approach can help ensure that HPC machines maintain high throughput as per-node compute power continues to increase faster than network bandwidth.
Lee Savoie, David K. Lowenthal, Bronis R. de Supinski, Kathryn Mohror
CLUSTER2
2018 Mitigating inter-job interference using adaptive flow-aware routing
Staci A. Smith, Clara E. Cromey, David K. Lowenthal, Jens Domke, Jayaraman J. Thiagarajan, Abhinav Bhatele
SC3
2016 I/O Aware Power Shifting
abstract
Power limits on future high-performance computing (HPC) systems will constrain applications. However, HPC applications do not consume constant power over their lifetimes. Thus, applications assigned a fixed power bound may be forced to slow down during high-power computation phases, but may not consume their full power allocation during low-power I/O phases. This paper explores algorithms that leverage application semantics -- phase frequency, duration and power needs -- to shift unused power from applications in I/O phases to applications in computation phases, thus improving system-wide performance. We design novel techniques that include explicit staggering of applications to improve power shifting. Compared to executing without power shifting, our algorithms can improve average performance by up to 8% or improve performance of a single, high-priority application by up to 32%.
Lee Savoie, David K. Lowenthal, Bronis R. de Supinski, Tanzima Z. Islam, Kathryn Mohror, Barry Rountree, Martin Schulz 0001
IPDPS2
2016 Exploiting Redundancy and Application Scalability for Cost-Effective, Time-Constrained Execution of HPC Applications on Amazon EC2
abstract
The use of clouds to execute high-performance computing (HPC) applications has greatly increased recently. Clouds provide several potential advantages over traditional supercomputers and in-house clusters. The most popular cloud is currently Amazon EC2, which provides fixed-cost and variable-cost, auction-based options. The auction market trades lower cost for potential interruptions that necessitate checkpointing; if the market price exceeds the bid price, a node is taken away from the user without warning. We explore techniques to maximize performance per dollar given a time constraint within which an application must complete. Specifically, we design and implement multiple techniques to reduce expected cost by exploiting redundancy in the EC2 auction market. We then design an adaptive algorithm that selects a scheduling algorithm and determines the bid price. We show that our adaptive algorithm executes programs up to seven times cheaper than using the on-demand market and up to 44 percent cheaper than the best non-redundant, auction-market algorithm. We extend our adaptive algorithm to incorporate application scalability characteristics for further cost savings. We show that the adaptive algorithm informed with scalability characteristics of applications achieves up to 56 percent cost savings compared to the expected cost for the base adaptive algorithm run at a fixed, user-defined scale.
Aniruddha Marathe, Rachel Harris, David K. Lowenthal, Bronis R. de Supinski, Barry Rountree, Martin Schulz 0001
IEEE Trans. Parallel Distributed Syst.3
2015 Practical Resource Management in Power-Constrained, High Performance Computing
abstract
Power management is one of the key research challenges on the path to exascale. Supercomputers today are designed to be worst-case power provisioned, leading to two main problems --- limited application performance and under-utilization of procured power.
Tapasya Patki, David K. Lowenthal, Anjana Sasidharan, Matthias Maiterth, Barry Rountree, Martin Schulz 0001, Bronis R. de Supinski
HPDC2
2015 Finding the limits of power-constrained application performance
abstract
As we approach exascale systems, power is turning from an optimization goal to a critical operating constraint. With power bounds imposed by both stakeholders and the limitations of existing infrastructure, we need to develop new techniques that work with limited power to extract maximum performance. In this paper, we explore this area and provide an approach to find the theoretical upper bound of computational performance on a per-application basis in hybrid MPI + OpenMP applications.
Peter E. Bailey, Aniruddha Marathe, David K. Lowenthal, Barry Rountree, Martin Schulz 0001
SC3
2015 Analyzing and mitigating the impact of manufacturing variability in power-constrained supercomputing
abstract
A key challenge in next-generation supercomputing is to effectively schedule limited power resources. Modern processors suffer from increasingly large power variations due to the chip manufacturing process. These variations lead to power inhomogeneity in current systems and manifest into performance inhomogeneity in power constrained environments, drastically limiting supercomputing performance. We present a first-of-its-kind study on manufacturing variability on four production HPC systems spanning four microarchitectures, analyze its impact on HPC applications, and propose a novel variation-aware power budgeting scheme to maximize effective application performance. Our low-cost and scalable budgeting algorithm strives to achieve performance homogeneity under a power constraint by deriving application-specific, module-level power allocations. Experimental results using a 1,920 socket system show up to 5.4X speedup, with an average speedup of 1.8X across all benchmarks when compared to a variation-unaware power allocation scheme.
Yuichi Inadomi, Tapasya Patki, Koji Inoue, Mutsumi Aoyagi, Barry Rountree, Martin Schulz 0001, David K. Lowenthal, Yasutaka Wada, Keiichiro Fukazawa, Masatsugu Ueda, Masaaki Kondo, Ikuo Miyoshi
SC7
2014 Exploiting redundancy for cost-effective, time-constrained execution of HPC applications on amazon EC2
abstract
The use of clouds to execute high-performance computing (HPC) applications has greatly increased recently. Clouds provide several potential advantages over traditional supercomputers and in-house clusters. The most popular cloud is currently Amazon EC2, which provides a fixed-cost option (called on-demand) and a variable-cost, auction-based option (called the spot market). The spot market trades lower cost for potential interruptions that necessitate checkpointing; if the market price exceeds the bid price, a node is taken away from the user without warning.
Aniruddha Marathe, Rachel Harris, David K. Lowenthal, Bronis R. de Supinski, Barry Rountree, Martin Schulz 0001
HPDC3
2014 Adaptive Configuration Selection for Power-Constrained Heterogeneous Systems
abstract
As power becomes an increasingly important design factor in high-end supercomputers, future systems will likely operate with power limitations significantly below their peak power specifications. These limitations will be enforced through a combination of software and hardware power policies, which will filter down from the system level to individual nodes. Hardware is already moving in this direction by providing power-capping interfaces to the user. The power/performance trade-off at the node level is critical in maximizing the performance of power-constrained cluster systems, but is also complex because of the many interacting architectural features and accelerators that comprise the hardware configuration of a node. The key to solving this challenge is an accurate power/performance model that will aid in selecting the right configuration from a large set of available configurations. In this paper, we present a novel approach to generate such a model offline using kernel clustering and multivariate linear regression. Our model requires only two iterations to select a configuration, which provides a significant advantage over exhaustive search-based strategies. We apply our model to predict power and performance for different applications using arbitrary configurations, and show that our model, when used with hardware frequency-limiting, selects configurations with significantly higher performance at a given power limit than those chosen by frequency-limiting alone. When applied to a set of 36 computational kernels from a range of applications, our model accurately predicts power and performance, it maintains 91% of optimal performance while meeting power constraints 88% of the time. When the model violates a power constraint, it exceeds the constraint by only 6% in the average case, while simultaneously achieving 54% more performance than an oracle.
Peter E. Bailey, David K. Lowenthal, Vignesh Ravi, Barry Rountree, Martin Schulz 0001, Bronis R. de Supinski
ICPP2
2013 A comparative study of high-performance computing on the cloud
Aniruddha Marathe, Rachel Harris, David K. Lowenthal, Bronis R. de Supinski, Barry Rountree, Martin Schulz 0001, Xin Yuan 0001
HPDC3
2013 Exploring hardware overprovisioning in power-constrained, high performance computing
abstract
Most recent research in power-aware supercomputing has focused on making individual nodes more efficient and measuring the results in terms of flops per watt. While this work is vital in order to reach exascale computing at 20 megawatts, there has been a dearth of work that explores efficiency at the whole system level. Traditional approaches in supercomputer design use worst-case power provisioning: the total power allocated to the system is determined by the maximum power draw possible per node. In a world where power is plentiful and nodes are scarce, this solution is optimal. However, as power becomes the limiting factor in supercomputer design, worst-case provisioning becomes a drag on performance.
Tapasya Patki, David K. Lowenthal, Barry Rountree, Martin Schulz 0001, Bronis R. de Supinski
ICS2
2013 Parallelizing heavyweight debugging tools with mpiecho
Barry Rountree, Todd Gamblin, Bronis R. de Supinski, Martin Schulz 0001, David K. Lowenthal, Guy Cobb, Henry M. Tufo
Parallel Comput.5
2012 Comet: Decentralized Complex Event Detection in Mobile Delay Tolerant Networks
abstract
Increased commodity use of mobile devices has the potential to enable mission-critical monitoring applications. However, these mobile-enabled monitoring applications have to often work in environments where a delay-tolerant network (DTN) is the only feasible communication paradigm. Detection of complex (composite) events is fundamental to monitoring applications. However, the existing plan-based CED techniques are mostly centralized, and hence are inherently unscalable for DTNs. In this paper, we create Comet â" a decentralized plan-based, efficient and scalable CED for DTNs. Comet shares the task of detecting complex events (CEs) among multiple nodes, with each node detecting a part of the CE by aggregating two or more primitive events or sub-CEs. Comet uses a unique h-function to construct cost and delay efficient CED trees. As finding an optimal CED plan requires exponential-time, Comet finds near-optimal detection plans for individual CEs through a novel multi-level push-pull conversion algorithm. Performance results show that Comet reduces cost by up to 89% compared to pushing all primitive events and over 60% compared to a two-level exhaustive search algorithm.
Jianxia Chen, Lakshmish Ramaswamy, David K. Lowenthal, Shivkumar Kalyanaraman
MDM3
2011 Adaptive, transparent CPU scaling algorithms leveraging inter-node MPI communication regions
Min Yeol Lim, Vincent W. Freeh, David K. Lowenthal
Parallel Comput.3
2010 CAEVA: A customizable and adaptive event aggregation framework for collaborative broker overlays
abstract
The publish-subscribe (pub-sub) paradigm is maturing and integrating into community-oriented collaborative applications. Because of this, pub-sub systems are faced with an event stream that may potentially contain large numbers of redundant and partial messages. Most pub-sub systems view partial and
Jianxia Chen, Lakshmish Ramaswamy, David K. Lowenthal, Shivkumar Kalyanaraman
CollaborateCom3
2010 Using focused regression for accurate time-constrained scaling of scientific applications
abstract
Many large-scale clusters now have hundreds of thousands of processors, and processor counts will be over one million within a few years. Computational scientists must scale their applications to exploit these new clusters. Time-constrained scaling, which is often used, tries to hold total execution time constant while increasing the problem size along with the processor count. However, complex interactions between parameters, the processor count, and execution time complicate determining the input parameters that achieve this goal. In this paper we develop a novel gray-box, focused regression-based approach that assists the computational scientist with maintaining constant run time on increasing processor counts. Combining application-level information from a small set of training runs, our approach allows prediction of the input parameters that result in similar per-processor execution time at larger scales. Our experimental validation across seven applications showed that median prediction errors are less than 13%.
Bradley J. Barnes, Jeonifer Garren, David K. Lowenthal, Jaxk Reeves, Bronis R. de Supinski, Martin Schulz 0001, Barry Rountree
IPDPS3
2009 Adagio: making DVS practical for complex HPC applications
abstract
Power and energy are first-order design constraints in high performance computing. Current research using dynamic voltage scaling (DVS) relies on trading increased execution time for energy savings, which is unacceptable for most high performance computing applications. We present Adagio, a novel runtime system that makes DVS practical for complex, real-world scientific applications by incurring only negligible delay while achieving significant energy savings. Adagio improves and extends previous state-of-the-art algorithms by combining the lessons learned from static energy-reducing CPU scheduling with a novel runtime mechanism for slack prediction. We present results using Adagio for two real-world programs, UMT2K and ParaDiS, along with the NAS Parallel Benchmark suite. While requiring no modification to the application source code, Adagio provides total system energy savings of 8% and 20% for UMT2K and ParaDiS, respectively, with less than 1% increase in execution time.
Barry Rountree, David K. Lowenthal, Bronis R. de Supinski, Martin Schulz 0001, Vincent W. Freeh, Tyler K. Bletsch
ICS2
2008 A regression-based approach to scalability prediction
abstract
Many applied scientific domains are increasingly relying on large-scale parallel computation. Consequently, many large clusters now have thousands of processors. However, the ideal number of processors to use for these scientific applications varies with both the input variables and the machine under consideration, and predicting this processor count is rarely straightforward. Accurate prediction mechanisms would provide many benefits, including improving cluster efficiency and identifying system configuration or hardware issues that impede performance.
Bradley J. Barnes, Barry Rountree, David K. Lowenthal, Jaxk Reeves, Bronis R. de Supinski, Martin Schulz 0001
ICS3
2008 Just-in-time dynamic voltage scaling: Exploiting inter-node slack to save energy in MPI programs
Vincent W. Freeh, Nandini Kappiah, David K. Lowenthal, Tyler K. Bletsch
J. Parallel Distributed Comput.3
2007 Bounding energy consumption in large-scale MPI programs
abstract
Power is now a first-order design constraint in large-scale parallel computing. Used carefully, dynamic voltage scaling can execute parts of a program at a slower CPU speed to achieve energy savings with a relatively small (possibly zero) time delay. However, the problem of when to change frequencies in order to optimize energy savings is NP-complete, which has led to many heuristic energy-saving algorithms. To determine how closely these algorithms approach optimal savings, we developed a system that determines a bound on the energy savings for an application. Our system uses a linear programming solver that takes as inputs the application communication trace and the cluster power characteristics and then outputs a schedule that realizes this bound. We apply our system to three scientific programs, two of which exhibit load imbalance---particle simulation and UMT2K. Results from our bounding technique show particle simulation is more amenable to energy savings than UMT2K.
Barry Rountree, David K. Lowenthal, Shelby H. Funk, Vincent W. Freeh, Bronis R. de Supinski, Martin Schulz 0001
SC2
2007 Analyzing the Energy-Time Trade-Off in High-Performance Computing Applications
abstract
Although users of high-performance computing are most interested in raw performance both energy and power consumption has become critical concerns. One approach to lowering energy and power is to use high-performance cluster nodes that have several power-performance states so that the energy-time trade-off can be dynamically adjusted. This paper analyzes the energy-time trade-off of a wide range of applications-serial and parallel-on a power-scalable cluster. We use a cluster of frequency and voltage-scalable AMD-64 nodes, each equipped with a power meter. We study the effects of memory and communication bottlenecks via direct measurement of time and energy. We also investigate metrics that can, at runtime, predict when each type of bottleneck occurs. Our results show that, for programs that have a memory or communication bottleneck, a power-scalable cluster can save significant energy with only a small time penalty. Furthermore, we find that, for some programs, it is possible to both consume less energy and execute in less time by increasing the number of nodes while reducing the frequency-voltage setting of each node
Vincent W. Freeh, David K. Lowenthal, Nandini Kappiah, Robert Springer, Barry Rountree, Mark E. Femal
IEEE Trans. Parallel Distributed Syst.2
2006 A Parallel, Out-of-Core Algorithm for RNA Secondary Structure Prediction
abstract
RNA pseudoknot prediction is an algorithm for RNA sequence search and alignment. An important building block towards pseudoknot prediction is RNA secondary structure prediction. The difficulty of extending the secondary structure prediction algorithm to a parallel program is (1) it has complicated data dependences, and (2) it has a large data set that typically cannot fit completely in main memory. In this paper, we propose a new out-of-core, distributed-memory algorithm for RNA secondary structure prediction. Its novelty lies in its redundant file scheme, I/O-reducing in-core buffer mechanism, and dynamic load balancing algorithm. Experimental results obtained on 16 Sun UltraSPARC Illi nodes provide evidence that our approach achieves good speedup. Furthermore, we found that counterintuitively, the size of the in-memory buffer is critical to efficiency of the parallel program
Wenduo Zhou, David K. Lowenthal
ICPP2
2006 STAR-MPI: self tuned adaptive routines for MPI collective operations
abstract
Message Passing Interface (MPI) collective communication routines are widely used in parallel applications. In order for a collective communication routine to achieve high performance for different applications on different platforms, it must be adaptable to both the system architecture and the application workload. Current MPI implementations do not support such software adaptability and are not able to achieve high performance on many platforms. In this paper, we present STAR-MPI (Self Tuned Adaptive Routines for MPI collective operations), a set of MPI collective communication routines that are capable of adapting to system architecture and application workload. For each operation, STAR-MPI maintains a set of communication algorithms that can potentially be efficient at different situations. As an application executes, a STAR-MPI routine applies the Automatic Empirical Optimization of Software (AEOS) technique at run time to dynamically select the best performing algorithm for the application on the platform. We describe the techniques used in STAR-MPI, analyze STAR-MPI overheads, and evaluate the performance of STAR-MPI with applications and benchmarks. The results of our study indicate that STAR-MPI is robust and efficient. It is able to and efficient algorithms with reasonable overheads, and it out-performs traditional MPI implementations to a large degree in many cases.
Ahmad Faraj, Xin Yuan 0001, David K. Lowenthal
ICS3
2006 Minimizing execution time in MPI programs on an energy-constrained, power-scalable cluster
abstract
Recently, the high-performance computing community has realized that power is a performance-limiting factor. One reason for this is that supercomputing centers have limited power capacity and machines are starting to hit that limit. In addition, the cost of energy has become increasingly significant, and the heat produced by higher-energy components tends to reduce their reliability. One way to reduce power (and therefore energy) requirements is to use high-performance cluster nodes that are frequency- and voltage-scalable (e.g., AMD-64 processors).The problem we address in this paper is: given a target program, a power-scalable cluster, and an upper limit for energy consumption, choose a schedule (number of nodes and CPU frequency) that simultaneously (1) satisfies an external upper limit for energy consumption and (2) minimizes execution time. There are too many schedules for an exhaustive search. Therefore, we find a schedule through a novel combination of performance modeling, performance prediction, and program execution. Using our technique, we are able to find a near-optimal schedule for all of our benchmarks in just a handful of partial program executions.
Robert Springer, David K. Lowenthal, Barry Rountree, Vincent W. Freeh
PPoPP2
2006 MPI and communication - Adaptive, transparent frequency and voltage scaling of communication phases in MPI programs
abstract
Although users of high-performance computing are most interested in raw performance, both energy and power consumption have become critical concerns. Some microprocessors allow frequency and voltage scaling, which enables a system to reduce CPU performance and power when the CPU is not on the critical path. When properly directed, such dynamic frequency and voltage scaling can produce significant energy savings with little performance penalty.This paper presents an MPI runtime system that dynamically reduces CPU performance during communication phases in MPI programs. It dynamically identifies such phases and, without profiling or training, selects the CPU frequency in order to minimize energy-delay product. All analysis and subsequent frequency and voltage scaling is within MPI and so is entirely transparent to the application. This means that the large number of existing MPI programs, as well as new ones being developed, can use our system without modification. Results show that the average reduction in energy-delay product over the NAS benchmark suite is 10%---the average energy reduction is 12% while the average execution time increase is only 2.1%.
Min Yeol Lim, Vincent W. Freeh, David K. Lowenthal
SC3
2006 Dyn-MPI: Supporting MPI on medium-scale, non-dedicated clusters
D. Brent Weatherly, David K. Lowenthal, Mario Nakazawa, Franklin Lowenthal
J. Parallel Distributed Comput.2
2006 Implicit array bounds checking on 64-bit architectures
abstract
Several programming languages guarantee that array subscripts are checked to ensure they are within the bounds of the array. While this guarantee improves the correctness and security of array-based code, it adds overhead to array references. This has been an obstacle to using higher-level languages, such as Java, for high-performance parallel computing, where the language specification requires that all array accesses must be checked to ensure they are within bounds. This is because, in practice, array-bounds checking in scientific applications may increase execution time by more than a factor of 2. Previous research has explored optimizations to statically eliminate bounds checks, but the dynamic nature of many scientific codes makes this difficult or impossible. Our approach is, instead, to create a compiler and operating system infrastructure that does not generate explicit bounds checks. It instead places arrays inside of Index Confinement Regions (ICRs), which are large, isolated, mostly unmapped virtual memory regions. Any array reference outside of its bounds will cause a protection violation; this provides implicit bounds checking. Our results show that when applying this infrastructure to high-performance computing programs written in Java, the overhead of bounds checking relative to a program with no bounds checks is reduced from an average of 63% to an average of 9%.
Chris Bentley, Scott A. Watterson, David K. Lowenthal, Barry Rountree
ACM Trans. Archit. Code Optim.3
2006 Client-Centered, Energy-Efficient Wireless Communication on IEEE 802.11b Networks
abstract
In mobile devices, the wireless network interface card (WNIC) consumes a significant portion of overall system energy. One way to reduce energy consumed by a device is to transition its WNIC to a lower-power sleep mode when data is not being received or transmitted. In this paper, we investigate client-centered techniques for energy efficient communication, using IEEE 802.11b, within the network layer. The basic idea is to conserve energy by keeping the WNIC in high-power mode only when necessary. We track each connection, which allows us to determine inactive intervals during which to transition the WNIC to sleep mode. Whenever necessary, we also shape the traffic from the client side to maximize sleep intervals—convincing the server to send data in bursts. This trades lower WNIC energy consumption for an increase in transmission time. Our techniques are compatible with standard TCP and do not rely on any assistance from the server or network infrastructure. Results show that during Web browsing, our client-centered technique saved 21 percent energy compared to PSM and incurred less than a 1 percent increase in transmission time compared to regular TCP. For a large file download, our scheme saved 27 percent energy on average with a transmission time increase of only 20 percent.
Haijin Yan, Scott A. Watterson, David K. Lowenthal, Kang Li 0001, Rupa Krishnan, Larry L. Peterson
IEEE Trans. Mob. Comput.3
2005 ACE: an active, client-directed method for reducing energy during web browsing
abstract
In mobile devices, the wireless network interface card (WNIC) consumes a significant portion of overall system energy. One way to reduce energy consumed by a device is to transition its WNIC to a lower-power sleep mode when data is not being received or transmitted.This paper develops ACE, an active, client-directed technique to improve energy efficiency during web browsing. ACE actively retrieves buffered packets from an access point based on predictions made through client-side connection tracking. The key novel implementation technique used in ACE is connection rescheduling, which results is a better energy/time tradeoff for interactive applications such as web browsing. We demonstrate the effectiveness of ACE through actual experiments to real Internet servers.
Haijin Yan, David K. Lowenthal, Kang Li 0001
NOSSDAV2
2005 Using multiple energy gears in MPI programs on a power-scalable cluster
abstract
Recently, system architects have built low-power, high-performance clusters, such as Green Destiny. The idea behind these clusters is to improve the energy efficiency of nodes. However, these clusters save power at the expense of performance. Our approach is instead to use high-performance cluster nodes that are frequency- and voltage-scalable; energy can than be saved by scaling down the CPU. Our prior work has examined the costs and benefits of executing an entire application at a single reduced frequency.This paper presents a framework for executing a single application in several frequency-voltage settings. The basic idea is to first divide programs into phases and then execute a series of experiments, with each phase assigned a prescribed frequency. During each experiment, we measure energy consumption and time and then use a heuristic to choose the assignment of frequency to phase for the next experiment.Our results show that significant energy can be saved without an undue performance penalty; particularly, our heuristic finds assignments of frequency to phase that is superior to any fixed-frequency solution. Specifically, this paper shows that more than half of the NAS benchmarks exhibit a better energy-time tradeoff using multiple gears than using a single gear. For example, IS using multiple gears uses 9% less energy and executes in 1% less time than the closest single-gear solution. Compared to no frequency scaling, multiple gear IS uses 16% less energy while executing only 1% longer.
Vincent W. Freeh, David K. Lowenthal
PPoPP2
2005 Just In Time Dynamic Voltage Scaling: Exploiting Inter-Node Slack to Save Energy in MPI Programs
abstract
Recently, improving the energy efficiency of HPC machines has become important. As a result, interest in using powerscalable clusters, where frequency and voltage can be dynamically modified, has increased. On power-scalable clusters, one opportunity for saving energy with little or no loss of performance exists when the computational load is not perfectly balanced. This situation occurs frequently, as balancing load between nodes is one of the long standing problems in parallel and distributed computing. In this paper we present a system called Jitter, which reduces the frequency on nodes that are assigned less computation and therefore have slack time. This saves energy on these nodes, and the goal of Jitter is to attempt to ensure that they arrive "just in time" so that they avoid increasing overall execution time. For example, in Aztec, from the ASCI Purple suite, our algorithm uses 8% less energy while increasing execution time by only 2.6%.
Nandini Kappiah, Vincent W. Freeh, David K. Lowenthal
SC3
2005 The MHETA Execution Model for Heterogeneous Clusters
abstract
The availability of inexpensive "off the shelf" machines increases the likelihood that parallel programs run on heterogeneous clusters of machines. These programs are increasingly likely to be out of core, meaning that portions of their datasets must be stored on disk during program execution. This results in significant, per-iteration, I/O cost. This paper describes an execution model, called MHETA, which is the key component to finding an effective data distribution on heterogeneous clusters. MHETA takes into account computation, communication, and I/O costs of iterative scientific applications. MHETA uses automatically extracted information from a single iteration to predict the execution time of the remaining iterations. Results show that MHETA predicts with on average 98% accuracy the execution time of several scientific benchmarks (with and without prefetching) and one full-scale scientific program that utilize pipelined and other communication. MHETA is thus an effective tool when searching for the most effective distribution on a heterogeneous cluster.
Mario Nakazawa, David K. Lowenthal, Wenduo Zhou
SC2
2005 Towards cooperation fairness in mobile ad hoc networks
abstract
For the sustainable operation of ad hoc networks, incentive mechanisms are required to encourage cooperation. More importantly, we must enforce available bandwidth fairness among nodes. Bandwidth sharing is fair if a node's available bandwidth is proportional to its forwarding contribution. In this paper, we achieve fair bandwidth sharing through a new packet scheduling algorithm, called cooperative queueing, on each node. In cooperative queueing, which is analogous to fair queueing, packet scheduling is based on a new abstraction that we call cooperation coefficient (CC). The CC quantifies how much a given node contributes to and consumes from the ad-hoc network; the larger the CC, the more bandwidth a node can obtain. We exploit the widely used dynamic source routing information to obtain the CC. We evaluate the effectiveness of cooperative queueing with different parameters and network configurations. We demonstrate that our algorithm is able to encourage cooperation and ensure fair sharing of bandwidth between nodes. We show that cooperative queueing is simple and has little overhead.
Haijin Yan, David K. Lowenthal
WCNC2
2005 An MPI prototype for compiled communication on Ethernet switched clusters
Amit Karwande, Xin Yuan 0001, David K. Lowenthal
J. Parallel Distributed Comput.3
2004 Dynamic, Power-Aware Scheduling for Mobile Clients Using a Transparent Proxy
abstract
Mobile computers consume significant amounts of energy when receiving large files. The wireless network interface card (WNIC) is the primary source of this energy consumption. One way to reduce the energy consumed is to transmit the packets to clients in a predictable fashion. Specifically, the packets can be sent in bursts to clients, who can then switch to a lower power sleep state between bursts. This technique is especially effective when the bandwidth of a stream is small. This work investigates techniques for saving energy in a multiple-client scenario, where clients may be receiving either UDP or TCP data. Energy is saved by using a transparent proxy that is invisible to both clients and servers. The proxy implementation maintains separate connections to the client and server so that a large increase in transmission time is avoided. The proxy also buffers data and dynamically generates a global transmission schedule that includes all active clients. Results show that energy savings within 10-15% of optimal are common, with little packet loss.
Michael Gundlach, Sarah Doster, Haijin Yan, David K. Lowenthal, Scott A. Watterson, Surendar Chandra
ICPP4
2004 Implicit java array bounds checking on 64-bit architecture
abstract
Interest in using Java for high-performance parallel computing has increased in recent years. One obstacle that has inhibited Java from widespread acceptance in the scientific community is the language requirement that all array accesses must be checked to ensure they are within bounds. In practice, array bounds checking in scientific applications may increase execution time by more than a factor of 2. Previous research has explored optimizations to statically eliminate bounds checks, but the dynamic nature of many scientific codes makes this difficult or impossible.Our approach is instead to create a new Java implementation that does not generate explicit bounds checks. It instead places arrays inside of Index Confinement Regions (ICRs), which are large, isolated, mostly unmapped virtual memory regions. Any array reference outside of its bounds will cause a protection violation; this provides implicit bounds checking. Our results show that our new Java implementation reduces the overhead of bounds checking from an average of 63% to an average of 9% on our benchmarks.
Chris Bentley, Scott A. Watterson, David K. Lowenthal, Barry Rountree
ICS3
2004 Client-centered energy and delay analysis for TCP downloads
abstract
In mobile devices, the wireless network interface card (WNIC) consumes a significant portion of overall system energy. One way to reduce energy consumed by a mobile device is to transition its WNIC to a lower-power sleep mode when data is not being received or transmitted. This paper investigates client-centered techniques for trading download time for energy savings during TCP downloads, in an attempt to reduce the energy' delay product. Effectively saving WNIC energy during a TCP download is difficult because TCP streams tend to be smooth, leaving little potential sleep time. The basic idea behind our technique is that the client increases the amount of time that can be spent in sleep mode by shaping the traffic. In particular, the client convinces the server to send data in predictable bursts, trading lower WNIC energy cost for increased transmission time. Our technique does not rely on any assistance from the server, a proxy, or IEEE 802.11b power-saving mode. Results show that in Internet experiments our scheme outperforms baseline TCP by 64% in the best case, with an average improvement of 19%.
Haijin Yan, Rupa Krishnan, Scott A. Watterson, David K. Lowenthal, Kang Li 0001, Larry L. Peterson
IWQoS4
2004 Client-centered energy savings for concurrent HTTP connections
abstract
In mobile devices, the wireless network interface card (WNIC) consumes a significant portion of overall system energy. One way to reduce energy consumed by a WNIC is to transition it to a lower-power sleep mode when data is not being received or transmitted.This paper investigates client-centered techniques for saving energy during web browsing. The basic idea is that the client predicts when packets will arrive, keeping the WNIC in high-power mode only when necessary. This is challenging because web browsing generally results in concurrent HTTP connections. To handle this, we maintain the state of each open connection on the client and then transition the WNIC to sleep mode when no connection is receiving data. Our technique is compatible with standard TCP and does not rely on any assistance from the server, a proxy, or IEEE 802.11b power-saving mode (PSM). Our technique combines the performance of regular TCP with nearly all the energy-saving of PSM during web downloads, and we save more energy than PSM during client think times. Results show that over an entire web browsing session (downloads and think times), our scheme saves up to 21% energy compared to PSM and incurs less than a 1% increase in transmission time compared to regular TCP.
Haijin Yan, Rupa Krishnan, Scott A. Watterson, David K. Lowenthal
NOSSDAV4
2003 CC-MPI: a compiled communication capable MPI prototype for ethernet switched clusters
abstract
No abstract available.
Amit Karwande, Xin Yuan 0001, David K. Lowenthal
PPoPP3
2003 CC-MPI: a compiled communication capable MPI prototype for ethernet switched clusters
abstract
Compiled communication has recently been proposed to improve communication performance for clusters of workstations. The idea of compiled communication is to apply more aggressive optimizations to communications whose information is known at compile time. Existing MPI libraries do not support compiled communication. In this paper, we present an MPI prototype, CC--MPI, that supports compiled communication on Ethernet switched clusters. The unique feature of CC--MPI is that it allows the user to manage network resources such as multicast groups directly and to optimize communications based on the availability of the communication information. CC--MPI optimizes one--to--all, one--to--many, all--to--all, and many--to--many collective communication routines using the compiled communication technique. We describe the techniques used in CC--MPI and report its performance. The results show that communication performance of Ethernet switched clusters can be significantly improved through compiled communication.
Amit Karwande, Xin Yuan 0001, David K. Lowenthal
PPoPP3
2003 Dyn-MPI: Supporting MPI on Non Dedicated Clusters
abstract
Distributing data is a fundamental problem in implementing efficient distributed-memory parallel programs. The problem becomes more difficult in environments where the participating nodes are not dedicated to a parallel application. We are investigating the data distribution problem in non dedicated environments in the context of explicit message-passing programs. To address this problem, we have designed and implemented an extension to MPI called Dynamic MPI (Dyn-MPI). The key component of Dyn-MPI is its run-time system, which efficiently and automatically redistributes data on the fly when there are changes in the application or the underlying environment. Dyn-MPI supports efficient memory allocation, precise measurement of system load and computation time, and node removal. Performance results show that programs that use Dyn-MPI execute efficiently in non dedicated environments, including up to almost a three-fold improvement compared to programs that do not redistribute data and a 25% improvement over standard adaptive load balancing techniques.
D. Brent Weatherly, David K. Lowenthal, Mario Nakazawa, Franklin Lowenthal
SC2
2003 A comparative analysis of fine-grain threads packages
Gregory W. Price, David K. Lowenthal
J. Parallel Distributed Comput.2
2001 Accurate data redistribution cost estimation in software distributed shared memory systems
abstract
Distributing data is one of the key problems in implementing efficient distributed-memory parallel programs. The problem becomes more difficult in programs where data redistribution between computational phases is considered. The global data distribution problem is to find the optimal distribution in multi-phase parallel programs. Solving this problem requires accurate knowledge of data redistribution cost.
Donald G. Morris, David K. Lowenthal
PPoPP2
2000 Architecture-independent parallelism for both shared- and distributed-memory machines using the Filaments package
David K. Lowenthal, Vincent W. Freeh
Parallel Comput.1
1998 Efficient support for fine-grain parallelism on shared-memory machines
abstract
A coarse-grain parallel program typically has one thread (task) per processor, whereas a fine-grain program has one thread for each independent unit of work. Although there are several advantages to fine-grain parallelism, conventional wisdom is that coarse-grain parallelism is more efficient. This paper illustrates the advantages of fine-grain parallelism and presents an efficient implementation for shared-memory machines. The approach has been implemented in a portable software package called Filaments, which employs a unique combination of techniques to achieve efficiency. The performance of the fine-grain programs discussed in this paper is always within 13% of a hand-coded coarse-grain program and is usually within 5%. © 1998 John Wiley & Sons, Ltd.
David K. Lowenthal, Vincent W. Freeh, Gregory R. Andrews
Concurr. Pract. Exp.1
1996 Using Fine-Grain Threads and Run-Time Decision Making in Parallel Computing
David K. Lowenthal, Vincent W. Freeh, Gregory R. Andrews
J. Parallel Distributed Comput.1
1994 Distributed Filaments: Efficient Fine-Grain Parallelism on a Cluster of Workstations
Vincent W. Freeh, David K. Lowenthal, Gregory R. Andrews
OSDI2