EDBT 2026 Demo / reviewers in the wild / expert
David K. Lowenthal
dblp:l/DavidKLowenthal
· DBLP profile ↗
50ranked-venue papers
3as first author
2since 2021 · last 2023
0000-0002-4235-1541ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 3 first-author · 1 since 2021Computer networks · 5Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
20 papers |
Cloud and datacenter computing · 30% Energy-efficient computing · 28% High-performance computing · 18% | |
| Computer networks
3 papers |
Datacenter networks · 75% Routing and switching · 9% Cellular and mobile networks · 9% |
Topics — the 30 heaviest of 60, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Datacenter networks › data center network topology
fat-tree |
0.5 | 1 | 2021 | Jigsaw: A High-Utilization, Interference-Free Job Scheduler for Fat-Tree Clusters · HPDC 2021 |
Cloud and datacenter computing
job scheduling |
0.5 | 1 | 2021 | Jigsaw: A High-Utilization, Interference-Free Job Scheduler for Fat-Tree Clusters · HPDC 2021 |
Cloud and datacenter computing
resource management |
0.5 | 2 | 2016 | Exploiting Redundancy and Application Scalability for Cost-Effective, Time-Constrained Execution of HPC Applications on Amazon EC2 · IEEE Trans. Parallel Distributed Syst. 2016 Practical Resource Management in Power-Constrained, High Performance Computing · HPDC 2015 |
High-performance computing › distributed computing infrastructure
cloud HPC |
0.4 | 2 | 2016 | Exploiting Redundancy and Application Scalability for Cost-Effective, Time-Constrained Execution of HPC Applications on Amazon EC2 · IEEE Trans. Parallel Distributed Syst. 2016 Exploiting redundancy for cost-effective, time-constrained execution of HPC applications on amazon EC2 · HPDC 2014 |
Energy-efficient computing
power-constrained computing |
0.4 | 2 | 2015 | Finding the limits of power-constrained application performance · SC 2015 Practical Resource Management in Power-Constrained, High Performance Computing · HPDC 2015 |
Energy-efficient computing
power management |
0.3 | 3 | 2015 | Practical Resource Management in Power-Constrained, High Performance Computing · HPDC 2015 Bounding energy consumption in large-scale MPI programs · SC 2007 Using multiple energy gears in MPI programs on a power-scalable cluster · PPoPP 2005 |
High-performance computing
performance optimization at scale |
0.3 | 2 | 2015 | Finding the limits of power-constrained application performance · SC 2015 Bounding energy consumption in large-scale MPI programs · SC 2007 |
Cloud and datacenter computing › job scheduling › cloud scheduling
spot instance scheduling |
0.2 | 1 | 2016 | Exploiting Redundancy and Application Scalability for Cost-Effective, Time-Constrained Execution of HPC Applications on Amazon EC2 · IEEE Trans. Parallel Distributed Syst. 2016 |
Energy-efficient computing › power management
power budgeting |
0.2 | 1 | 2015 | Analyzing and mitigating the impact of manufacturing variability in power-constrained supercomputing · SC 2015 |
Cloud and datacenter computing › resource management
cloud resource management |
0.2 | 1 | 2014 | Exploiting redundancy for cost-effective, time-constrained execution of HPC applications on amazon EC2 · HPDC 2014 |
Hardware reliability and fault tolerance
redundancy |
0.2 | 1 | 2014 | Exploiting redundancy for cost-effective, time-constrained execution of HPC applications on amazon EC2 · HPDC 2014 |
Cloud and datacenter computing › utility computing › cloud pricing
spot market |
0.2 | 1 | 2014 | Exploiting redundancy for cost-effective, time-constrained execution of HPC applications on amazon EC2 · HPDC 2014 |
Energy-efficient computing › power management
dynamic voltage and frequency scaling |
0.2 | 3 | 2006 | MPI and communication - Adaptive, transparent frequency and voltage scaling of communication phases in MPI programs · SC 2006 Just In Time Dynamic Voltage Scaling: Exploiting Inter-Node Slack to Save Energy in MPI Programs · SC 2005 Using multiple energy gears in MPI programs on a power-scalable cluster · PPoPP 2005 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.1 | 1 | 2021 | Jigsaw: A High-Utilization, Interference-Free Job Scheduler for Fat-Tree Clusters · HPDC 2021 |
Routing and switching
adaptive routing |
0.1 | 1 | 2018 | Mitigating inter-job interference using adaptive flow-aware routing · SC 2018 |
Cellular and mobile networks › interference management
interference mitigation |
0.1 | 1 | 2018 | Mitigating inter-job interference using adaptive flow-aware routing · SC 2018 |
Parallel and multicore computing
data distribution |
0.1 | 2 | 2005 | The MHETA Execution Model for Heterogeneous Clusters · SC 2005 Accurate data redistribution cost estimation in software distributed shared memory systems · PPoPP 2001 |
Energy-efficient computing › voltage scaling
dynamic voltage scaling |
0.1 | 1 | 2007 | Bounding energy consumption in large-scale MPI programs · SC 2007 |
Energy-efficient computing
energy-aware scheduling |
0.1 | 1 | 2007 | Bounding energy consumption in large-scale MPI programs · SC 2007 |
Energy-efficient computing › power-performance tradeoff
energy-delay tradeoff |
0.1 | 1 | 2007 | Analyzing the Energy-Time Trade-Off in High-Performance Computing Applications · IEEE Trans. Parallel Distributed Syst. 2007 |
High-performance computing
performance optimization |
0.1 | 1 | 2007 | Analyzing the Energy-Time Trade-Off in High-Performance Computing Applications · IEEE Trans. Parallel Distributed Syst. 2007 |
High-performance computing › supercomputing
exascale computing |
0.1 | 1 | 2015 | Practical Resource Management in Power-Constrained, High Performance Computing · HPDC 2015 |
Parallel and multicore computing › parallel programming models › hybrid programming models
hybrid MPI/OpenMP |
0.1 | 1 | 2015 | Finding the limits of power-constrained application performance · SC 2015 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2015 | Finding the limits of power-constrained application performance · SC 2015 |
Hardware reliability and fault tolerance
process variation |
0.1 | 1 | 2015 | Analyzing and mitigating the impact of manufacturing variability in power-constrained supercomputing · SC 2015 |
Internet of things and sensor networks › wireless sensor network
energy-efficient communication |
0.1 | 1 | 2006 | Client-Centered, Energy-Efficient Wireless Communication on IEEE 802.11b Networks · IEEE Trans. Mob. Comput. 2006 |
Program analysis › error detection
array bounds checking |
0.1 | 1 | 2006 | Implicit array bounds checking on 64-bit architectures · ACM Trans. Archit. Code Optim. 2006 |
Operating systems › resource management › memory management
virtual memory |
0.1 | 1 | 2006 | Implicit array bounds checking on 64-bit architectures · ACM Trans. Archit. Code Optim. 2006 |
Energy-efficient computing
energy-constrained computing |
0.1 | 1 | 2006 | Minimizing execution time in MPI programs on an energy-constrained, power-scalable cluster · PPoPP 2006 |
Parallel and multicore computing › parallel programming models › message passing
MPI runtime |
0.1 | 1 | 2006 | MPI and communication - Adaptive, transparent frequency and voltage scaling of communication phases in MPI programs · SC 2006 |
Methods — techniques the papers use, named apart from their topics
checkpointing · 0.4flow-aware routing · 0.3performance modeling · 0.3adaptive scheduling · 0.2variation-aware power budgeting · 0.2upper bound analysis · 0.2power provisioning · 0.2dynamic voltage scaling · 0.2redundancy · 0.2comparative study · 0.2linear programming · 0.1traffic shaping · 0.1connection tracking · 0.1compiler and OS infrastructure · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Evaluating the Potential of Coscheduling on High-Performance Computing Systems
Jason Hall, Arjun Lathi, David K. Lowenthal, Tapasya Patki |
JSSPP | 3 |
| 2021 | Jigsaw: A High-Utilization, Interference-Free Job Scheduler for Fat-Tree ClustersabstractJobs on HPC clusters can suffer significant performance degradation due to inter-job network interference. Approaches to mitigating this interference primarily focus on reactive routing schemes. A better approach---in that it completely eliminates inter-job interference---is to implement scheduling policies that proactively enforce network isolation for every job. However, existing schedulers that allocate isolated partitions lead to lowered system utilization, which creates a barrier to adoption. Staci A. Smith, David K. Lowenthal |
HPDC | 2 |
| 2020 | The Case of Performance Variability on Dragonfly-based SystemsabstractPerformance of a parallel code running on a large supercomputer can vary significantly from one run to another even when the executable and its input parameters are left unchanged. Such variability can occur due to perturbation of the computation and/or communication in the code. In this paper, we investigate the case of performance variability arising due to network effects on supercomputers that use a dragonfly topology - specifically, Cray XC systems equipped with the Aries interconnect. We perform post-mortem analysis of network hardware counters, profiling output, job queue logs, and placement information, all gathered from periodic representative application runs. We investigate the causes of performance variability using deviation prediction and recursive feature elimination. Additionally, using time-stepped performance data of individual applications, we train machine learning models that can forecast the execution time of future time steps. Abhinav Bhatele, Jayaraman J. Thiagarajan, Taylor L. Groves, Rushil Anirudh, Staci A. Smith, Brandon Cook 0001, David K. Lowenthal |
IPDPS | 7 |
| 2019 | Mitigating Inter-Job Interference via Process-Level Quality-of-ServiceabstractJobs on most high-performance computing (HPC) systems share the network with other concurrently executing jobs. This sharing creates contention that can severely degrade performance. We investigate the use of Quality of Service (QoS) mechanisms to reduce the negative impacts of network contention. Our results show that careful use of QoS reduces the impact of contention for specific jobs, resulting in up to a 27% performance improvement. In some cases the impact of contention is completely eliminated. These improvements are achieved with limited negative impact to other jobs; any job that experiences performance loss typically degrades less than 5%, often much less. Our approach can help ensure that HPC machines maintain high throughput as per-node compute power continues to increase faster than network bandwidth. Lee Savoie, David K. Lowenthal, Bronis R. de Supinski, Kathryn Mohror |
CLUSTER | 2 |
| 2018 | Mitigating inter-job interference using adaptive flow-aware routing
Staci A. Smith, Clara E. Cromey, David K. Lowenthal, Jens Domke, Jayaraman J. Thiagarajan, Abhinav Bhatele |
SC | 3 |
| 2016 | I/O Aware Power ShiftingabstractPower limits on future high-performance computing (HPC) systems will constrain applications. However, HPC applications do not consume constant power over their lifetimes. Thus, applications assigned a fixed power bound may be forced to slow down during high-power computation phases, but may not consume their full power allocation during low-power I/O phases. This paper explores algorithms that leverage application semantics -- phase frequency, duration and power needs -- to shift unused power from applications in I/O phases to applications in computation phases, thus improving system-wide performance. We design novel techniques that include explicit staggering of applications to improve power shifting. Compared to executing without power shifting, our algorithms can improve average performance by up to 8% or improve performance of a single, high-priority application by up to 32%. Lee Savoie, David K. Lowenthal, Bronis R. de Supinski, Tanzima Z. Islam, Kathryn Mohror, Barry Rountree, Martin Schulz 0001 |
IPDPS | 2 |
| 2016 | Exploiting Redundancy and Application Scalability for Cost-Effective, Time-Constrained Execution of HPC Applications on Amazon EC2abstractThe use of clouds to execute high-performance computing (HPC) applications has greatly increased recently. Clouds provide several potential advantages over traditional supercomputers and in-house clusters. The most popular cloud is currently Amazon EC2, which provides fixed-cost and variable-cost, auction-based options. The auction market trades lower cost for potential interruptions that necessitate checkpointing; if the market price exceeds the bid price, a node is taken away from the user without warning. We explore techniques to maximize performance per dollar given a time constraint within which an application must complete. Specifically, we design and implement multiple techniques to reduce expected cost by exploiting redundancy in the EC2 auction market. We then design an adaptive algorithm that selects a scheduling algorithm and determines the bid price. We show that our adaptive algorithm executes programs up to seven times cheaper than using the on-demand market and up to 44 percent cheaper than the best non-redundant, auction-market algorithm. We extend our adaptive algorithm to incorporate application scalability characteristics for further cost savings. We show that the adaptive algorithm informed with scalability characteristics of applications achieves up to 56 percent cost savings compared to the expected cost for the base adaptive algorithm run at a fixed, user-defined scale. Aniruddha Marathe, Rachel Harris, David K. Lowenthal, Bronis R. de Supinski, Barry Rountree, Martin Schulz 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | Practical Resource Management in Power-Constrained, High Performance ComputingabstractPower management is one of the key research challenges on the path to exascale. Supercomputers today are designed to be worst-case power provisioned, leading to two main problems --- limited application performance and under-utilization of procured power. Tapasya Patki, David K. Lowenthal, Anjana Sasidharan, Matthias Maiterth, Barry Rountree, Martin Schulz 0001, Bronis R. de Supinski |
HPDC | 2 |
| 2015 | Finding the limits of power-constrained application performanceabstractAs we approach exascale systems, power is turning from an optimization goal to a critical operating constraint. With power bounds imposed by both stakeholders and the limitations of existing infrastructure, we need to develop new techniques that work with limited power to extract maximum performance. In this paper, we explore this area and provide an approach to find the theoretical upper bound of computational performance on a per-application basis in hybrid MPI + OpenMP applications. Peter E. Bailey, Aniruddha Marathe, David K. Lowenthal, Barry Rountree, Martin Schulz 0001 |
SC | 3 |
| 2015 | Analyzing and mitigating the impact of manufacturing variability in power-constrained supercomputingabstractA key challenge in next-generation supercomputing is to effectively schedule limited power resources. Modern processors suffer from increasingly large power variations due to the chip manufacturing process. These variations lead to power inhomogeneity in current systems and manifest into performance inhomogeneity in power constrained environments, drastically limiting supercomputing performance. We present a first-of-its-kind study on manufacturing variability on four production HPC systems spanning four microarchitectures, analyze its impact on HPC applications, and propose a novel variation-aware power budgeting scheme to maximize effective application performance. Our low-cost and scalable budgeting algorithm strives to achieve performance homogeneity under a power constraint by deriving application-specific, module-level power allocations. Experimental results using a 1,920 socket system show up to 5.4X speedup, with an average speedup of 1.8X across all benchmarks when compared to a variation-unaware power allocation scheme. Yuichi Inadomi, Tapasya Patki, Koji Inoue, Mutsumi Aoyagi, Barry Rountree, Martin Schulz 0001, David K. Lowenthal, Yasutaka Wada, Keiichiro Fukazawa, Masatsugu Ueda, Masaaki Kondo, Ikuo Miyoshi |
SC | 7 |
| 2014 | Exploiting redundancy for cost-effective, time-constrained execution of HPC applications on amazon EC2abstractThe use of clouds to execute high-performance computing (HPC) applications has greatly increased recently. Clouds provide several potential advantages over traditional supercomputers and in-house clusters. The most popular cloud is currently Amazon EC2, which provides a fixed-cost option (called on-demand) and a variable-cost, auction-based option (called the spot market). The spot market trades lower cost for potential interruptions that necessitate checkpointing; if the market price exceeds the bid price, a node is taken away from the user without warning. Aniruddha Marathe, Rachel Harris, David K. Lowenthal, Bronis R. de Supinski, Barry Rountree, Martin Schulz 0001 |
HPDC | 3 |
| 2014 | Adaptive Configuration Selection for Power-Constrained Heterogeneous SystemsabstractAs power becomes an increasingly important design factor in high-end supercomputers, future systems will likely operate with power limitations significantly below their peak power specifications. These limitations will be enforced through a combination of software and hardware power policies, which will filter down from the system level to individual nodes. Hardware is already moving in this direction by providing power-capping interfaces to the user. The power/performance trade-off at the node level is critical in maximizing the performance of power-constrained cluster systems, but is also complex because of the many interacting architectural features and accelerators that comprise the hardware configuration of a node. The key to solving this challenge is an accurate power/performance model that will aid in selecting the right configuration from a large set of available configurations. In this paper, we present a novel approach to generate such a model offline using kernel clustering and multivariate linear regression. Our model requires only two iterations to select a configuration, which provides a significant advantage over exhaustive search-based strategies. We apply our model to predict power and performance for different applications using arbitrary configurations, and show that our model, when used with hardware frequency-limiting, selects configurations with significantly higher performance at a given power limit than those chosen by frequency-limiting alone. When applied to a set of 36 computational kernels from a range of applications, our model accurately predicts power and performance, it maintains 91% of optimal performance while meeting power constraints 88% of the time. When the model violates a power constraint, it exceeds the constraint by only 6% in the average case, while simultaneously achieving 54% more performance than an oracle. Peter E. Bailey, David K. Lowenthal, Vignesh Ravi, Barry Rountree, Martin Schulz 0001, Bronis R. de Supinski |
ICPP | 2 |
| 2013 | A comparative study of high-performance computing on the cloud
Aniruddha Marathe, Rachel Harris, David K. Lowenthal, Bronis R. de Supinski, Barry Rountree, Martin Schulz 0001, Xin Yuan 0001 |
HPDC | 3 |
| 2013 | Exploring hardware overprovisioning in power-constrained, high performance computingabstractMost recent research in power-aware supercomputing has focused on making individual nodes more efficient and measuring the results in terms of flops per watt. While this work is vital in order to reach exascale computing at 20 megawatts, there has been a dearth of work that explores efficiency at the whole system level. Traditional approaches in supercomputer design use worst-case power provisioning: the total power allocated to the system is determined by the maximum power draw possible per node. In a world where power is plentiful and nodes are scarce, this solution is optimal. However, as power becomes the limiting factor in supercomputer design, worst-case provisioning becomes a drag on performance. Tapasya Patki, David K. Lowenthal, Barry Rountree, Martin Schulz 0001, Bronis R. de Supinski |
ICS | 2 |
| 2013 | Parallelizing heavyweight debugging tools with mpiecho
Barry Rountree, Todd Gamblin, Bronis R. de Supinski, Martin Schulz 0001, David K. Lowenthal, Guy Cobb, Henry M. Tufo |
Parallel Comput. | 5 |
| 2012 | Comet: Decentralized Complex Event Detection in Mobile Delay Tolerant NetworksabstractIncreased commodity use of mobile devices has the potential to enable mission-critical monitoring applications. However, these mobile-enabled monitoring applications have to often work in environments where a delay-tolerant network (DTN) is the only feasible communication paradigm. Detection of complex (composite) events is fundamental to monitoring applications. However, the existing plan-based CED techniques are mostly centralized, and hence are inherently unscalable for DTNs. In this paper, we create Comet â" a decentralized plan-based, efficient and scalable CED for DTNs. Comet shares the task of detecting complex events (CEs) among multiple nodes, with each node detecting a part of the CE by aggregating two or more primitive events or sub-CEs. Comet uses a unique h-function to construct cost and delay efficient CED trees. As finding an optimal CED plan requires exponential-time, Comet finds near-optimal detection plans for individual CEs through a novel multi-level push-pull conversion algorithm. Performance results show that Comet reduces cost by up to 89% compared to pushing all primitive events and over 60% compared to a two-level exhaustive search algorithm. Jianxia Chen, Lakshmish Ramaswamy, David K. Lowenthal, Shivkumar Kalyanaraman |
MDM | 3 |
| 2011 | Adaptive, transparent CPU scaling algorithms leveraging inter-node MPI communication regions
Min Yeol Lim, Vincent W. Freeh, David K. Lowenthal |
Parallel Comput. | 3 |
| 2010 | CAEVA: A customizable and adaptive event aggregation framework for collaborative broker overlaysabstractThe publish-subscribe (pub-sub) paradigm is maturing and integrating into community-oriented collaborative applications. Because of this, pub-sub systems are faced with an event stream that may potentially contain large numbers of redundant and partial messages. Most pub-sub systems view partial and Jianxia Chen, Lakshmish Ramaswamy, David K. Lowenthal, Shivkumar Kalyanaraman |
CollaborateCom | 3 |
| 2010 | Using focused regression for accurate time-constrained scaling of scientific applicationsabstractMany large-scale clusters now have hundreds of thousands of processors, and processor counts will be over one million within a few years. Computational scientists must scale their applications to exploit these new clusters. Time-constrained scaling, which is often used, tries to hold total execution time constant while increasing the problem size along with the processor count. However, complex interactions between parameters, the processor count, and execution time complicate determining the input parameters that achieve this goal. In this paper we develop a novel gray-box, focused regression-based approach that assists the computational scientist with maintaining constant run time on increasing processor counts. Combining application-level information from a small set of training runs, our approach allows prediction of the input parameters that result in similar per-processor execution time at larger scales. Our experimental validation across seven applications showed that median prediction errors are less than 13%. Bradley J. Barnes, Jeonifer Garren, David K. Lowenthal, Jaxk Reeves, Bronis R. de Supinski, Martin Schulz 0001, Barry Rountree |
IPDPS | 3 |
| 2009 | Adagio: making DVS practical for complex HPC applicationsabstractPower and energy are first-order design constraints in high performance computing. Current research using dynamic voltage scaling (DVS) relies on trading increased execution time for energy savings, which is unacceptable for most high performance computing applications. We present Adagio, a novel runtime system that makes DVS practical for complex, real-world scientific applications by incurring only negligible delay while achieving significant energy savings. Adagio improves and extends previous state-of-the-art algorithms by combining the lessons learned from static energy-reducing CPU scheduling with a novel runtime mechanism for slack prediction. We present results using Adagio for two real-world programs, UMT2K and ParaDiS, along with the NAS Parallel Benchmark suite. While requiring no modification to the application source code, Adagio provides total system energy savings of 8% and 20% for UMT2K and ParaDiS, respectively, with less than 1% increase in execution time. Barry Rountree, David K. Lowenthal, Bronis R. de Supinski, Martin Schulz 0001, Vincent W. Freeh, Tyler K. Bletsch |
ICS | 2 |
| 2008 | A regression-based approach to scalability predictionabstractMany applied scientific domains are increasingly relying on large-scale parallel computation. Consequently, many large clusters now have thousands of processors. However, the ideal number of processors to use for these scientific applications varies with both the input variables and the machine under consideration, and predicting this processor count is rarely straightforward. Accurate prediction mechanisms would provide many benefits, including improving cluster efficiency and identifying system configuration or hardware issues that impede performance. Bradley J. Barnes, Barry Rountree, David K. Lowenthal, Jaxk Reeves, Bronis R. de Supinski, Martin Schulz 0001 |
ICS | 3 |
| 2008 | Just-in-time dynamic voltage scaling: Exploiting inter-node slack to save energy in MPI programs
Vincent W. Freeh, Nandini Kappiah, David K. Lowenthal, Tyler K. Bletsch |
J. Parallel Distributed Comput. | 3 |
| 2007 | Bounding energy consumption in large-scale MPI programsabstractPower is now a first-order design constraint in large-scale parallel computing. Used carefully, dynamic voltage scaling can execute parts of a program at a slower CPU speed to achieve energy savings with a relatively small (possibly zero) time delay. However, the problem of when to change frequencies in order to optimize energy savings is NP-complete, which has led to many heuristic energy-saving algorithms. To determine how closely these algorithms approach optimal savings, we developed a system that determines a bound on the energy savings for an application. Our system uses a linear programming solver that takes as inputs the application communication trace and the cluster power characteristics and then outputs a schedule that realizes this bound. We apply our system to three scientific programs, two of which exhibit load imbalance---particle simulation and UMT2K. Results from our bounding technique show particle simulation is more amenable to energy savings than UMT2K. Barry Rountree, David K. Lowenthal, Shelby H. Funk, Vincent W. Freeh, Bronis R. de Supinski, Martin Schulz 0001 |
SC | 2 |
| 2007 | Analyzing the Energy-Time Trade-Off in High-Performance Computing ApplicationsabstractAlthough users of high-performance computing are most interested in raw performance both energy and power consumption has become critical concerns. One approach to lowering energy and power is to use high-performance cluster nodes that have several power-performance states so that the energy-time trade-off can be dynamically adjusted. This paper analyzes the energy-time trade-off of a wide range of applications-serial and parallel-on a power-scalable cluster. We use a cluster of frequency and voltage-scalable AMD-64 nodes, each equipped with a power meter. We study the effects of memory and communication bottlenecks via direct measurement of time and energy. We also investigate metrics that can, at runtime, predict when each type of bottleneck occurs. Our results show that, for programs that have a memory or communication bottleneck, a power-scalable cluster can save significant energy with only a small time penalty. Furthermore, we find that, for some programs, it is possible to both consume less energy and execute in less time by increasing the number of nodes while reducing the frequency-voltage setting of each node Vincent W. Freeh, David K. Lowenthal, Nandini Kappiah, Robert Springer, Barry Rountree, Mark E. Femal |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2006 | A Parallel, Out-of-Core Algorithm for RNA Secondary Structure PredictionabstractRNA pseudoknot prediction is an algorithm for RNA sequence search and alignment. An important building block towards pseudoknot prediction is RNA secondary structure prediction. The difficulty of extending the secondary structure prediction algorithm to a parallel program is (1) it has complicated data dependences, and (2) it has a large data set that typically cannot fit completely in main memory. In this paper, we propose a new out-of-core, distributed-memory algorithm for RNA secondary structure prediction. Its novelty lies in its redundant file scheme, I/O-reducing in-core buffer mechanism, and dynamic load balancing algorithm. Experimental results obtained on 16 Sun UltraSPARC Illi nodes provide evidence that our approach achieves good speedup. Furthermore, we found that counterintuitively, the size of the in-memory buffer is critical to efficiency of the parallel program Wenduo Zhou, David K. Lowenthal |
ICPP | 2 |
| 2006 | STAR-MPI: self tuned adaptive routines for MPI collective operationsabstractMessage Passing Interface (MPI) collective communication routines are widely used in parallel applications. In order for a collective communication routine to achieve high performance for different applications on different platforms, it must be adaptable to both the system architecture and the application workload. Current MPI implementations do not support such software adaptability and are not able to achieve high performance on many platforms. In this paper, we present STAR-MPI (Self Tuned Adaptive Routines for MPI collective operations), a set of MPI collective communication routines that are capable of adapting to system architecture and application workload. For each operation, STAR-MPI maintains a set of communication algorithms that can potentially be efficient at different situations. As an application executes, a STAR-MPI routine applies the Automatic Empirical Optimization of Software (AEOS) technique at run time to dynamically select the best performing algorithm for the application on the platform. We describe the techniques used in STAR-MPI, analyze STAR-MPI overheads, and evaluate the performance of STAR-MPI with applications and benchmarks. The results of our study indicate that STAR-MPI is robust and efficient. It is able to and efficient algorithms with reasonable overheads, and it out-performs traditional MPI implementations to a large degree in many cases. Ahmad Faraj, Xin Yuan 0001, David K. Lowenthal |
ICS | 3 |
| 2006 | Minimizing execution time in MPI programs on an energy-constrained, power-scalable clusterabstractRecently, the high-performance computing community has realized that power is a performance-limiting factor. One reason for this is that supercomputing centers have limited power capacity and machines are starting to hit that limit. In addition, the cost of energy has become increasingly significant, and the heat produced by higher-energy components tends to reduce their reliability. One way to reduce power (and therefore energy) requirements is to use high-performance cluster nodes that are frequency- and voltage-scalable (e.g., AMD-64 processors).The problem we address in this paper is: given a target program, a power-scalable cluster, and an upper limit for energy consumption, choose a schedule (number of nodes and CPU frequency) that simultaneously (1) satisfies an external upper limit for energy consumption and (2) minimizes execution time. There are too many schedules for an exhaustive search. Therefore, we find a schedule through a novel combination of performance modeling, performance prediction, and program execution. Using our technique, we are able to find a near-optimal schedule for all of our benchmarks in just a handful of partial program executions. Robert Springer, David K. Lowenthal, Barry Rountree, Vincent W. Freeh |
PPoPP | 2 |
| 2006 | MPI and communication - Adaptive, transparent frequency and voltage scaling of communication phases in MPI programsabstractAlthough users of high-performance computing are most interested in raw performance, both energy and power consumption have become critical concerns. Some microprocessors allow frequency and voltage scaling, which enables a system to reduce CPU performance and power when the CPU is not on the critical path. When properly directed, such dynamic frequency and voltage scaling can produce significant energy savings with little performance penalty.This paper presents an MPI runtime system that dynamically reduces CPU performance during communication phases in MPI programs. It dynamically identifies such phases and, without profiling or training, selects the CPU frequency in order to minimize energy-delay product. All analysis and subsequent frequency and voltage scaling is within MPI and so is entirely transparent to the application. This means that the large number of existing MPI programs, as well as new ones being developed, can use our system without modification. Results show that the average reduction in energy-delay product over the NAS benchmark suite is 10%---the average energy reduction is 12% while the average execution time increase is only 2.1%. Min Yeol Lim, Vincent W. Freeh, David K. Lowenthal |
SC | 3 |
| 2006 | Dyn-MPI: Supporting MPI on medium-scale, non-dedicated clusters
D. Brent Weatherly, David K. Lowenthal, Mario Nakazawa, Franklin Lowenthal |
J. Parallel Distributed Comput. | 2 |
| 2006 | Implicit array bounds checking on 64-bit architecturesabstractSeveral programming languages guarantee that array subscripts are checked to ensure they are within the bounds of the array. While this guarantee improves the correctness and security of array-based code, it adds overhead to array references. This has been an obstacle to using higher-level languages, such as Java, for high-performance parallel computing, where the language specification requires that all array accesses must be checked to ensure they are within bounds. This is because, in practice, array-bounds checking in scientific applications may increase execution time by more than a factor of 2. Previous research has explored optimizations to statically eliminate bounds checks, but the dynamic nature of many scientific codes makes this difficult or impossible. Our approach is, instead, to create a compiler and operating system infrastructure that does not generate explicit bounds checks. It instead places arrays inside of Index Confinement Regions (ICRs), which are large, isolated, mostly unmapped virtual memory regions. Any array reference outside of its bounds will cause a protection violation; this provides implicit bounds checking. Our results show that when applying this infrastructure to high-performance computing programs written in Java, the overhead of bounds checking relative to a program with no bounds checks is reduced from an average of 63% to an average of 9%. Chris Bentley, Scott A. Watterson, David K. Lowenthal, Barry Rountree |
ACM Trans. Archit. Code Optim. | 3 |
| 2006 | Client-Centered, Energy-Efficient Wireless Communication on IEEE 802.11b NetworksabstractIn mobile devices, the wireless network interface card (WNIC) consumes a significant portion of overall system energy. One way to reduce energy consumed by a device is to transition its WNIC to a lower-power sleep mode when data is not being received or transmitted. In this paper, we investigate client-centered techniques for energy efficient communication, using IEEE 802.11b, within the network layer. The basic idea is to conserve energy by keeping the WNIC in high-power mode only when necessary. We track each connection, which allows us to determine inactive intervals during which to transition the WNIC to sleep mode. Whenever necessary, we also shape the traffic from the client side to maximize sleep intervals—convincing the server to send data in bursts. This trades lower WNIC energy consumption for an increase in transmission time. Our techniques are compatible with standard TCP and do not rely on any assistance from the server or network infrastructure. Results show that during Web browsing, our client-centered technique saved 21 percent energy compared to PSM and incurred less than a 1 percent increase in transmission time compared to regular TCP. For a large file download, our scheme saved 27 percent energy on average with a transmission time increase of only 20 percent. Haijin Yan, Scott A. Watterson, David K. Lowenthal, Kang Li 0001, Rupa Krishnan, Larry L. Peterson |
IEEE Trans. Mob. Comput. | 3 |
| 2005 | ACE: an active, client-directed method for reducing energy during web browsingabstractIn mobile devices, the wireless network interface card (WNIC) consumes a significant portion of overall system energy. One way to reduce energy consumed by a device is to transition its WNIC to a lower-power sleep mode when data is not being received or transmitted.This paper develops ACE, an active, client-directed technique to improve energy efficiency during web browsing. ACE actively retrieves buffered packets from an access point based on predictions made through client-side connection tracking. The key novel implementation technique used in ACE is connection rescheduling, which results is a better energy/time tradeoff for interactive applications such as web browsing. We demonstrate the effectiveness of ACE through actual experiments to real Internet servers. Haijin Yan, David K. Lowenthal, Kang Li 0001 |
NOSSDAV | 2 |
| 2005 | Using multiple energy gears in MPI programs on a power-scalable clusterabstractRecently, system architects have built low-power, high-performance clusters, such as Green Destiny. The idea behind these clusters is to improve the energy efficiency of nodes. However, these clusters save power at the expense of performance. Our approach is instead to use high-performance cluster nodes that are frequency- and voltage-scalable; energy can than be saved by scaling down the CPU. Our prior work has examined the costs and benefits of executing an entire application at a single reduced frequency.This paper presents a framework for executing a single application in several frequency-voltage settings. The basic idea is to first divide programs into phases and then execute a series of experiments, with each phase assigned a prescribed frequency. During each experiment, we measure energy consumption and time and then use a heuristic to choose the assignment of frequency to phase for the next experiment.Our results show that significant energy can be saved without an undue performance penalty; particularly, our heuristic finds assignments of frequency to phase that is superior to any fixed-frequency solution. Specifically, this paper shows that more than half of the NAS benchmarks exhibit a better energy-time tradeoff using multiple gears than using a single gear. For example, IS using multiple gears uses 9% less energy and executes in 1% less time than the closest single-gear solution. Compared to no frequency scaling, multiple gear IS uses 16% less energy while executing only 1% longer. Vincent W. Freeh, David K. Lowenthal |
PPoPP | 2 |
| 2005 | Just In Time Dynamic Voltage Scaling: Exploiting Inter-Node Slack to Save Energy in MPI ProgramsabstractRecently, improving the energy efficiency of HPC machines has become important. As a result, interest in using powerscalable clusters, where frequency and voltage can be dynamically modified, has increased. On power-scalable clusters, one opportunity for saving energy with little or no loss of performance exists when the computational load is not perfectly balanced. This situation occurs frequently, as balancing load between nodes is one of the long standing problems in parallel and distributed computing. In this paper we present a system called Jitter, which reduces the frequency on nodes that are assigned less computation and therefore have slack time. This saves energy on these nodes, and the goal of Jitter is to attempt to ensure that they arrive "just in time" so that they avoid increasing overall execution time. For example, in Aztec, from the ASCI Purple suite, our algorithm uses 8% less energy while increasing execution time by only 2.6%. Nandini Kappiah, Vincent W. Freeh, David K. Lowenthal |
SC | 3 |
| 2005 | The MHETA Execution Model for Heterogeneous ClustersabstractThe availability of inexpensive "off the shelf" machines increases the likelihood that parallel programs run on heterogeneous clusters of machines. These programs are increasingly likely to be out of core, meaning that portions of their datasets must be stored on disk during program execution. This results in significant, per-iteration, I/O cost. This paper describes an execution model, called MHETA, which is the key component to finding an effective data distribution on heterogeneous clusters. MHETA takes into account computation, communication, and I/O costs of iterative scientific applications. MHETA uses automatically extracted information from a single iteration to predict the execution time of the remaining iterations. Results show that MHETA predicts with on average 98% accuracy the execution time of several scientific benchmarks (with and without prefetching) and one full-scale scientific program that utilize pipelined and other communication. MHETA is thus an effective tool when searching for the most effective distribution on a heterogeneous cluster. Mario Nakazawa, David K. Lowenthal, Wenduo Zhou |
SC | 2 |
| 2005 | Towards cooperation fairness in mobile ad hoc networksabstractFor the sustainable operation of ad hoc networks, incentive mechanisms are required to encourage cooperation. More importantly, we must enforce available bandwidth fairness among nodes. Bandwidth sharing is fair if a node's available bandwidth is proportional to its forwarding contribution. In this paper, we achieve fair bandwidth sharing through a new packet scheduling algorithm, called cooperative queueing, on each node. In cooperative queueing, which is analogous to fair queueing, packet scheduling is based on a new abstraction that we call cooperation coefficient (CC). The CC quantifies how much a given node contributes to and consumes from the ad-hoc network; the larger the CC, the more bandwidth a node can obtain. We exploit the widely used dynamic source routing information to obtain the CC. We evaluate the effectiveness of cooperative queueing with different parameters and network configurations. We demonstrate that our algorithm is able to encourage cooperation and ensure fair sharing of bandwidth between nodes. We show that cooperative queueing is simple and has little overhead. Haijin Yan, David K. Lowenthal |
WCNC | 2 |
| 2005 | An MPI prototype for compiled communication on Ethernet switched clusters
Amit Karwande, Xin Yuan 0001, David K. Lowenthal |
J. Parallel Distributed Comput. | 3 |
| 2004 | Dynamic, Power-Aware Scheduling for Mobile Clients Using a Transparent ProxyabstractMobile computers consume significant amounts of energy when receiving large files. The wireless network interface card (WNIC) is the primary source of this energy consumption. One way to reduce the energy consumed is to transmit the packets to clients in a predictable fashion. Specifically, the packets can be sent in bursts to clients, who can then switch to a lower power sleep state between bursts. This technique is especially effective when the bandwidth of a stream is small. This work investigates techniques for saving energy in a multiple-client scenario, where clients may be receiving either UDP or TCP data. Energy is saved by using a transparent proxy that is invisible to both clients and servers. The proxy implementation maintains separate connections to the client and server so that a large increase in transmission time is avoided. The proxy also buffers data and dynamically generates a global transmission schedule that includes all active clients. Results show that energy savings within 10-15% of optimal are common, with little packet loss. Michael Gundlach, Sarah Doster, Haijin Yan, David K. Lowenthal, Scott A. Watterson, Surendar Chandra |
ICPP | 4 |
| 2004 | Implicit java array bounds checking on 64-bit architectureabstractInterest in using Java for high-performance parallel computing has increased in recent years. One obstacle that has inhibited Java from widespread acceptance in the scientific community is the language requirement that all array accesses must be checked to ensure they are within bounds. In practice, array bounds checking in scientific applications may increase execution time by more than a factor of 2. Previous research has explored optimizations to statically eliminate bounds checks, but the dynamic nature of many scientific codes makes this difficult or impossible.Our approach is instead to create a new Java implementation that does not generate explicit bounds checks. It instead places arrays inside of Index Confinement Regions (ICRs), which are large, isolated, mostly unmapped virtual memory regions. Any array reference outside of its bounds will cause a protection violation; this provides implicit bounds checking. Our results show that our new Java implementation reduces the overhead of bounds checking from an average of 63% to an average of 9% on our benchmarks. Chris Bentley, Scott A. Watterson, David K. Lowenthal, Barry Rountree |
ICS | 3 |
| 2004 | Client-centered energy and delay analysis for TCP downloadsabstractIn mobile devices, the wireless network interface card (WNIC) consumes a significant portion of overall system energy. One way to reduce energy consumed by a mobile device is to transition its WNIC to a lower-power sleep mode when data is not being received or transmitted. This paper investigates client-centered techniques for trading download time for energy savings during TCP downloads, in an attempt to reduce the energy' delay product. Effectively saving WNIC energy during a TCP download is difficult because TCP streams tend to be smooth, leaving little potential sleep time. The basic idea behind our technique is that the client increases the amount of time that can be spent in sleep mode by shaping the traffic. In particular, the client convinces the server to send data in predictable bursts, trading lower WNIC energy cost for increased transmission time. Our technique does not rely on any assistance from the server, a proxy, or IEEE 802.11b power-saving mode. Results show that in Internet experiments our scheme outperforms baseline TCP by 64% in the best case, with an average improvement of 19%. Haijin Yan, Rupa Krishnan, Scott A. Watterson, David K. Lowenthal, Kang Li 0001, Larry L. Peterson |
IWQoS | 4 |
| 2004 | Client-centered energy savings for concurrent HTTP connectionsabstractIn mobile devices, the wireless network interface card (WNIC) consumes a significant portion of overall system energy. One way to reduce energy consumed by a WNIC is to transition it to a lower-power sleep mode when data is not being received or transmitted.This paper investigates client-centered techniques for saving energy during web browsing. The basic idea is that the client predicts when packets will arrive, keeping the WNIC in high-power mode only when necessary. This is challenging because web browsing generally results in concurrent HTTP connections. To handle this, we maintain the state of each open connection on the client and then transition the WNIC to sleep mode when no connection is receiving data. Our technique is compatible with standard TCP and does not rely on any assistance from the server, a proxy, or IEEE 802.11b power-saving mode (PSM). Our technique combines the performance of regular TCP with nearly all the energy-saving of PSM during web downloads, and we save more energy than PSM during client think times. Results show that over an entire web browsing session (downloads and think times), our scheme saves up to 21% energy compared to PSM and incurs less than a 1% increase in transmission time compared to regular TCP. Haijin Yan, Rupa Krishnan, Scott A. Watterson, David K. Lowenthal |
NOSSDAV | 4 |
| 2003 | CC-MPI: a compiled communication capable MPI prototype for ethernet switched clustersabstractNo abstract available. Amit Karwande, Xin Yuan 0001, David K. Lowenthal |
PPoPP | 3 |
| 2003 | CC-MPI: a compiled communication capable MPI prototype for ethernet switched clustersabstractCompiled communication has recently been proposed to improve communication performance for clusters of workstations. The idea of compiled communication is to apply more aggressive optimizations to communications whose information is known at compile time. Existing MPI libraries do not support compiled communication. In this paper, we present an MPI prototype, CC--MPI, that supports compiled communication on Ethernet switched clusters. The unique feature of CC--MPI is that it allows the user to manage network resources such as multicast groups directly and to optimize communications based on the availability of the communication information. CC--MPI optimizes one--to--all, one--to--many, all--to--all, and many--to--many collective communication routines using the compiled communication technique. We describe the techniques used in CC--MPI and report its performance. The results show that communication performance of Ethernet switched clusters can be significantly improved through compiled communication. Amit Karwande, Xin Yuan 0001, David K. Lowenthal |
PPoPP | 3 |
| 2003 | Dyn-MPI: Supporting MPI on Non Dedicated ClustersabstractDistributing data is a fundamental problem in implementing efficient distributed-memory parallel programs. The problem becomes more difficult in environments where the participating nodes are not dedicated to a parallel application. We are investigating the data distribution problem in non dedicated environments in the context of explicit message-passing programs. To address this problem, we have designed and implemented an extension to MPI called Dynamic MPI (Dyn-MPI). The key component of Dyn-MPI is its run-time system, which efficiently and automatically redistributes data on the fly when there are changes in the application or the underlying environment. Dyn-MPI supports efficient memory allocation, precise measurement of system load and computation time, and node removal. Performance results show that programs that use Dyn-MPI execute efficiently in non dedicated environments, including up to almost a three-fold improvement compared to programs that do not redistribute data and a 25% improvement over standard adaptive load balancing techniques. D. Brent Weatherly, David K. Lowenthal, Mario Nakazawa, Franklin Lowenthal |
SC | 2 |
| 2003 | A comparative analysis of fine-grain threads packages
Gregory W. Price, David K. Lowenthal |
J. Parallel Distributed Comput. | 2 |
| 2001 | Accurate data redistribution cost estimation in software distributed shared memory systemsabstractDistributing data is one of the key problems in implementing efficient distributed-memory parallel programs. The problem becomes more difficult in programs where data redistribution between computational phases is considered. The global data distribution problem is to find the optimal distribution in multi-phase parallel programs. Solving this problem requires accurate knowledge of data redistribution cost. Donald G. Morris, David K. Lowenthal |
PPoPP | 2 |
| 2000 | Architecture-independent parallelism for both shared- and distributed-memory machines using the Filaments package
David K. Lowenthal, Vincent W. Freeh |
Parallel Comput. | 1 |
| 1998 | Efficient support for fine-grain parallelism on shared-memory machinesabstractA coarse-grain parallel program typically has one thread (task) per processor, whereas a fine-grain program has one thread for each independent unit of work. Although there are several advantages to fine-grain parallelism, conventional wisdom is that coarse-grain parallelism is more efficient. This paper illustrates the advantages of fine-grain parallelism and presents an efficient implementation for shared-memory machines. The approach has been implemented in a portable software package called Filaments, which employs a unique combination of techniques to achieve efficiency. The performance of the fine-grain programs discussed in this paper is always within 13% of a hand-coded coarse-grain program and is usually within 5%. © 1998 John Wiley & Sons, Ltd. David K. Lowenthal, Vincent W. Freeh, Gregory R. Andrews |
Concurr. Pract. Exp. | 1 |
| 1996 | Using Fine-Grain Threads and Run-Time Decision Making in Parallel Computing
David K. Lowenthal, Vincent W. Freeh, Gregory R. Andrews |
J. Parallel Distributed Comput. | 1 |
| 1994 | Distributed Filaments: Efficient Fine-Grain Parallelism on a Cluster of Workstations
Vincent W. Freeh, David K. Lowenthal, Gregory R. Andrews |
OSDI | 2 |