Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jay P. Hoeflinger

dblp:09/573 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
0since 2021 · last 2009
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 2 first-authorSoftware engineering, systems software and programming languages · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Parallel and multicore computing · 71% Memory systems · 12% High-performance computing · 10%
Software engineering, system software, and programming languages
5 papers
Compilers and program optimization · 88% Runtime systems and virtual machines · 12%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 28 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
parallel programming models
0.122009
The Design of OpenMP Tasks · IEEE Trans. Parallel Distributed Syst. 2009
A sampling-based framework for parallel data mining · PPoPP 2005
Parallel and multicore computing › parallel programming models › task parallelism
OpenMP tasking
0.112009
The Design of OpenMP Tasks · IEEE Trans. Parallel Distributed Syst. 2009
Parallel and multicore computing › parallel programming models
task parallelism
0.112009
The Design of OpenMP Tasks · IEEE Trans. Parallel Distributed Syst. 2009
Parallel and multicore computing › parallel programming models
shared-memory parallelization
0.122007
Programming with cluster openMP · PPoPP 2007
An Advanced Compiler Framework for Non-Cache-Coherent Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2002
High-performance computing
cluster computing
0.112007
Programming with cluster openMP · PPoPP 2007
Memory systems › shared memory
distributed shared memory
0.112007
Programming with cluster openMP · PPoPP 2007
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP
0.112007
Programming with cluster openMP · PPoPP 2007
Compilers and program optimization › dependence analysis
array access analysis
0.122002
Efficient and precise array access analysis · ACM Trans. Program. Lang. Syst. 2002
Simplification of Array Access Patterns for Compiler Optimizations · PLDI 1998
Compilers and program optimization › parallelization
automatic parallelization
0.122002
An Advanced Compiler Framework for Non-Cache-Coherent Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2002
On the Automatic Parallelization of the Perfect Benchmarks · IEEE Trans. Parallel Distributed Syst. 1998
Compilers and program optimization
loop optimization
0.122002
Efficient and precise array access analysis · ACM Trans. Program. Lang. Syst. 2002
Simplification of Array Access Patterns for Compiler Optimizations · PLDI 1998
Data mining › pattern mining › itemset mining
frequent itemset mining
0.112005
A sampling-based framework for parallel data mining · PPoPP 2005
Data mining › big data analytics › large-scale data mining
parallel data mining
0.112005
A sampling-based framework for parallel data mining · PPoPP 2005
Data mining
pattern mining
0.112005
A sampling-based framework for parallel data mining · PPoPP 2005
Data mining › pattern mining
sequential pattern mining
0.112005
A sampling-based framework for parallel data mining · PPoPP 2005
Runtime systems and virtual machines
parallel runtime systems
0.012009
The Design of OpenMP Tasks · IEEE Trans. Parallel Distributed Syst. 2009
Compilers and program optimization › parallelization › automatic parallelization
array privatization
0.011998
On the Automatic Parallelization of the Perfect Benchmarks · IEEE Trans. Parallel Distributed Syst. 1998
Compilers and program optimization
dependence analysis
0.011998
Simplification of Array Access Patterns for Compiler Optimizations · PLDI 1998
Parallel and multicore computing › parallel algorithms › parallel algorithm design
divide-and-conquer parallelization
0.012005
A sampling-based framework for parallel data mining · PPoPP 2005
Memory systems
non-uniform memory access
0.012002
An Advanced Compiler Framework for Non-Cache-Coherent Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2002
Performance modeling and evaluation
benchmarking
0.011993
The Cedar System and an Initial Performance Study · ISCA 1993
Performance modeling and evaluation › parallel system performance
multiprocessor performance evaluation
0.011993
The Cedar System and an Initial Performance Study · ISCA 1993
Performance modeling and evaluation
parallel performance evaluation
0.011993
The Cedar System and an Initial Performance Study · ISCA 1993
Parallel and multicore computing
multiprocessor system
0.011998
On the Automatic Parallelization of the Perfect Benchmarks · IEEE Trans. Parallel Distributed Syst. 1998
Parallel and multicore computing › parallel computing
parallel programming languages
0.011988
Cedar Fortran and other Vector and parallel Fortran dialects · SC 1988
Processor architecture and microarchitecture
vector processor
0.011988
Cedar Fortran and other Vector and parallel Fortran dialects · SC 1988
Performance modeling and evaluation › benchmarking
parallel benchmark
0.011993
The Cedar System and an Initial Performance Study · ISCA 1993
Performance modeling and evaluation
workload characterization
0.011993
The Cedar System and an Initial Performance Study · ISCA 1993
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor
0.011988
Cedar Fortran and other Vector and parallel Fortran dialects · SC 1988

Methods — techniques the papers use, named apart from their topics

task-based parallelism · 0.2selective sampling · 0.1load balancing · 0.1performance tuning · 0.1debugging · 0.1compiler framework design · 0.1hand parallelization · 0.0compiler transformation · 0.0abstract interpretation · 0.0static analysis · 0.0performance measurement · 0.0benchmarking methodology · 0.0
YearPublicationVenuePosition
2009 The Design of OpenMP Tasks
abstract
OpenMP has been very successful in exploiting structured parallelism in applications. With increasing application complexity, there is a growing need for addressing irregular parallelism in the presence of complicated control structures. This is evident in various efforts by the industry and research communities to provide a solution to this challenging problem. One of the primary goals of OpenMP 3.0 is to define a standard dialect to express and efficiently exploit unstructured parallelism. This paper presents the design of the OpenMP tasking model by members of the OpenMP 3.0 tasking sub-committee which was formed for this purpose. The paper summarizes the efforts of the sub-committee (spanning over two years) in designing, evaluating and seamlessly integrating the tasking model into the OpenMP specification. In this paper, we present the design goals and key features of the tasking model, including a rich set of examples and an in-depth discussion of the rationale behind various design choices. We compare a prototype implementation of the tasking model with existing models, and evaluate it on a wide range of applications. The comparison shows that the OpenMP tasking model provides expressiveness, flexibility, and huge potential for performance and scalability.
Eduard Ayguadé, Nawal Copty, Alejandro Duran, Jay P. Hoeflinger, Federico Massaioli, Xavier Teruel, Priya Unnikrishnan, Guansong Zhang
IEEE Trans. Parallel Distributed Syst.4
2007 Programming with cluster openMP
abstract
This full-day tutorial will teach the attendees about Cluster OpenMP and the tools that are available to assist the programmer in debugging and tuning. Cluster OpenMP is an Intel® programming system that allows the user to run an OpenMP program on a cluster of computers without a common hardware shared memory. The tutorial will consist of a short tutorial on OpenMP, a longer description of Cluster OpenMP, its concepts, mechanisms and tools, a set of short hands-on porting exercises for the participants, and a set of exercises with the Cluster OpenMP debugging and tuning tools.
Jay P. Hoeflinger
PPoPP1
2006 Debugging Distributed Shared Memory Applications
Jeffrey Olivier, Chih-Ping Chen, Jay P. Hoeflinger
ISPA3
2005 A sampling-based framework for parallel data mining
abstract
The goal of data mining algorithm is to discover useful information embedded in large databases. Frequent itemset mining and sequential pattern mining are two important data mining problems with broad applications. Perhaps the most efficient way to solve these problems sequentially is to apply a pattern-growth algorithm, which is a divide-and-conquer algorithm [9, 10]. In this paper, we present a framework for parallel mining frequent itemsets and sequential patterns based on the divide-and-conquer strategy of pattern growth. Then, we discuss the load balancing problem and introduce a sampling technique, called selective sampling, to address this problem. We implemented parallel versions of both frequent itemsets and sequential pattern mining algorithms following our framework. The experimental results show that our parallel algorithms usually achieve excellent speedups.
Shengnan Cong, Jiawei Han 0001, Jay P. Hoeflinger, David A. Padua
PPoPP3
2005 A compiler for exploiting nested parallelism in OpenMP programs
Xinmin Tian, Jay P. Hoeflinger, Grant Haab, Yen-Kuang Chen, Milind Girkar, Sanjiv Shah
Parallel Comput.2
2002 Hybrid analysis: static & dynamic memory reference analysis
abstract
We present a novel Hybrid Analysis technology which can efficiently and seamlessly integrate all static and run-time analysis of memory references into a single framework that is capable of performing all data dependence analysis and can generate necessary information for most associated memory related optimizations. We use HA to perform automatic parallelization by extracting run-time assertions from any loop and generating appropriate run-time tests that range from a low cost scalar comparison to a full, reference by reference run-time analysis. Moreover we can order the run-time tests in increasing order of complexity (overhead) and thus risk the minimum necessary overhead. We accomplish this by both extending compile time IP analysis techniques and by incorporating speculative run-time techniques when necessary. Our solution is to bridge 'free' compile time techniques with exhaustive run-time techniques through a continuum of simple to complex solutions. We have implemented our framework in the Polaris compiler by introducing an innovative intermediate representation called RT_LMAD and a run-time library that can operate on it. Based on the experimental results obtained to date we hope to to automatically parallelize most and possibly all PERFECT codes, a significant accomplishment.
Silvius Vasile Rus, Lawrence Rauchwerger, Jay P. Hoeflinger
ICS3
2002 Efficient and precise array access analysis
abstract
A number of existing compiler techniques hinge on the analysis of array accesses in a program. The most important task in array access analysis is to collect the information about array accesses of interest and summarize it in some standard form. Traditional forms used in array access analysis are sensitive to the complexity of array subscripts; that is, they are usually quite accurate and efficient for simple array subscripting expressions, but lose accuracy or require potentially expensive algorithms for complex subscripts. Our study has revealed that in many programs, particularly numerical applications, many access patterns are simple in nature even when the subscripting expressions are complex. Based on this analysis, we have developed a new, general array region representational form, called the linear memory access descriptor (LMAD). The key idea of the LMAD is to relate all memory accesses to the linear machine memory rather than to the shape of the logical data structures of a programming language. This form helps us expose the simplicity of the actual patterns of array accesses in memory, which may be hidden by complex array subscript expressions. Our recent experimental studies show that our new representation simplifies array access analysis and, thus, enables efficient and accurate compiler analysis.
Yunheung Paek, Jay P. Hoeflinger, David A. Padua
ACM Trans. Program. Lang. Syst.2
2002 An Advanced Compiler Framework for Non-Cache-Coherent Multiprocessors
abstract
The Cray T3D and T3E are non-cache-coherent (NCC) computers with a NUMA structure. They have been shown to exhibit a very stable and scalable performance for a variety of application programs. Considerable evidence suggests that they are more stable and scalable than many other shared-memory multiprocessors. However, the principal drawback of these machines is a lack of programmability, caused by the absence of the global cache coherence that is necessary to provide a convenient shared view of memory in hardware. This forces the programmer to keep careful track of where each piece of data is stored, a complication that is unnecessary when a pure shared-memory view is presented to the user. We believe that a remedy for this problem is advanced compiler technology. In this paper, we present our experience with a compiler framework for automatic parallelization and communication generation that has the potential to reduce the time-consuming hand-tuning that would otherwise be necessary to achieve good performance with this type of machine. From our experiments, we learned that our compiler performs well for a variety of applications on the T3D and T3E and we found a few sophisticated techniques that could improve performance even more once they are fully implemented in the compiler.
Yunheung Paek, Angeles G. Navarro, Emilio L. Zapata, Jay P. Hoeflinger, David A. Padua
IEEE Trans. Parallel Distributed Syst.4
2001 A Parallel Programming Environment for a V-Busbased PC-cluste
abstract
Nowadays PC-cluster architectures are widely accepted for parallel computing. In a PC-cluster system, memories are physically distributed. To harness the computational power of a distributed-memory PC-cluster, a user must write efficient software for the machine by hand. The absence of global address space makes the process laborious because the user must manually assign computations to processors, distribute data across processors and explicitly manage communication among processors. This paper introduces our recent three related studies on parallel computing. The first is a V-Bus based PC-cluster in which all PCs are interconnected through V-Bus network cards. The second is the implementation of an one-sided communication library on our PC-cluster system, which provides the user with a view of global address space on top of the distributed-memory PC-cluster. This encapsulation of global address allows the user to write shared memory code on our cluster system, which simplifies programming our cluster because the user does not need to explicitly cope with distributed memories. The third is a parallelizing compiler that automatically translates sequential code to shared-memory code for the PC-cluster. This compiler not only enables legacy code to be compiled for our cluster, but also further simplifies programming our cluster by allowing the user to continue using conventional sequential languages such as Fortran 77. In this work, the compiler was optimized particularly for our V-Bus based PC-cluster. The paper also reports our experimental results with the compiler on our PC-cluster system.
Sang Seok Lim, Yunheung Paek, Jay P. Hoeflinger
CLUSTER4
2001 Monotonic evolution: an alternative to induction variable substitution for dependence analysis
abstract
We present a new approach to dependence testing in the presence of induction variables. Instead of looking for closed form expressions, our method computes monotonic evolution which captures the direction in which the value of a variable changes. This information is then used in the dependence test to help determine whether array references are dependence-free. Under this scheme, closed form computation and induction variable substitution can be delayed until after the dependence test and be performed on-demand. To improve computative efficiency, we also propose an optimized (non-iterative) data-flow algorithm to compute evolution. Experimental results show that dependence tests based on evolution information matches the accuracy of that based on closed-form computation (implemented in Polaris), and when no closed form expressions can be calculated, our method is more accurate than that of Polaris.
Peng Wu 0001, Albert Cohen 0001, Jay P. Hoeflinger, David A. Padua
ICS3
2001 A synthesis of memory mechanisms for distributed architectures
abstract
Producing efficient parallel programs for distributed memory multiprocessors is a difficult task. Hand-coding efficient parallel programs for these systems can be extremely difficult, time consuming and error-prone, so people have turned to the shared memory abstraction and automatic parallelizing compilers to ease the task. The two main approaches to this are using compilers that 1) generate message passing code, or 2) generate code for a distributed shared memory software layer. Neither has been completely successful for all types of programs. In this paper, we discuss the use of a combination of these mechanisms to produce a compiler code generation paradigm that can be successful for many user programs. The experimental results indicate that our new paradigm would be able to support both regular and irregular code efficiently.
Jiajing Zhu, Jay P. Hoeflinger, David A. Padua
ICS2
2001 Producing scalable performance with OpenMP: Experiments with two CFD applications
Jay P. Hoeflinger, Prasad Alavilli, Bob Kuhn
Parallel Comput.1
1998 Simplification of Array Access Patterns for Compiler Optimizations
abstract
Existing array region representation techniques are sensitive to the complexity of array subscripts. In general, these techniques are very accurate and efficient for simple subscript expressions, but lose accuracy or require potentially expensive algorithms for complex subscripts. We found that in scientific applications, many access patterns are simple even when the subscript expressions are complex. In this work, we present a new, general array access representation and define operations for it. This allows us to aggregate and simplify the representation enough that precise region operations may be applied to enable compiler optimizations. Our experiments show that these techniques hold promise for speeding up applications.
Yunheung Paek, Jay P. Hoeflinger, David A. Padua
PLDI2
1998 On the Automatic Parallelization of the Perfect Benchmarks
abstract
This paper presents the results of the Cedar Hand-Parallelization Experiment conducted from 1989 through 1992, within the Center for Supercomputing Research and Development (CSRD) at the University of Illinois. In this experiment, we manually transformed the Perfect Benchmarks(R) into parallel program versions. In doing so, we used techniques that may be automated in an optimizing compiler. We then ran these programs on the Cedar multiprocessor (built at CSRD during the 1980s) and measured the speed improvement due to each technique. The results presented here extend the findings previously reported. The techniques credited most for the performance gains include array privatization, parallelization of reduction operations, and the substitution of generalized induction variables. All these techniques can be considered extensions of transformations that were available in vectorizers and commercial restructuring compilers of the late 1980s. We applied these transformations by hand to the given programs, in a mechanical manner, similar to that of a parallelizing compiler. Because of our success with these transformations, we believed that it would be possible to implement many of these techniques in a new parallelizing compiler. Such a compiler has been completed in the meantime and we show preliminary results.
Rudolf Eigenmann, Jay P. Hoeflinger, David A. Padua
IEEE Trans. Parallel Distributed Syst.2
1993 The Cedar System and an Initial Performance Study
abstract
In this paper, we give an overview of the Cedar multiprocessor and present recent performance results. These include the performance of some computational kernels and the Perfect Benchmarks. We also present a methodology for judging parallel system performance and apply this methodology to Cedar, Cray YMP-8, and Thinking Machines CM-5.
David J. Kuck, Edward S. Davidson, Duncan H. Lawrie, Ahmed H. Sameh, Chuanqi Zhu, Alexander V. Veidenbaum, Jeff Konicek, Pen-Chung Yew, Kyle A. Gallivan, William Jalby, Harry A. G. Wijshoff, Randall Bramley, Ulrike Meier Yang, Perry A. Emrath, David A. Padua, Rudolf Eigenmann, Jay P. Hoeflinger, Greg P. Jaxon, Zhiyuan Li 0001, T. Murphy, John T. Andrews, Stephen W. Turner
ISCA17
1993 Restructuring Fortran programs for Cedar
abstract
Abstract The paper reports on the status of the Fortran translator for the Cedar computer at the end of March 1991. A brief description of the Cedar Fortran language is followed by a discussion of the Fortran‐77 to Cedar Fortran parallelizer that describes the techniques currently being implemented. A collection of experiments illustrate the effectiveness of the current implementation, and point toward new approaches to be incorporated into the system in the near future.
Rudolf Eigenmann, Jay P. Hoeflinger, Greg P. Jaxon, Zhiyuan Li 0001, David A. Padua
Concurr. Pract. Exp.2
1991 Restructuring Fortran Programs for Cedar
Rudolf Eigenmann, Jay P. Hoeflinger, Greg P. Jaxon, Zhiyuan Li 0001, David A. Padua
ICPP (1)2
1990 Cedar Fortran and other vector and parallel Fortran dialects
Mark D. Guzzi, David A. Padua, Jay P. Hoeflinger, Duncan H. Lawrie
J. Supercomput.3
1988 Cedar Fortran and other Vector and parallel Fortran dialects
abstract
The development of vector and multiprocessor language constructs in Fortran is outlined. The significant architectures, their languages, and optimizers are described. A description is given of Cedar Fortran, the language for the Cedar multiprocessor, a hierarchical, shared-memory, vector multiprocessor currently under development.>
Mark D. Guzzi, Jay P. Hoeflinger, David A. Padua, Duncan H. Lawrie
SC2