Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jerry C. Yan

dblp:91/2253 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
0since 2021 · last 2000
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Parallel and multicore computing · 36% Performance modeling and evaluation · 35% Distributed systems · 22%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
grid computing
0.012000
An Evaluation of Alternative Designs for a Grid Information Service · HPDC 2000
Performance modeling and evaluation › simulation › discrete-event simulation
trace-driven simulation
0.012000
An Evaluation of Alternative Designs for a Grid Information Service · HPDC 2000
Compilers and program optimization
parallelizing compiler
0.011998
A Comparison of Automatic Parallelization Tools/Compilers on the SGI Origin 2000 · SC 1998
Parallel and multicore computing › parallel programming models
automatic parallelization
0.011998
A Comparison of Automatic Parallelization Tools/Compilers on the SGI Origin 2000 · SC 1998
Parallel and multicore computing
parallel programming models
0.011998
A Comparison of Automatic Parallelization Tools/Compilers on the SGI Origin 2000 · SC 1998
Performance modeling and evaluation › parallel system performance
parallel performance modeling
0.011994
A Comparison of Two Model-Based Performance-Prediction Techniques for Message-Passing Parallel Programs · SIGMETRICS 1994
Performance modeling and evaluation
performance prediction
0.011994
A Comparison of Two Model-Based Performance-Prediction Techniques for Message-Passing Parallel Programs · SIGMETRICS 1994
Distributed systems › distributed resource management
resource discovery
0.012000
An Evaluation of Alternative Designs for a Grid Information Service · HPDC 2000
Parallel and multicore computing › task allocation
process placement
0.011991
Intelligent mapping of communicating processes in distributed computing systems · SC 1991
Performance modeling and evaluation
benchmarking
0.011998
A Comparison of Automatic Parallelization Tools/Compilers on the SGI Origin 2000 · SC 1998
Cloud and datacenter computing
resource management
0.011989
The Post-Game Analysis Framework - Developing Resource Management Strategies for Concurrent Systems · IEEE Trans. Knowl. Data Eng. 1989
Parallel and multicore computing
software partitioning
0.011989
The Post-Game Analysis Framework - Developing Resource Management Strategies for Concurrent Systems · IEEE Trans. Knowl. Data Eng. 1989
Parallel and multicore computing
message-passing programs
0.011994
A Comparison of Two Model-Based Performance-Prediction Techniques for Message-Passing Parallel Programs · SIGMETRICS 1994
Embedded and real-time systems
distributed real-time systems
0.011991
Intelligent mapping of communicating processes in distributed computing systems · SC 1991
Electronic design automation › high-level synthesis
scheduling
0.011989
The Post-Game Analysis Framework - Developing Resource Management Strategies for Concurrent Systems · IEEE Trans. Knowl. Data Eng. 1989

Methods — techniques the papers use, named apart from their topics

directive-based parallelization · 0.0HPF compilation · 0.0trace-driven simulation · 0.0simulation · 0.0statistical regression · 0.0profiling · 0.0statistical methods · 0.0post-game analysis · 0.0rule-based architecture · 0.0heuristic optimization · 0.0
YearPublicationVenuePosition
2000 An Evaluation of Alternative Designs for a Grid Information Service
abstract
Computational grids consisting of large and diverse sets of distributed resources have recently been adopted by organizations such as NASA and the NSF. One key component of a computational grid is an information service that provides information about resources, services and applications to users and their tools. This information is required to use a computational grid and therefore should be available in a timely and reliable manner. In this work, we describe the Globus information service, describe how this service is used, analyze its current performance, and perform trace-driven simulations to evaluate alternative implementations of this grid information service. We find that the majority of the transactions with the information service are changes to the data maintained by the service. We also find that of the three servers we evaluate, one of the commercial products provides the best performance for our workload and that the response time of the information service was not improved during the single experiment we performed with data distributed across two servers.
Warren Smith, David Meyers, Jerry C. Yan
HPDC4
1999 Parallelization of NAS benchmarks for shared memory multiprocessors
Jerry C. Yan
Future Gener. Comput. Syst.2
1998 Performance Modeling and Measurement of Parallelized Code for Distributed Shared Memory Multiprocessors
abstract
This paper presents a model to evaluate the performance and overhead of parallelizing sequential code using compiler directives for multiprocessing on distributed shared memory (DSM) systems. We parallelized the sequential implementation of NAS benchmarks using native Fortran77 compiler directives on an Origin2000, which is a DSM system. We report measurement based performance of these parallelized benchmarks from four perspectives: efficacy of parallelization process; scalability; parallelization overhead; and comparison with hand-parallelized and -optimized version of the same benchmarks. Our results indicate that sequential programs can conveniently be parallelized for DSM systems using compiler directives but realizing performance gains as predicted by the performance model depends primarily on minimizing architecture-specific data locality overhead.
Jerry C. Yan
MASCOTS2
1998 A Comparison of Automatic Parallelization Tools/Compilers on the SGI Origin 2000
abstract
Porting applications to new high performance parallel and distributed computing platforms is a challenging task. Since writing parallel code by hand is time consuming and costly, porting codes would ideally be automated by using some parallelization tools and compilers. In this paper, we compare the performance of three parallelization tools and compilers based on the NAS Parallel Benchmark and a CFD application, ARC3D, on the SGI Origin2000 multiprocessor. The tools and compilers compared include: 1) CAPTools: an interactive computer aided parallelization toolkit, 2) Portland Group's HPF compiler, and 3) the MIPSPro FORTRAN compiler available on the Origin2000, with support for shared memory multiprocessing directives and MP runtime library. The tools and compilers are evaluated in four areas: 1) required user interaction, 2) limitations, 3) portability and 4) performance. Based on these results, a discussion on the feasibility of computer-aided parallelization of aerospace applications is presented along with suggestions for future work.
Michael A. Frumkin, Michelle R. Hribar, Haoqiang Jin, Jerry C. Yan
SC5
1996 Analyzing Parallel Program Performance Using Normalized Performance Indices and Trace Transformation Techniques
Jerry C. Yan, Sekhar R. Sarukkai
Parallel Comput.1
1995 Performance Measurement, Visualization and Modeling of Parallel and Distributed Programs using the AIMS Toolkit
abstract
Abstract Writing large‐scale parallel and distributed scientific applications that make optimum use of the multiprocessor is a challenging problem. Typically, computational resources are underused due to performance failures in the application being executed. Performance‐tuning tools are essential for exposing these performance failures and for suggesting ways to improve program performance. In this paper, we first address fundamental issues in building useful performance‐tuning tools and then describe our experience with the AIMS toolkit for tuning parallel and distributed programs on a variety of platforms. AIMS supports source‐code instrumentation, run‐time monitoring, graphical execution profiles, performance indices and automated modeling techniques as ways to expose performance problems of programs. Using several examples representing a broad range of scientific applications, we illustrate AIMS' effectiveness in exposing performance problems in parallel and distributed programs.
Jerry C. Yan, Sekhar R. Sarukkai, Pankaj Mehra
Softw. Pract. Exp.1
1994 Normalized performance indices for message passing parallel programs
abstract
Existing tools for locating performance bottlenecks of message passing parallel programs either provide visualizations or profiles of program executions only; they do not highlight the cause of poor program performance. From the perspective of the application, the location and cause of performance problems in terms of procedures, processors and data structures are all important. Identifying the cause of poor performance necessitates the need to expose how well the underlying algorithm has been mapped onto the parallel machine.
Sekhar R. Sarukkai, Jerry C. Yan, Jacob Gotwals
International Conference on Supercomputing2
1994 A Comparison of Two Model-Based Performance-Prediction Techniques for Message-Passing Parallel Programs
abstract
This paper describes our experience in modeling two significant parallel applications: ARC2D, a 2-dimensional Euler solver; and, Xtrid, a tridiagonal linear solver. Both of these models were expressed in BDL (Behavior Description language) and simulated on an iPSC/860 Hypercube modeled using Axe (Abstract eXecution Environment). BDL models consist of abstract communicating objects: blocks of sequential code are modeled by single RUN statements; all communication operations in the original code are mirrored by corresponding BDL operations in the model. Our ARC2D model was built by first profiling the program to locate the significant loops and then timing the basic blocks within those loops. Simulated completion times were (except in one case) within 8% of measured execution times. Lengthy simulations were necessary for predicting the performance of large-scale runs. For Xtrid, only the loops surrounding communications were modeled; other loops were absorbed into large sequential blocks whose complexity was estimated using statistical regression. This approach yielded a much smaller model whose computation and communication complexities were clearly manifest. Analysis of complexity allowed rapid prediction of large-scale performance without lengthy simulations! Analytically predicted speed-ups were within 7% of those predicted by simulation. Simulated completion times were within 5% of measured execution times. The second approach provides a more effective methodology for simulation-based performance-tuning.
Pankaj Mehra, Catherine H. Schulbach, Jerry C. Yan
SIGMETRICS3
1992 Intelligent Process Mapping through Systematic Improvement of Heuristics
Arthur Ieumwananonthachai, Akiko Aizawa, Steven R. Schwartz, Benjamin W. Wah, Jerry C. Yan
J. Parallel Distributed Comput.5
1991 New "Post-game Analysis" Heuristics for Mapping Parallel Computations to Hypercubes
Jerry C. Yan
ICPP (2)1
1991 Intelligent mapping of communicating processes in distributed computing systems
abstract
In this paper we present TEACHER 4.1, a system for designing automatically heuristics that map a set of cmnmunicating processes on a real-time distributed computing system.The problem of optimal promas mapping is NP-hard and involves the optimal placement of precesses on the distributed system and the optimaf routing of messages from one computer to another.The design of efficient and robust heuristics is often ad hoc and is guided by intuition and experience of the designers.In this paper we develop a statistical method to explore systematically the space of possible heuristics for process mapp.mg.The method operates under a specified time constraint and intends to get the best possible heuristics whale trading between the solution quality and the execution time of the process mapping heuristics.Our prototype for process mWPing is extended hom post-game analysis, a system that uses a set of user-specified roles for generating new mappings.It tunes pamneters of these rules and proposea new heuristics for process mapping.Simulations show that there is significant improvement m performance through systematic and automatic exploration of the spare of heuristics.
Arthur Ieumwananonthachai, Akiko Aizawa, Steven R. Schwartz, Benjamin W. Wah, Jerry C. Yan
SC5
1989 The Post-Game Analysis Framework - Developing Resource Management Strategies for Concurrent Systems
abstract
Research has been conducted to determine how distributed computations can be mapped to multiprocessors to minimize execution time. The approach described here, known as post-game analysis, incrementally changes the program partitioning in between program execution time in subsequent runs. Post-game analysis differs from conventional iterative refinement or controlled opportunistic perturbation in that no abstract program models or any single objective function are employed to determine the relative merits of two alternative mappings. Multiple optimization subgoals are formulated, based on actual timing data gathered during program execution. Heuristics, based on various optimization subgoals, are then applied to propose changes to the current mapping. Finally, a mapping generation process which prioritizes and resolves conflicting proposals is applied. Results obtained from simulations show that post-game analysis consistently out-performs random placement, load-balancing, and clustering algorithms by 15%. Few iterations are required for simulations involving more than 200 processes and 64 sites. A rule-based architecture enables incremental strategy refinement, thus making post-game analysis easily tailorable to programs written in many concurrent programming paradigms and multiprocessor architectures.>
Jerry C. Yan, Stefen F. Lundstrom
IEEE Trans. Knowl. Data Eng.1