Jhy-Chun Wang

dblp:52/4444 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
0since 2021 · last 1995
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Parallel and multicore computing · 35% Interconnection networks and networks-on-chip · 27% Performance modeling and evaluation · 15%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization › parallel language compilation
data-parallel compilation
0.011995
An Integrated Compilation and Performance Analysis Environment for Data Parallel Programs · SC 1995
Performance modeling and evaluation
performance analysis tools
0.011995
An Integrated Compilation and Performance Analysis Environment for Data Parallel Programs · SC 1995
High-performance computing
collective communication
0.011994
Static and Run-Time Algorithms for All-to-Many Personalized Communication on Permutation Networks · IEEE Trans. Parallel Distributed Syst. 1994
Parallel and multicore computing › parallel scheduling
communication scheduling
0.011994
Scheduling of unstructured communication on the Intel iPSC/860 · SC 1994
Parallel and multicore computing
contention reduction
0.011994
Scheduling of unstructured communication on the Intel iPSC/860 · SC 1994
Interconnection networks and networks-on-chip
permutation network
0.011994
Static and Run-Time Algorithms for All-to-Many Personalized Communication on Permutation Networks · IEEE Trans. Parallel Distributed Syst. 1994
Electronic design automation › physical design
routing
0.011994
Static and Run-Time Algorithms for All-to-Many Personalized Communication on Permutation Networks · IEEE Trans. Parallel Distributed Syst. 1994
Interconnection networks and networks-on-chip
graph embedding
0.011990
Embedding meshes on the star graph · SC 1990
Interconnection networks and networks-on-chip
network topology
0.011990
Embedding meshes on the star graph · SC 1990
Parallel and multicore computing
parallel algorithms
0.011990
Embedding meshes on the star graph · SC 1990
Interconnection networks and networks-on-chip › network topology › cayley graph
star graph
0.011990
Embedding meshes on the star graph · SC 1990
Parallel and multicore computing › data parallelism
data-parallel program
0.011995
An Integrated Compilation and Performance Analysis Environment for Data Parallel Programs · SC 1995
Performance modeling and evaluation
performance instrumentation
0.011995
An Integrated Compilation and Performance Analysis Environment for Data Parallel Programs · SC 1995
Parallel and multicore computing › task scheduling
compile-time and runtime scheduling
0.011994
Static and Run-Time Algorithms for All-to-Many Personalized Communication on Permutation Networks · IEEE Trans. Parallel Distributed Syst. 1994
Parallel and multicore computing › parallel programming models
message passing
0.011994
Scheduling of unstructured communication on the Intel iPSC/860 · SC 1994
Electronic design automation › high-level synthesis
scheduling
0.011994
Static and Run-Time Algorithms for All-to-Many Personalized Communication on Permutation Networks · IEEE Trans. Parallel Distributed Syst. 1994

Methods — techniques the papers use, named apart from their topics

performance correlation · 0.0compiler instrumentation · 0.0static and runtime scheduling · 0.0performance measurement · 0.0partial permutation decomposition · 0.0approximate analysis · 0.0dilation and expansion analysis · 0.0
YearPublicationVenuePosition
1995 An Integrated Compilation and Performance Analysis Environment for Data Parallel Programs
abstract
Supporting source-level performance analysis of programs written in data-parallel languages requires a unique degree of integration between compilers and performance analysis tools. Compilers for languages such as High Performance Fortran infer parallelism and communication from data distribution directives, thus, performance tools cannot meaningfully relate measurements about these key aspects of execution performance to source-level constructs without substantial compiler support. This paper describes an integrated system for performance analysis of data-parallel programs based on the Rice Fortran 77D compiler and the Illinois Pablo performance analysis toolkit. During code generation, the Fortran D compiler records mapping information and semantic analysis results describing the relationship between performance instrumentation and the original source program. An integrated performance analysis system based on the Pablo toolkit uses this information to correlate the program's dynamic behavior with the data parallel source code. The integrated system provides detailed source-level performance feedback to programmers via a pair of graphical interfaces. Our strategy serves as a model for integration of data-parallel compilers and performance tools.
Vikram S. Adve, John M. Mellor-Crummey, Ken Kennedy, Jhy-Chun Wang, Daniel A. Reed
SC5
1995 Irregular Personalized Communication on Distributed Memory Machines
abstract
In this paper, we present several algorithms for performing all-to-many personalized communication on distributed memory parallel machines. We assume that each processor sends a different message (of potentially different size) to a subset of all the processors involved in the collective communication. The algorithms are based on decomposing the communication matrix into a set of partial permutations. We study the effectiveness of our algorithms from both the view of static scheduling and runtime scheduling.
Sanjay Ranka, Jhy-Chun Wang
J. Parallel Distributed Comput.2
1994 Scheduling of unstructured communication on the Intel iPSC/860
abstract
We present several algorithms for decomposing all-to-many personalized communication into a set of disjoint partial permutations. These partial permutations avoid node contention as well as link contention. We discuss the theoretical complexity of these algorithms and study their effectiveness both from the view of static scheduling and from runtime scheduling. Experimental results for our algorithms are presented on the iPSC/860.>
Jhy-Chun Wang, Sanjay Ranka
SC1
1994 Static and Run-Time Algorithms for All-to-Many Personalized Communication on Permutation Networks
abstract
With the advent of new routing methods, the distance that a message is sent is becoming relatively less and less important. Thus, assuming no link contention, permutation seems to be an efficient collective communication primitive. In this paper, we present several algorithms for decomposing all-to-many personalized communication into a set of disjoint partial permutations. We discuss several algorithms and study their effectiveness from the view of static scheduling as well as run-time scheduling. An approximate analysis shows that with n processors, and assuming that every processor sends and receives d messages to random destinations, our algorithm can perform the scheduling in O(dn In d) time, on average, and can use an expected number of d+log d partial permutations to carry out the communication. We present experimental results of our algorithms on the CM-5.>
Sanjay Ranka, Jhy-Chun Wang, Geoffrey C. Fox
IEEE Trans. Parallel Distributed Syst.2
1993 Personalized Communication Avoiding Node Contention on Distributed Memory Systems
abstract
In this paper, we present several algorithms for per forming all-to-many personalized communication on distributed memory parallel machines. Each proces sor sends a different message (of potentially different size) to a subset of all the processors involved in the collective communication. The algorithms are based on decomposing the communication matrix into a set of partial permutations. We study the effectiveness of our algorithms both from the view of static scheduling as well as runtime scheduling.
Sanjay Ranka, Jhy-Chun Wang
ICPP (1)2
1993 Embedding Meshes on the Star Graph
Sanjay Ranka, Jhy-Chun Wang, Nangkang Yeh
J. Parallel Distributed Comput.2
1990 Embedding meshes on the star graph
abstract
Algorithms for mapping n-dimensional meshes on a star graph of degree n with expansion 1 and dilation 3 are developed. It is shown that an n-degree star graph can efficiently simulate an n-dimensional mesh. The analysis indicates that the algorithms developed for uniform meshes may not be efficiently simulated on the star graph.>
Sanjay Ranka, Jhy-Chun Wang, Nangkang Yeh
SC2