EDBT 2026 Demo / reviewers in the wild / expert
Choming Wang
dblp:61/1823
· DBLP profile ↗
2ranked-venue papers
0as first author
0since 2021 · last 1999
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
High-performance computing · 37% Performance modeling and evaluation · 33% Interconnection networks and networks-on-chip · 16% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
benchmarking |
0.0 | 2 | 1999 | Resource Scaling Effects on MPP Performance: The STAP Benchmark Implications · IEEE Trans. Parallel Distributed Syst. 1999 Evaluating MPI Collective Communication on the SP2, T3D, and Paragon Multicomputers · HPCA 1997 |
Parallel and multicore computing › parallel architecture
massively parallel processing |
0.0 | 1 | 1999 | Resource Scaling Effects on MPP Performance: The STAP Benchmark Implications · IEEE Trans. Parallel Distributed Syst. 1999 |
Performance modeling and evaluation › benchmarking
parallel benchmark |
0.0 | 1 | 1999 | Resource Scaling Effects on MPP Performance: The STAP Benchmark Implications · IEEE Trans. Parallel Distributed Syst. 1999 |
High-performance computing
space-time adaptive processing |
0.0 | 1 | 1999 | Resource Scaling Effects on MPP Performance: The STAP Benchmark Implications · IEEE Trans. Parallel Distributed Syst. 1999 |
High-performance computing
collective communication |
0.0 | 1 | 1997 | Evaluating MPI Collective Communication on the SP2, T3D, and Paragon Multicomputers · HPCA 1997 |
High-performance computing › collective communication
MPI collective communication |
0.0 | 1 | 1997 | Evaluating MPI Collective Communication on the SP2, T3D, and Paragon Multicomputers · HPCA 1997 |
Interconnection networks and networks-on-chip › interconnection networks
multicomputer network |
0.0 | 1 | 1997 | Evaluating MPI Collective Communication on the SP2, T3D, and Paragon Multicomputers · HPCA 1997 |
Interconnection networks and networks-on-chip
network bandwidth |
0.0 | 1 | 1999 | Resource Scaling Effects on MPP Performance: The STAP Benchmark Implications · IEEE Trans. Parallel Distributed Syst. 1999 |
Methods — techniques the papers use, named apart from their topics
scaling model · 0.0closed-form performance modeling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 1999 | Resource Scaling Effects on MPP Performance: The STAP Benchmark ImplicationsabstractPresently, massively parallel processors (MPPs) are available only in a few commercial models. A sequence of three ASCI Teraflops MPPs has appeared before the new millenium. This paper evaluates six MPP systems through STAP benchmark experiments. The STAP is a radar signal processing benchmark which exploits regularly structured SPMD data parallelism. We reveal the resource scaling effects on MPP performance along orthogonal dimensions of machine size, processor speed, memory capacity messaging latency, and network bandwidth. We show how to achieve balanced resources scaling against enlarged workload (problem size). Among three commercial MPPs, the IBM SP2 shows the highest speed and efficiency, attributed to its well-designed network with middleware support for single system image. The Cray T3D demonstrates a high network bandwidth with a good NUMA memory hierarchy. The Intel Paragon trails far behind due to slow processors used and excessive latency experienced in passing messages. Our analysis projects the lowest STAP speed on the ASCI Red, compared with the projected speed of two ASCI Blue machines. This is attributed to slow processors used in ASCI Red and the mismatch between its hardware and software. The Blue Pacific shows the highest potential to deliver scalable performance up to thousands of nodes. The Blue Mountain is designed to have the highest network bandwidth. Our results suggest a limit on the scalability of the distributed shared-memory (DSM) architecture adopted in Blue Mountain. The scaling model offers a quantitative method to match resource scaling with problem scaling to yield a truly scalable performance. The model helps MPP designers optimize the processors, memory, network, and I/O subsystems of an MPP. For MPP users, the scaling results can be applied to partition a large workload for SPMD execution or to minimize the software overhead in collective communication or remote memory update operations. Finally, our scaling model is assessed to evaluate MPPs with benchmarks other than STAP. Kai Hwang 0001, Choming Wang, Cho-Li Wang |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1997 | Evaluating MPI Collective Communication on the SP2, T3D, and Paragon MulticomputersabstractWe evaluate the architectural support of collective communication operations on the IBM SP2, Cray T3D, and Intel Paragon. The MPI performance data are obtained from the STAP benchmark experiments jointly performed at the USC and HKU. The T3D demonstrated clearly the best timing performance in almost all collective operations. This is attributed to the special hardware built in the T3D for fast messaging and block data transfer. With hardwired barriers, the T3D performs the barrier synchronization in 3 /spl mu/s at least 30 times faster than the SP2 or Paragon. The startup latency of collective operations increases either linearly or logarithmically in three multicomputers. For short messages, the SP2 outperforms the Paragon in the barrier, total exchange, scatter, and gather operations. Various collective operations with 64 KBytes per message over 64 nodes of the three machines can be completed in the time range (5.12 ms, 675 ms). The Paragon outperforms the SP2 in almost all collective operations with long messages. We have derived closed-form expressions to quantify the collective messaging times and aggregated bandwidth on all three machines. For total exchange with 64 nodes, the T3D, Paragon, and SP2 achieved an aggregated bandwidth of 1.745, 0.879, and 0.818 GBytes/s, respectively. These findings are useful to those who wish to predict the MPP performance or to optimize parallel applications by trade-offs between divided computation and collective communication. Kai Hwang 0001, Choming Wang, Cho-Li Wang |
HPCA | 2 |