EDBT 2026 Demo / reviewers in the wild / expert
Robert Blackmore
dblp:10/2044
· DBLP profile ↗
5ranked-venue papers
0as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
High-performance computing · 56% Parallel and multicore computing · 23% Performance modeling and evaluation · 16% | |
| Software engineering, system software, and programming languages
1 paper |
Operating systems · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › supercomputing
supercomputer deployment |
0.3 | 1 | 2018 | The design, deployment, and evaluation of the CORAL pre-exascale systems · SC 2018 |
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2018 | The design, deployment, and evaluation of the CORAL pre-exascale systems · SC 2018 |
Operating systems › resource management › process management › CPU scheduling
kernel scheduling |
0.0 | 1 | 2003 | Improving the Scalability of Parallel Jobs by adding Parallel Awareness to the Operating System · SC 2003 |
Parallel and multicore computing › parallel scheduling
coscheduling |
0.0 | 1 | 2003 | Improving the Scalability of Parallel Jobs by adding Parallel Awareness to the Operating System · SC 2003 |
Parallel and multicore computing
parallel scheduling |
0.0 | 1 | 2003 | Improving the Scalability of Parallel Jobs by adding Parallel Awareness to the Operating System · SC 2003 |
Interconnection networks and networks-on-chip
interconnection networks |
0.0 | 1 | 2001 | MPI-LAPI: An Efficient Implementation of MPI for IBM RS/6000 SP Systems · IEEE Trans. Parallel Distributed Syst. 2001 |
Parallel and multicore computing › parallel programming models
message passing |
0.0 | 1 | 2001 | MPI-LAPI: An Efficient Implementation of MPI for IBM RS/6000 SP Systems · IEEE Trans. Parallel Distributed Syst. 2001 |
Parallel and multicore computing › parallel programming models › message passing
MPI implementation |
0.0 | 1 | 2001 | MPI-LAPI: An Efficient Implementation of MPI for IBM RS/6000 SP Systems · IEEE Trans. Parallel Distributed Syst. 2001 |
High-performance computing › supercomputing
supercomputing systems |
0.0 | 1 | 2001 | MPI-LAPI: An Efficient Implementation of MPI for IBM RS/6000 SP Systems · IEEE Trans. Parallel Distributed Syst. 2001 |
Methods — techniques the papers use, named apart from their topics
kernel modification · 0.1co-scheduling · 0.1zero-copy communication · 0.0polling and interrupt modes · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | The design, deployment, and evaluation of the CORAL pre-exascale systems
Sudharshan S. Vazhkudai, Bronis R. de Supinski, Arthur S. Bland, Al Geist, James C. Sexton, James A. Kahle, Christopher Zimmer 0001, Scott Atchley, Sarp Oral, Don E. Maxwell, Verónica G. Vergara Larrea, Adam Bertsch, Robin Goldstone, Wayne Joubert, Christopher M. Chambreau, David Appelhans, Robert Blackmore, Ben Casses, George Chochia, Gene Davison, Matthew Ezell, Thomas Gooding, Elsa Gonsiorowski, Leopold Grinberg, Bill Hanson, Bill Hartner, Ian Karlin, Matthew L. Leininger, Dustin Leverman, Chris Marroquin, Adam Moody, Martin Ohmacht, Ramesh Pankajakshan, Fernando Pizzano, James H. Rogers, Bryan S. Rosenburg, Drew Schmidt, Mallikarjun Shankar, Feiyi Wang, Py Watson, Bob Walkup, Lance D. Weems, Junqi Yin |
SC | 17 |
| 2016 | Optimization of Message Passing Services on POWER8 InfiniBand ClustersabstractWe present scalability and performance enhancements to MPI libraries on POWER8 InfiniBand clusters. We explore optimizations in the Parallel Active Messaging Interface (PAMI) libraries. We bypass IB VERBS via low level inline calls resulting in low latencies and high message rates. MPI is enabled on POWER8 by extension of both MPICH and Open MPI to call PAMI libraries. The IBM POWER8 nodes have GPU accelerators to optimize floating throughput of the node. We explore optimized algorithms for GPU-to-GPU communication with minimal processor involvement. We achieve a peak MPI message rate of 186 million messages per second. We also present scalable performance in the QBOX and AMG applications. Sameer Kumar 0001, Robert Blackmore, Sameh Sharkawi, K. A. Nysal Jan, Amith R. Mamidala, T. J. Christopher Ward |
EuroMPI | 2 |
| 2004 | Architecture and Early Performance of the New IBM HPS Fabric and Adapter
Rama Govindaraju, Peter Hochschild, Don G. Grice, Kevin J. Gildea, Robert Blackmore, Carl A. Bender, Chulho Kim, Piyush Chaudhary, Jason Goscinski, Jay Herring, John Houston |
HiPC | 5 |
| 2003 | Improving the Scalability of Parallel Jobs by adding Parallel Awareness to the Operating SystemabstractA parallel application benefits from scheduling policies that include a global perspective of the application's process working set. As the interactions among cooperating processes increase, mechanisms to ameliorate waiting within one or more of the processes become more important. In particular, collective operations such as barriers and reductions are extremely sensitive to even usually harmless events such as context switches among members of the process working set. For the last 18 months, we have been researching the impact of random short-lived interruptions such as timer-decrement processing and periodic daemon activity, and developing strategies to minimize their impact on large processor-count SPMD bulk-synchronous programming styles. We present a novel co-scheduling scheme for improving performance of fine-grain collective activities such as barriers and reductions, describe an implementation consisting of operating system kernel modifications and run-time system, and present a set of empirical results comparing the technique with traditional operating system scheduling. Our results indicate a speedup of over 300% on synchronizing collectives. Terry R. Jones, Shawn Dawson, Rob Neely, William G. Tuel Jr., Larry Brenner, Jeffrey Fier, Robert Blackmore, Patrick Caffrey, Brian Maskell, Paul Tomlinson, Mark Roberts |
SC | 7 |
| 2001 | MPI-LAPI: An Efficient Implementation of MPI for IBM RS/6000 SP SystemsabstractThe IBM RS/6000 SP system is one of the most cost-effective commercially available high performance machines. IBM RS/6000 SP systems support the Message Passing Interface standard (MPI) and LAPI. LAPI is a low level, reliable and efficient one-sided communication API library implemented on IBM RS/6000 SP systems. This paper explains how the high performance of the LAPI library has been exploited in order to implement the MPI standard more efficiently than the existing MPI. It describes how to avoid unnecessary data copies at both the sending and receiving sides for such an implementation. The resolution of problems arising from the mismatches between the requirements of the MPI standard and the features of LAPI is discussed. As a result of this exercise, certain enhancements to LAPI are identified to enable an efficient implementation of MPI on LAPI. The performance of the new implementation of MPI is compared with that of the underlying LAPI itself. The latency (in polling and interrupt modes) and bandwidth of our new implementation is compared with that of the native MPI implementation on RS/6000 SP systems. The results indicate that the MPI implementation on LAPI performs comparably to or better than the original MPI implementation in most cases. Improvements of up to 17.3 percent in polling mode latency, 35.8 percent in interrupt mode latency, and 20.9 percent in bandwidth are obtained for certain message sizes. The implementation of MPI on top of LAPI also outperforms the native MPI implementation for the NAS Parallel Benchmarks. Mohammad Banikazemi, Rama Govindaraju, Robert Blackmore, Dhabaleswar K. Panda 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |