VLDB 2026 Research / reviewers in the wild / expert
Chris Holt
dblp:72/385
· DBLP profile ↗
5ranked-venue papers
1as first author
0since 2021 · last 2003
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-authorTheory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Parallel and multicore computing · 57% Memory systems · 18% High-performance computing · 15% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › multiprocessor system › shared-memory multiprocessor
distributed shared-memory multiprocessor |
0.0 | 1 | 2003 | Latency, Occupancy, and Bandwidth in DSM Multiprocessors: A Performance Evaluation · IEEE Trans. Computers 2003 |
Parallel and multicore computing › parallel computing
parallel application performance |
0.0 | 1 | 2003 | Latency, Occupancy, and Bandwidth in DSM Multiprocessors: A Performance Evaluation · IEEE Trans. Computers 2003 |
Memory systems › cache coherence
cache-coherent shared memory |
0.0 | 1 | 1996 | Application and Architectural Bottlenecks in Large Scale Distributed Shared Memory Machines · ISCA 1996 |
Memory systems › shared memory
distributed shared memory |
0.0 | 1 | 1996 | Application and Architectural Bottlenecks in Large Scale Distributed Shared Memory Machines · ISCA 1996 |
Performance modeling and evaluation › analytical modeling
logp model |
0.0 | 1 | 2003 | Latency, Occupancy, and Bandwidth in DSM Multiprocessors: A Performance Evaluation · IEEE Trans. Computers 2003 |
High-performance computing
n-body simulation |
0.0 | 1 | 1993 | A parallel adaptive fast multipole method · SC 1993 |
Parallel and multicore computing › parallel computing
parallel scientific computing |
0.0 | 1 | 1993 | A parallel adaptive fast multipole method · SC 1993 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 1996 | Application and Architectural Bottlenecks in Large Scale Distributed Shared Memory Machines · ISCA 1996 |
Parallel and multicore computing › load balancing
adaptive load balancing |
0.0 | 1 | 1993 | A parallel adaptive fast multipole method · SC 1993 |
Parallel and multicore computing
load balancing |
0.0 | 1 | 1993 | A parallel adaptive fast multipole method · SC 1993 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.0analytical modeling · 0.0parallelization · 0.0fast multipole method · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2003 | Latency, Occupancy, and Bandwidth in DSM Multiprocessors: A Performance EvaluationabstractWhile the desire to use commodity parts in the communication architecture of a DSM multiprocessor offers advantages in cost and design time, the impact on application performance is unclear. We study this performance impact through detailed simulation, analytical modeling, and experiments on a flexible DSM prototype, using a range of parallel applications. We adapt the logP model to characterize the communication architectures of DSM machines. The l (network latency) and o (controller occupancy) parameters are the keys to performance in these machines, with the g (node-to-network bandwidth) parameter becoming important only for the fastest controllers. We show that, of all the logP parameters, controller occupancy has the greatest impact on application performance. Of the two contributions of occupancy to performance degradation-the latency it adds and the contention it induces-it is the contention component that governs performance regardless of network latency, showing a quadratic dependence on o. As expected, techniques to reduce the impact of latency make controller occupancy a greater bottleneck. Surprisingly, the performance impact of occupancy is substantial, even for highly-tuned applications and even in the absence of latency hiding techniques. Scaling the problem size is often used as a technique to overcome limitations in communication latency and bandwidth. Through experiments on a DSM prototype, we show that there are important classes of applications for which the performance lost by using higher occupancy controllers cannot be regained easily, if at all, by scaling the problem size. Mainak Chaudhuri, Mark A. Heinrich, Chris Holt, Jaswinder Pal Singh, Edward Rothberg, John L. Hennessy |
IEEE Trans. Computers | 3 |
| 1996 | Application and Architectural Bottlenecks in Large Scale Distributed Shared Memory MachinesabstractMany of the programming challenges encountered in small to moderate-scale hardware cache-coherent shared memory machines have been extensively studied. While work remains to be done, the basic techniques needed to efficiently program such machines have been well explored. Recently, a number of researchers have presented architectural techniques for scaling a cache coherent shared address space to much larger processor counts. In this paper, we examine the extent to which applications can achieve reasonable performance on such large-scale, cache-coherent, distributed shared address space machines, by determining the problems sizes needed to achieve a reasonable level of efficiency. We also look at how much programming effort and optimization is needed to achieve high efficiency, beyond that needed at small processor counts. For each application, we discuss the main architectural bottlenecks that prevent smaller problem sizes or less optimized programs from achieving good efficiency. Our results show that while there are some applications that either do not scale or must be heavily optimized to do so, for most of the applications we studied it is not necessary to heavily modify the code or restructure algorithms to scale well upto several hundred processors, once the basic techniques for load balancing and data locality are used that are needed for small-scale systems as well. Programs written with some care perform well without substantially compromising the ease of programming advantage of a shared address space, and the problem sizes required to achieve good performance are surprisingly small. It is important to be careful about how data structures and layouts interact with system granularities, but these optimizations are usually needed for moderate-scale machines as well. Chris Holt, Jaswinder Pal Singh, John L. Hennessy |
ISCA | 1 |
| 1995 | Load Balancing and Data locality in Adaptive Hierarchical N-Body Methods: Barnes-Hut, Fast Multipole, and Rasiosity
Jaswinder Pal Singh, Chris Holt, Takashi Totsuka, Anoop Gupta, John L. Hennessy |
J. Parallel Distributed Comput. | 2 |
| 1994 | Projection in Temporal Logic Programming
Maciej Koutny, Chris Holt |
LPAR | 3 |
| 1993 | A parallel adaptive fast multipole methodabstractWe present parallel versions of a representative N-body application that uses Greengard and Rokhlin's adaptive Fast Multipole Method (FMMJ While parallel implementations of the umform FMM are straightforward and have been developed on alfferent architectures, the aalzptive version complicates the task of obtaining eflective parallel peflormance owing to the nonunz~onn and dynamically changing nature of the problem &mains to which it is applied.We propose and evaluate two techniques for providing load balancing and data locality, both of which take advantageof key insights into the method and its typical applications.Using the better of these techm"ques, we demonstrate 45-fold speedups on galactic sinudations on a 48-processor Stanford DASH machine, a state-of-the-art shared address space multiprocessor even for relatively small problems.We also show good speedups on a 2-ring Kendall Square Research KSR-I.Finally we summarize some key architectural implications of this important computational method.Permission to copy without fee aft or pan of Ibis material is Sranted, provided that the copies am not made or distrit!uted for direct ccmtmerciaf advantage, the ACM copyright ndice and the tiUe of the ~bficstion and 54 its date appear, and notice is given that copyins is by permis$icm of the Association for Com@ing Machinery.To copy dheww, m to repubfish, requires 8 fee andh specific pxmission. Jaswinder Pal Singh, Chris Holt, John L. Hennessy, Anoop Gupta |
SC | 2 |