VLDB 2026 Research / reviewers in the wild / expert
Ben Verghese
dblp:68/2333
· DBLP profile ↗
6ranked-venue papers
2as first author
0since 2021 · last 2000
0009-0000-4458-400XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-authorSoftware engineering, systems software and programming languages · 5 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Memory systems · 43% Processor architecture and microarchitecture · 18% Performance modeling and evaluation · 16% | |
| Software engineering, system software, and programming languages
3 papers |
Operating systems · 100% |
Topics — the 19 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
data locality |
0.0 | 2 | 1998 | Flexible Use of Memory for Replication/Migration in Cache-Coherent DSM Multiprocessors · ISCA 1998 Operating System Support for Improving Data Locality on CC-NUMA Compute Servers · ASPLOS 1996 |
Memory systems › memory hierarchy
cache hierarchy |
0.0 | 2 | 2000 | Impact of Chip-Level Integration on Performance of OLTP Workloads · HPCA 2000 Piranha: a scalable architecture based on single-chip multiprocessing · ISCA 2000 |
Memory systems
cache coherence |
0.0 | 3 | 2000 | Flexible Use of Memory for Replication/Migration in Cache-Coherent DSM Multiprocessors · ISCA 1998 Piranha: a scalable architecture based on single-chip multiprocessing · ISCA 2000 Operating System Support for Improving Data Locality on CC-NUMA Compute Servers · ASPLOS 1996 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2000 | Piranha: a scalable architecture based on single-chip multiprocessing · ISCA 2000 |
Performance modeling and evaluation › simulation › architectural simulation
full-system simulation |
0.0 | 1 | 2000 | Impact of Chip-Level Integration on Performance of OLTP Workloads · HPCA 2000 |
Performance modeling and evaluation › workload characterization › commercial workloads
OLTP workload characterization |
0.0 | 1 | 2000 | Impact of Chip-Level Integration on Performance of OLTP Workloads · HPCA 2000 |
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor |
0.0 | 2 | 1998 | Performance Isolation: Sharing and Isolation in Shared-Memory Multiprocessors · ASPLOS 1998 Flexible Use of Memory for Replication/Migration in Cache-Coherent DSM Multiprocessors · ISCA 1998 |
Memory systems › non-uniform memory access
CC-NUMA |
0.0 | 2 | 1998 | Operating System Support for Improving Data Locality on CC-NUMA Compute Servers · ASPLOS 1996 Flexible Use of Memory for Replication/Migration in Cache-Coherent DSM Multiprocessors · ISCA 1998 |
Operating systems › system security › operating system security › protection mechanism › isolation
performance isolation |
0.0 | 1 | 1998 | Performance Isolation: Sharing and Isolation in Shared-Memory Multiprocessors · ASPLOS 1998 |
Operating systems
resource management |
0.0 | 1 | 1998 | Performance Isolation: Sharing and Isolation in Shared-Memory Multiprocessors · ASPLOS 1998 |
Distributed systems
data replication and migration |
0.0 | 1 | 1998 | Flexible Use of Memory for Replication/Migration in Cache-Coherent DSM Multiprocessors · ISCA 1998 |
Distributed systems
resource sharing |
0.0 | 1 | 1998 | Performance Isolation: Sharing and Isolation in Shared-Memory Multiprocessors · ASPLOS 1998 |
Operating systems › resource management
memory management |
0.0 | 1 | 1996 | Operating System Support for Improving Data Locality on CC-NUMA Compute Servers · ASPLOS 1996 |
Operating systems › resource management › process management
CPU scheduling |
0.0 | 1 | 1994 | Scheduling and Page Migration for Multiprocessor Compute Servers · ASPLOS 1994 |
Operating systems › resource management › process management › CPU scheduling
multiprocessor scheduling |
0.0 | 1 | 1994 | Scheduling and Page Migration for Multiprocessor Compute Servers · ASPLOS 1994 |
Memory systems › virtual memory management
page migration |
0.0 | 1 | 1994 | Scheduling and Page Migration for Multiprocessor Compute Servers · ASPLOS 1994 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.0 | 1 | 1998 | Performance Isolation: Sharing and Isolation in Shared-Memory Multiprocessors · ASPLOS 1998 |
Processor architecture and microarchitecture › multiprocessor architecture
cache-coherent multiprocessor |
0.0 | 1 | 1994 | Scheduling and Page Migration for Multiprocessor Compute Servers · ASPLOS 1994 |
Memory systems › shared memory
distributed shared memory |
0.0 | 1 | 1994 | Scheduling and Page Migration for Multiprocessor Compute Servers · ASPLOS 1994 |
Methods — techniques the papers use, named apart from their topics
trace-driven simulation · 0.0cache miss sampling · 0.0simulation · 0.0full-system simulation · 0.0space-sharing · 0.0cluster affinity · 0.0cache affinity · 0.0programmable memory controller · 0.0directory protocol comparison · 0.0timeslicing · 0.0time-slicing · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2000 | Impact of Chip-Level Integration on Performance of OLTP WorkloadsabstractWith increasing chip densities, future microprocessor designs have the opportunity to integrate many of the traditional system-level modules onto the same chip as the processor. Some current designs already integrate extremely large on-chip caches, and there are aggressive next-generation designs that attempt to also integrate the memory controller, coherence hardware, and network router all onto a single chip. The tight coupling of these modules will enable efficient memory systems with substantially better latency and bandwidth characteristics relative to current designs. Among the important application areas for high-performance servers, online transaction processing (OLTP) workloads are likely to benefit most from these trends due to their large instruction and data footprints and high communication miss rates. This paper examines the design trade-offs that arise as more system functionality is integrated onto the processor chip, and identifies a number of important architectural choices that are influenced by chip-level integration. In addition, the paper presents a detailed study of the performance impact of chip-level integration in the context of OLTP workloads. Our results are based on full system simulations of the Oracle commercial database engine running on both in-order and out-of-order issue processors used in uniprocessor and multiprocessor configurations. The results show that chip-level integration can improve the performance of both configurations by about 1.4 to 1.5 times, though for different reasons. For uniprocessors, integration of the L2 cache and the resulting lower hit latency is the primary factor in performance improvement. For multiprocessors, the improvement comes from both the integration of the L2 cache (lower L2 hit latency) and the integration of the other memory system components (better dirty remote latency). Furthermore, we find that the higher associativity afforded by integrating the L2 cache plays a critical role in counteracting the loss of capacity relative to larger off-chip caches. Finally, we find that the relative gains from chip-level integration are virtually identical for in-order and out-of-order processors. Luiz André Barroso, Kourosh Gharachorloo, Andreas Nowatzyk, Ben Verghese |
HPCA | 4 |
| 2000 | Piranha: a scalable architecture based on single-chip multiprocessingabstractThis paper describes the Piranha system, a research prototype being developed at Compaq that aggressively exploits chip multiprocessing by integrating eight simple Alpha processor cores along with a two-level cache hierarchy onto a single chip. Piranha also integrates further on-chip functionality to allow for scalable multiprocessor configurations to be built in a glueless and modular fashion. The use of simple processor cores combined with an industry-standard ASIC design methodology allow us to complete our prototype within a short time-frame, with a team size and investment that are an order of magnitude smaller than that of a commercial microprocessor. Our detailed simulation results show that while each Piranha processor core is substantially slower than an aggressive next-generation processor, the integration of eight cores onto a single chip allows Piranha to outperform next-generation processors by up to 2.9 times (on a per chip basis) on important workloads such as OLTP. This performance advantage can approach a factor of five by using full-custom instead of ASIC logic. In addition to exploiting chip multiprocessing, the Piranha prototype incorporates several other unique design choices including a shared second-level cache with no inclusion, a highly optimized cache coherence protocol, and a novel I/O architecture. Luiz André Barroso, Kourosh Gharachorloo, Robert McNamara, Andreas Nowatzyk, Shaz Qadeer, Barton Sano, Robert Stets, Ben Verghese |
ISCA | 9 |
| 1998 | Performance Isolation: Sharing and Isolation in Shared-Memory MultiprocessorsabstractShared-memory multiprocessors (SMPs) are being extensively used as general-purpose servers. The tight coupling of multiple processors, memory, and I/O provides enormous computing power in a single system, and enables the efficient sharing of these resources.The operating systems for these machines (UNIX or Windows NT) provide very few controls for sharing the resources of the system among the active tasks or users. This unconstrained sharing model is a serious limitation for a server because the load placed by one user can adversely affect other users' performance in an unpredictable manner. We show that this lack of isolation is caused by the resource allocation scheme (or lack thereof) carried over from singleuser workstations. Multi-user multiprocessor systems require more sophisticated resource management, and we show how the proposed "performance isolation" scheme can address the current weaknesses of these systems. We have implemented performance isolation in the Silicon Graphics IRIX operating system for three important system resources: CPU time, memory, and disk bandwidth. Running a number of workloads we show that our proposed scheme is successful at providing workstation-like isolation under heavy load, SMP-like latency under light load, and SMP-like throughput in all cases. Ben Verghese, Anoop Gupta, Mendel Rosenblum |
ASPLOS | 1 |
| 1998 | Flexible Use of Memory for Replication/Migration in Cache-Coherent DSM MultiprocessorsabstractGiven the limitations of bus-based multiprocessors, CC-NUMA is the scalable architecture of choice for shared-memory machines. The most important characteristic of the CC-NUMA architecture is that the latency to access data on a remote node is considerably larger than the latency to access local memory. On such machines, good data locality can reduce memory stall time and is therefore a critical factor in application performance. In this paper we study the various options available to system designers to transparently decrease the fraction of data misses serviced remotely. This work is done in the context of the Stanford FLASH multiprocessor. FLASH is unique in that each node has a single pool of DRAM that can be used in a variety of ways by the programmable memory controller. We use the programmability of FLASH to explore different options for cache-coherence and data-locality in compute-server workloads. First, we consider two protocols for providing base cache-coherence, one with centralized directory information (dynamic pointer allocation) and another with distributed directory information (SCI). While several commercial systems are based on SCI, we find that a centralized scheme has superior performance. Next, we consider different hardware and software techniques that use some or all of the local memory in a node to improve data locality. Finally, we propose a hybrid scheme that combines hardware and software techniques. These schemes work on the same base platform with both user and kernel references from the workloads. The paper thus offers a realistic and fair comparison of replication/migration techniques that has not previously been feasible. Vijayaraghavan Soundararajan, Mark A. Heinrich, Ben Verghese, Kourosh Gharachorloo, Anoop Gupta, John L. Hennessy |
ISCA | 3 |
| 1996 | Operating System Support for Improving Data Locality on CC-NUMA Compute ServersabstractThe dominant architecture for the next generation of shared-memory multiprocessors is CC-NUMA (cache-coherent non-uniform memory architecture). These machines are attractive as compute servers because they provide transparent access to local and remote memory. However, the access latency to remote memory is 3 to 5 times the latency to local memory. CC-NOW machines provide the benefits of cache coherence to networks of workstations, at the cost of even higher remote access latency. Given the large remote access latencies of these architectures, data locality is potentially the most important performance issue. Using realistic workloads, we study the performance improvements provided by OS supported dynamic page migration and replication. Analyzing our kernel-based implementation, we provide a detailed breakdown of the costs. We show that sampling of cache misses can be used to reduce cost without compromising performance, and that TLB misses may not be a consistent approximation for cache misses. Finally, our experiments show that dynamic page migration and replication can substantially increase application performance, as much as 30%, and reduce contention for resources in the NUMA memory system. Ben Verghese, Scott Devine, Anoop Gupta, Mendel Rosenblum |
ASPLOS | 1 |
| 1994 | Scheduling and Page Migration for Multiprocessor Compute ServersabstractSeveral cache-coherent shared-memory multiprocessors have been developed that are scalable and offer a very tight coupling between the processing resources. They are therefore quite attractive for use as compute servers for multiprogramming and parallel application workloads. Process scheduling and memory management, however, remain challenging due to the distributed main memory found on such machines. This paper examines the effects of OS scheduling and page migration policies on the performance of such compute servers. Our experiments are done on the Stanford DASH, a distributed-memory cache-coherent multiprocessor. We show that for our multiprogramming workloads consisting of sequential jobs, the traditional Unix scheduling policy does very poorly. In contrast, a policy incorporating cluster and cache affinity along with a simple page-migration algorithm offers up to two-fold performance improvement. For our workloads consisting of multiple parallel applications, we compare space-sharing policies that divide the processors among the applications to time-slicing policies such as standard Unix or gang scheduling. We show that space-sharing policies can achieve better processor utilization due to the operating point effect, but time-slicing policies benefit strongly from user-level data distribution. Our initial experience with automatic page migration suggests that policies based only on TLB miss information can be quite effective, and useful for addressing the data distribution problems of space-sharing schedulers. Rohit Chandra, Scott Devine, Ben Verghese, Anoop Gupta, Mendel Rosenblum |
ASPLOS | 3 |