VLDB 2026 Research / reviewers in the wild / expert
Sahan Gamage
dblp:18/5247
· DBLP profile ↗
10ranked-venue papers
3as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Cloud and datacenter computing · 97% Processor architecture and microarchitecture · 3% | |
| Computer networks
2 papers |
Transport protocols and congestion control · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Programming languages and type systems · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing
virtualization |
0.6 | 4 | 2013 | Protocol Responsibility Offloading to Improve TCP Throughput in Virtualized Environments · ACM Trans. Comput. Syst. 2013 vTurbo: Accelerating Virtual Machine I/O Processing Using Designated Turbo-Sliced Core · USENIX ATC 2013 vSlicer: latency-aware virtual machine scheduling via differentiated-frequency CPU slicing · HPDC 2012 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.6 | 2 | 2020 | Building Scalable and Flexible Cluster Managers Using Declarative Programming · OSDI 2020 vSlicer: latency-aware virtual machine scheduling via differentiated-frequency CPU slicing · HPDC 2012 |
Programming languages and type systems › programming paradigms
declarative programming |
0.4 | 1 | 2020 | Building Scalable and Flexible Cluster Managers Using Declarative Programming · OSDI 2020 |
Transport protocols and congestion control › transport protocol implementation
congestion control offloading |
0.2 | 1 | 2013 | Protocol Responsibility Offloading to Improve TCP Throughput in Virtualized Environments · ACM Trans. Comput. Syst. 2013 |
Transport protocols and congestion control
TCP congestion control |
0.2 | 1 | 2013 | Protocol Responsibility Offloading to Improve TCP Throughput in Virtualized Environments · ACM Trans. Comput. Syst. 2013 |
Cloud and datacenter computing › virtualization
virtual machine networking |
0.2 | 1 | 2013 | Protocol Responsibility Offloading to Improve TCP Throughput in Virtualized Environments · ACM Trans. Comput. Syst. 2013 |
Cloud and datacenter computing › virtualization › virtual machine management
virtual machine scheduling |
0.1 | 1 | 2012 | vSlicer: latency-aware virtual machine scheduling via differentiated-frequency CPU slicing · HPDC 2012 |
Transport protocols and congestion control
TCP |
0.1 | 1 | 2010 | vSnoop: Improving TCP Throughput in Virtualized Environments via Acknowledgement Offload · SC 2010 |
Cloud and datacenter computing › virtualization › virtual machine management
virtual machine consolidation |
0.1 | 2 | 2013 | Protocol Responsibility Offloading to Improve TCP Throughput in Virtualized Environments · ACM Trans. Comput. Syst. 2013 vSnoop: Improving TCP Throughput in Virtualized Environments via Acknowledgement Offload · SC 2010 |
Processor architecture and microarchitecture › chip multiprocessor
core scheduling |
0.0 | 1 | 2013 | vTurbo: Accelerating Virtual Machine I/O Processing Using Designated Turbo-Sliced Core · USENIX ATC 2013 |
Methods — techniques the papers use, named apart from their topics
declarative programming · 0.9acknowledgement offload · 0.2turbo-sliced core · 0.2micro time slice scheduling · 0.1CPU slicing · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 3D IGZO Charge-Coupled Memory DTCO & STCO Analysis for Compute-near-Memory ApplicationsabstractThe demand for high-capacity and energy-efficient memory solutions has surged in the era of data-centric computing, particularly for Artificial Intelligence (AI) and Machine Learning (ML) workloads. This paper introduces a novel memory architecture leveraging Charge-Coupled Device (CCD) technology, engineered in a sequential-access block memory configuration, to enhance Compute-near-Memory (CnM) systems. We propose an optimized 3D IGZO CCD block memory as an on-chip weight buffer for high-capacity CnM systems. Our approach achieves 2.95−131.26× improvement in area efficiency and 1.32−4.33× improvement in energy efficiency compared to SRAM solutions. Khakim Akhunov, Hyungrock Oh, Fernando García-Redondo, Yukai Chen, Arvind Sharma, Jiacong Sun, Sahan Gamage, Maarten Rosmeulen, Swaraj Bandhu Mahato, Rishabh Kishore, Subhali Subhechha, Jaydeep P. Kulkarni, Marian Verhelst, Dwaipayan Biswas, Marie Garcia Bardon, Wim Dehaene, Julien Ryckaert |
ISCAS | 8 |
| 2020 | Building Scalable and Flexible Cluster Managers Using Declarative Programming
Lalith Suresh 0001, João Loff, Faria Kalim, Sangeetha Abdu Jyothi, Nina Narodytska, Leonid Ryzhyk, Sahan Gamage, Brian Oki, Pranshu Jain, Michael Gasch |
OSDI | 7 |
| 2015 | vRead: Efficient Data Access for Hadoop in Virtualized CloudsabstractWith its unlimited scalability and on-demand access to computation and storage, a virtualized cloud platform is the perfect match for big data systems such as Hadoop. However, virtualization introduces a significant amount of overhead to I/O intensive applications due to device virtualization and VMs or I/O threads scheduling delay. In particular, device virtualization causes significant CPU overhead as I/O data needs to be moved across several protection boundaries. We observe that such overhead especially affects the I/O performance of the Hadoop distributed file system (HDFS). In fact, data read from an HDFS datanode VM must go through virtual devices multiple times --- incurring non-negligible virtualization overhead --- even though both client VM and datanode VM may be running on the same machine. In this paper, we propose vRead, a programmable framework which connects I/O flows from HDFS applications directly to their data. vRead enables direct "reads" to the disk images of datanode VMs from the hypervisor. By doing so, vRead can significantly avoid device virtualization overhead, resulting in improved I/O throughput as well as CPU savings for Hadoop workloads and other applications relying on HDFS. Cong Xu 0010, Brendan Saltaformaggio, Sahan Gamage, Ramana Rao Kompella, Dongyan Xu |
Middleware | 3 |
| 2014 | vPipe: Piped I/O Offloading for Efficient Data Movement in Virtualized CloudsabstractVirtualization introduces a significant amount of overhead for I/O intensive applications running inside virtual machines (VMs). Such overhead is caused by two main sources: (1) device virtualization and (2) VM scheduling. Device virtualization causes significant CPU overhead as I/O data need to be moved across several protection boundaries. VM scheduling introduces delays to the overall I/O processing path due to the wait time of VMs' virtual CPUs in the run queue. We observe that such overhead particularly affects many applications involving piped I/O data movements, such as web servers, streaming servers, big data analytics, and storage, because the data has to be transferred first into the application from the source I/O device and then back to the sink I/O device, incurring the virtualization overhead twice. In this paper, we propose vPipe, a programmable framework to mitigate this problem for a wide range of applications running in virtualized clouds. vPipe enables direct "piping" of application I/O data from source to sink devices, either files or TCP sockets, at virtual machine monitor (VMM) level. By doing so, vPipe can avoid both device virtualization overhead and VM scheduling delays, resulting in improved I/O throughput and application performance as well as significant CPU savings. Sahan Gamage, Cong Xu 0010, Ramana Rao Kompella, Dongyan Xu |
SoCC | 1 |
| 2014 | CLUE: System trace analytics for cloud service performance diagnosisabstractIn this paper, we present CLUE, a system event analytics tool for black-box performance diagnosis in production Cloud Computing systems. CLUE provides an unified and extensible means of profiling service transactional behaviors, and builds structured data called event sketches. CLUE further offers a set of analytic tools for summarizing and analyzing event sketches by integrating data mining and statistical analysis. CLUE has been developed in NEC as an internal tool and applied in diagnosing a diverse set of real performance problems for multi-tiered IT applications running on multi-core servers of major platforms including Linux (Redhat, Fedora), Unix (HP-UX), and Windows (Windows Server 2008). We demonstrated the evaluation of our framework on real-world IT systems, and showed how it can enable visibility and effective diagnosis of service system performance problems. Hui Zhang 0002, Junghwan Rhee, Nipun Arora, Sahan Gamage, Guofei Jiang, Kenji Yoshihira, Dongyan Xu |
NOMS | 4 |
| 2013 | vTurbo: Accelerating Virtual Machine I/O Processing Using Designated Turbo-Sliced Core
Cong Xu 0010, Sahan Gamage, Hui Lu 0001, Ramana Rao Kompella, Dongyan Xu |
USENIX ATC | 2 |
| 2013 | Protocol Responsibility Offloading to Improve TCP Throughput in Virtualized EnvironmentsabstractVirtualization is a key technology that powers cloud computing platforms such as Amazon EC2. Virtual machine (VM) consolidation, where multiple VMs share a physical host, has seen rapid adoption in practice, with increasingly large numbers of VMs per machine and per CPU core. Our investigations, however, suggest that the increasing degree of VM consolidation has serious negative effects on the VMs’ TCP performance. As multiple VMs share a given CPU, the scheduling latencies, which can be in the order of tens of milliseconds, substantially increase the typically submillisecond round-trip times (RTTs) for TCP connections in a datacenter, causing significant degradation in throughput. In this article, we propose a lightweight solution, called vPRO, that (a) offloads the VM’s TCP congestion control function to the driver domain to improve TCP transmit performance; and (b) offloads TCP acknowledgment functionality to the driver domain to improve the TCP receive performance. Our evaluation of a vPRO prototype on Xen suggests that vPRO substantially improves TCP receive and transmit throughputs with minimal per-packet CPU overhead. We further show that the higher TCP throughput leads to improvement in application-level performance, via experiments with Apache Olio, a Web 2.0 cloud application, and Intel MPI benchmark. Sahan Gamage, Ramana Rao Kompella, Dongyan Xu, Ardalan Kangarlou |
ACM Trans. Comput. Syst. | 1 |
| 2012 | vSlicer: latency-aware virtual machine scheduling via differentiated-frequency CPU slicingabstractRecent advances in virtualization technologies have made it feasible to host multiple virtual machines (VMs) in the same physical host and even the same CPU core, with fair share of the physical resources among the VMs. However, as more VMs share the same core/CPU, the CPU access latency experienced by each VM increases substantially, which translates into longer I/O processing latency perceived by I/O-bound applications. To mitigate such impact while retaining the benefit of CPU sharing, we introduce a new class of VMs called latency-sensitive VMs (LSVMs), which achieve better performance for I/O-bound applications while maintaining the same resource share (and thus cost) as other CPU-sharing VMs. LSVMs are enabled by vSlicer, a hypervisor-level technique that schedules each LSVM more frequently but with a smaller micro time slice. vSlicer enables more timely processing of I/O events by LSVMs, without violating the CPU share fairness among all sharing VMs. Our evaluation of a vSlicer prototype in Xen shows that vSlicer substantially reduces network packet round-trip times and jitter and improves application-level performance. For example, vSlicer doubles both the connection rate and request processing throughput of an Apache web server; reduces a VoIP server's upstream jitter by 62%; and shortens the execution times of Intel MPI benchmark programs by half or more. Cong Xu 0010, Sahan Gamage, Pawan N. Rao, Ardalan Kangarlou, Ramana Rao Kompella, Dongyan Xu |
HPDC | 2 |
| 2011 | Opportunistic flooding to improve TCP transmit performance in virtualized cloudsabstractVirtualization is a key technology that powers cloud computing platforms such as Amazon EC2. Virtual machine (VM) consolidation, where multiple VMs share a physical host, has seen rapid adoption in practice with increasingly large number of VMs per machine and per CPU core. Our investigations, however, suggest that the increasing degree of VM consolidation has serious negative effects on the VMs' TCP transport performance. As multiple VMs share a given CPU, the scheduling latencies, which can be in the order of tens of milliseconds, substantially increase the typically sub-millisecond round-trip times (RTTs) for TCP connections in a datacenter, causing significant degradation in throughput. In this paper, we propose a light-weight solution called vFlood that (a) allows a TCP sender VM to opportunistically flood the driver domain in the same host, and (b) offloads the VM's TCP congestion control function to the driver domain in order to mask the effects of VM consolidation. Our evaluation of a vFlood prototype on Xen suggests that vFlood substantially improves TCP transmit throughput with minimal per-packet CPU overhead. Further, our application-level evaluation using Apache Olio, a web 2.0 cloud application, indicates a 33% improvement in the number of operations per second. Sahan Gamage, Ardalan Kangarlou, Ramana Rao Kompella, Dongyan Xu |
SoCC | 1 |
| 2010 | vSnoop: Improving TCP Throughput in Virtualized Environments via Acknowledgement OffloadabstractVirtual machine (VM) consolidation has become a common practice in clouds, Grids, and datacenters. While this practice leads to higher CPU utilization, we observe its negative impact on the TCP throughput of the consolidated VMs: As more VMs share the same core/CPU, the CPU scheduling latency for each VM increases significantly. Such increase leads to slower progress of TCP transmissions to the VMs. To address this problem, we propose an approach called vSnoop, where the driver domain of a host acknowledges TCP packets on behalf of the guest VMs - whenever it is safe to do so. Our evaluation of a Xen-based prototype indicates that vSnoop constantly achieves TCP throughput improvement for VMs (of orders of magnitude in some scenarios). We further show that the higher TCP throughput leads to improvement in application- level performance, via experiments with a two-tier online auction application and two suites of MPI benchmarks. Ardalan Kangarlou, Sahan Gamage, Ramana Rao Kompella, Dongyan Xu |
SC | 2 |