VLDB 2026 Research / reviewers in the wild / expert
Scott Hahn
dblp:41/3310
· DBLP profile ↗
9ranked-venue papers
0as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Processor architecture and microarchitecture · 33% Electronic design automation · 22% GPUs and heterogeneous computing · 22% | |
| Software engineering, system software, and programming languages
4 papers |
Operating systems · 82% Concurrent programming · 18% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Operating systems › resource management › process management
CPU scheduling |
0.3 | 3 | 2010 | Operating system support for overlapping-ISA heterogeneous multi-core architectures · HPCA 2010 Efficient and scalable multiprocessor fair scheduling using distributed weighted round-robin · PPoPP 2009 Efficient operating system scheduling for performance-asymmetric multi-core architectures · SC 2007 |
Processor architecture and microarchitecture › multicore design › heterogeneous multicore
asymmetric multicore |
0.2 | 2 | 2012 | QuickIA: Exploring heterogeneous architectures on real prototypes · HPCA 2012 Efficient operating system scheduling for performance-asymmetric multi-core architectures · SC 2007 |
Electronic design automation
design space exploration |
0.1 | 1 | 2012 | QuickIA: Exploring heterogeneous architectures on real prototypes · HPCA 2012 |
GPUs and heterogeneous computing
heterogeneous architecture |
0.1 | 1 | 2012 | QuickIA: Exploring heterogeneous architectures on real prototypes · HPCA 2012 |
Processor architecture and microarchitecture › multicore design
heterogeneous cores |
0.1 | 1 | 2012 | The Forgotten 'Uncore': On the Energy-Efficiency of Heterogeneous Cores · USENIX ATC 2012 |
Operating systems › resource management › process management › CPU scheduling
thread scheduling |
0.1 | 1 | 2010 | Bias scheduling in heterogeneous multi-core architectures · EuroSys 2010 |
GPUs and heterogeneous computing › heterogeneous architecture
heterogeneous processors |
0.1 | 1 | 2010 | Operating system support for overlapping-ISA heterogeneous multi-core architectures · HPCA 2010 |
Electronic design automation › high-level synthesis
scheduling |
0.1 | 1 | 2010 | Bias scheduling in heterogeneous multi-core architectures · EuroSys 2010 |
Concurrent programming
fair scheduling |
0.1 | 1 | 2009 | Efficient and scalable multiprocessor fair scheduling using distributed weighted round-robin · PPoPP 2009 |
Embedded and real-time systems › real-time scheduling
multiprocessor scheduling |
0.1 | 1 | 2009 | Efficient and scalable multiprocessor fair scheduling using distributed weighted round-robin · PPoPP 2009 |
Parallel and multicore computing › task scheduling › process scheduling
proportional share scheduling |
0.1 | 1 | 2009 | Efficient and scalable multiprocessor fair scheduling using distributed weighted round-robin · PPoPP 2009 |
Processor architecture and microarchitecture
multicore design |
0.1 | 2 | 2012 | The Forgotten 'Uncore': On the Energy-Efficiency of Heterogeneous Cores · USENIX ATC 2012 Efficient operating system scheduling for performance-asymmetric multi-core architectures · SC 2007 |
Energy-efficient computing
energy-aware scheduling |
0.0 | 1 | 2010 | Bias scheduling in heterogeneous multi-core architectures · EuroSys 2010 |
Operating systems › resource management › process management › CPU scheduling
NUMA-aware scheduling |
0.0 | 1 | 2007 | Efficient operating system scheduling for performance-asymmetric multi-core architectures · SC 2007 |
Methods — techniques the papers use, named apart from their topics
linux kernel implementation · 0.2big.LITTLE-style core asymmetry · 0.2proportional fairness analysis · 0.2distributed thread queues · 0.2prototyping · 0.1load balancing · 0.1case study · 0.1CPU clock modulation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | A protected block device for Persistent MemoryabstractPersistent Memory (PM) technologies, such as Phase Change Memory, STT-RAM, and memristors, are receiving increasingly high interest in academia and industry. PM provides many attractive features, such as DRAM-like speed and storage-like persistence. Yet, because it draws a blurry line between memory and storage, neither a memory- or storage-based model is a natural fit. Best integrating PM into existing systems has become challenging and is now a top priority for many. In this paper we share our initial approach to integrating PM into computer systems, with minimal impact to the core operating system. By adopting a hybrid storage model, all of our changes are confined to a block storage driver, called PMBD, which directly accesses PM attached to the memory bus and exposes a logical block I/O interface to users. We explore the design space by examining a variety of options to achieve performance, protection from stray writes, ordered persistence, and compatibility for legacy file systems and applications. All told, we find that by using a combination of existing OS mechanisms (per-core page table mappings, non-temporal store instructions, memory fences, and I/O barriers), we are able to achieve each of these goals with small performance overhead for both micro-benchmarks and real world applications (e.g., file server and database workloads). Our experience suggests that determining the right combination of existing platform and OS mechanisms is a non-trivial exercise. In this paper, we share both our failed and successful attempts. The final solution that we propose represents an evolution of our initial approach. We have also open-sourced our software prototype with all attempted design options to encourage further research in this area. Feng Chen 0005, Michael P. Mesnier, Scott Hahn |
MSST | 3 |
| 2014 | Client-aware cloud storageabstractCloud storage is receiving high interest in both academia and industry. As a new storage model, it provides many attractive features, such as high availability, resilience, and cost efficiency. Yet, cloud storage also brings many new challenges. In particular, it widens the already-significant semantic gap between applications, which generate data, and storage systems, which manage data. This widening semantic gap makes end-to-end differentiated services extremely difficult. In this paper, we present a client-aware cloud storage framework, which allows semantic information to flow from clients, across multiple intermediate layers, to the cloud storage system. In turn, the storage system can differentiate various data classes and enforce predefined policies. We showcase the effectiveness of enabling such client awareness by using Intel's Differentiated Storage Services (DSS) to enhance persistent disk caching and to control I/O traffic to different storage devices. We find that we can significantly outperform LRU-style caching, improving upload bandwidth by 5x and download bandwidth by 1.6x. Further, we can achieve 85% of the performance of a full-SSD solution at only a fraction (14%) of the cost. Feng Chen 0005, Michael P. Mesnier, Scott Hahn |
MSST | 3 |
| 2012 | QuickIA: Exploring heterogeneous architectures on real prototypesabstractOver the last decade, homogeneous multi-core processors emerged and became the de-facto approach for offering high parallelism, high performance and scalability for a wide range of platforms. We are now at an interesting juncture where several critical factors (smaller form factor devices, power challenges, need for specialization, etc) are guiding architects to consider heterogeneous chips and platforms for the next decade and beyond. Exploring heterogeneous architectures is challenging since it involves re-evaluating architecture options, OS implications and application development. In this paper, we describe these research challenges and then introduce a heterogeneous prototype platform called QuickIA that enables rapid exploration of heterogeneous architectures employing multiple generations of Intel processors for evaluating the implications of asymmetry and FPGAs to experiment with specialized processors or accelerators. We also show example case studies using the QuickIA research prototype to highlight its value in conducting heterogeneous architecture, OS and applications research. Bhushan Chitlur, Ganapati Srinivasa, Scott Hahn, Dheeraj Reddy, David A. Koufaty, Paul Brett, Abirami Prabhakaran, Li Zhao 0002, Nelson Ijih, Suchit Subhaschandra, Sabina Grover, Xiaowei Jiang, Ravi R. Iyer 0001 |
HPCA | 3 |
| 2012 | The Forgotten 'Uncore': On the Energy-Efficiency of Heterogeneous Cores
Vishal Gupta 0001, Paul Brett, David A. Koufaty, Dheeraj Reddy, Scott Hahn, Karsten Schwan, Ganapati Srinivasa |
USENIX ATC | 5 |
| 2010 | Bias scheduling in heterogeneous multi-core architecturesabstractHeterogeneous architectures that integrate a mix of big and small cores are very attractive because they can achieve high single-threaded performance while enabling high performance thread-level parallelism with lower energy costs. Despite their benefits, they pose significant challenges to the operating system software. Thread scheduling is one of the most critical challenges. David A. Koufaty, Dheeraj Reddy, Scott Hahn |
EuroSys | 3 |
| 2010 | Operating system support for overlapping-ISA heterogeneous multi-core architecturesabstractA heterogeneous processor consists of cores that are asymmetric in performance and functionality. Such a design provides a cost-effective solution for processor manufacturers to continuously improve both single-thread performance and multi-thread throughput. This design, however, faces significant challenges in the operating system, which traditionally assumes only homogeneous hardware. This paper presents a comprehensive study of OS support for heterogeneous architectures in which cores have asymmetric performance and overlapping, but non-identical instruction sets. Our algorithms allow applications to transparently execute and fairly share different types of cores. We have implemented these algorithms in the Linux 2.6.24 kernel and evaluated them on an actual heterogeneous platform. Evaluation results demonstrate that our designs efficiently manage heterogeneous hardware and enable significant performance improvements for a range of applications. Tong Li 0003, Paul Brett, Rob C. Knauerhase, David A. Koufaty, Dheeraj Reddy, Scott Hahn |
HPCA | 6 |
| 2009 | Efficient and scalable multiprocessor fair scheduling using distributed weighted round-robinabstractFairness is an essential requirement of any operating system scheduler. Unfortunately, existing fair scheduling algorithms are either inaccurate or inefficient and non-scalable for multiprocessors. This problem is becoming increasingly severe as the hardware industry continues to produce larger scale multi-core processors. This paper presents Distributed Weighted Round-Robin (DWRR), a new scheduling algorithm that solves this problem. With distributed thread queues and small additional overhead to the underlying scheduler, DWRR achieves high efficiency and scalability. Besides conventional priorities, DWRR enables users to specify weights to threads and achieve accurate proportional CPU sharing with constant error bounds. DWRR operates in concert with existing scheduler policies targeting other system attributes, such as latency and throughput. As a result, it provides a practical solution for various production OSes. To demonstrate the versatility of DWRR,we have implemented it in Linux kernels 2.6.22.15 and 2.6.24, which represent two vastly different scheduler designs. Our evaluation shows that DWRR achieves accurate proportional fairness and high performance for a diverse set of workloads. Tong Li 0003, Dan P. Baumberger, Scott Hahn |
PPoPP | 3 |
| 2007 | Soft Real-Time Scheduling on Performance Asymmetric Multicore PlatformsabstractThis paper discusses an approach for supporting soft real-time periodic tasks in Linux on performance asymmetric multicore platforms (AMPs). Such architectures consist of a large number of processing units on one or several chips, where each processing unit is capable of executing the same instruction set at a different performance level. We discuss deficiencies of Linux in supporting periodic real-time tasks, particularly when cores are asymmetric, and how such deficiencies were overcome. We also investigate how to provide good performance for non-real-time tasks in the presence of a real-time workload. We show that this can be done by using deferrable servers to explicitly reserve a share of each core for non-real-time tasks. This allows non-real-time tasks to have priority over real-time tasks when doing so will not cause timing requirements to be violated, thus improving non-real-time response times. Experiments show that even small deferrable servers can have a dramatic impact on non-real-time task performance John M. Calandrino, Dan P. Baumberger, Tong Li 0003, Scott Hahn, James H. Anderson |
IEEE Real-Time and Embedded Technology and Applications Symposium | 4 |
| 2007 | Efficient operating system scheduling for performance-asymmetric multi-core architecturesabstractRecent research advocates asymmetric multi-core architectures, where cores in the same processor can have different performance. These architectures support single-threaded performance and multithreaded throughput at lower costs (e.g., die size and power). However, they also pose unique challenges to operating systems, which traditionally assume homogeneous hardware. This paper presents AMPS, an operating system scheduler that efficiently supports both SMP-and NUMA-style performance-asymmetric architectures. AMPS contains three components: asymmetry-aware load balancing, faster-core-first scheduling, and NUMA-aware migration. We have implemented AMPS in Linux kernel 2.6.16 and used CPU clock modulation to emulate performance asymmetry on an SMP and NUMA system. For various workloads, we show that AMPS achieves a median speedup of 1.16 with a maximum of 1.44 over stock Linux on the SMP, and a median of 1.07 with a maximum of 2.61 on the NUMA system. Our results also show that AMPS improves fairness and repeatability of application performance measurements. Tong Li 0003, Dan P. Baumberger, David A. Koufaty, Scott Hahn |
SC | 4 |