Ross McIlroy

dblp:38/6971 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 3 first-authorSystems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
4 papers
Runtime systems and virtual machines · 51% Operating systems · 26% Programming languages and type systems · 15%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 62% GPUs and heterogeneous computing · 19% Processor architecture and microarchitecture · 19%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Runtime systems and virtual machines
garbage collection
0.212016
Idle time garbage collection scheduling · PLDI 2016
Operating systems › i/o
asynchronous i/o
0.112011
AC: composable asynchronous IO for native languages · OOPSLA 2011
Programming languages and type systems › concurrent programming languages
language constructs for concurrency
0.112011
AC: composable asynchronous IO for native languages · OOPSLA 2011
Runtime systems and virtual machines › virtual machine implementation
java virtual machine
0.112010
Hera-JVM: a runtime system for heterogeneous multi-core architectures · OOPSLA 2010
Parallel and multicore computing
parallel programming models and runtimes
0.112010
Hera-JVM: a runtime system for heterogeneous multi-core architectures · OOPSLA 2010
Runtime systems and virtual machines
managed runtime
0.112016
Idle time garbage collection scheduling · PLDI 2016
Concurrent programming
message passing
0.012011
AC: composable asynchronous IO for native languages · OOPSLA 2011
Processor architecture and microarchitecture › chip multiprocessor
cell processor
0.012010
Hera-JVM: a runtime system for heterogeneous multi-core architectures · OOPSLA 2010
GPUs and heterogeneous computing › heterogeneous architecture
heterogeneous multicore architecture
0.012010
Hera-JVM: a runtime system for heterogeneous multi-core architectures · OOPSLA 2010

Methods — techniques the papers use, named apart from their topics

frame time discrepancy metric · 0.2concurrent garbage collection · 0.2thread migration · 0.2java memory model · 0.2operational semantics · 0.1affinity metric · 0.1
YearPublicationVenuePosition
2018 Phase-Aware Web Browser Power Management on HMP Platforms
abstract
Over the last years, web browsing has been steadily shifting from desktop computers to mobile devices like smartphones and tablets. However, mobile browsers available today have mainly focused on performance rather than power consumption, although the battery life of a mobile device is one of the most important usability metrics. This is because many of these browsers have originated in the desktop domain and have been ported to the mobile domain. Such browsers have multiple power hungry components such as the rendering engine, and the JavaScript engine, and generate high workload without considering the capabilities and the power consumption characteristics of the underlying hardware platform. Also, the lack of coordination between a browser application and the power manager in the operating system (such as Android) results in poor power savings. In this paper, we propose a power manager that takes into account the internal state of a browser -- that we refer to as a phase -- and show with Google's Chrome running on Android that up to 57.4% more energy can be saved over Android's default power managers. We implemented and evaluated our technique on a heterogeneous multiprocessing (HMP) ARM big.LITTLE platform such as the ones found in most modern smartphones.
Nadja Heitmann, Sangyoung Park, Daniel Clifford, S. Kyostila, Ross McIlroy, Benedikt Meurer, Hannes Payer, Samarjit Chakraborty
ICS5
2016 Idle time garbage collection scheduling
abstract
Efficient garbage collection is increasingly important in today's managed language runtime systems that demand low latency, low memory consumption, and high throughput. Garbage collection may pause the application for many milliseconds to identify live memory, free unused memory, and compact fragmented regions of memory, even when employing concurrent garbage collection. In animation-based applications that require 60 frames per second, these pause times may be observable, degrading user experience. This paper introduces idle time garbage collection scheduling to increase the responsiveness of applications by hiding expensive garbage collection operations inside of small, otherwise unused idle portions of the application's execution, resulting in smoother animations. Additionally we take advantage of idleness to reduce memory consumption while allowing higher memory use when high throughput is required. We implemented idle time garbage collection scheduling in V8, an open-source, production JavaScript virtual machine running within Chrome. We present performance results on various benchmarks running popular webpages and show that idle time garbage collection scheduling can significantly improve latency and memory consumption. Furthermore, we introduce a new metric called frame time discrepancy to quantify the quality of the user experience and precisely measure the improvements that idle time garbage collection provides for a WebGL-based game benchmark. Idle time garbage collection is shipped and enabled by default in Chrome.
Ulan Degenbaev, Jochen Eisinger, Manfred Ernst, Ross McIlroy, Hannes Payer
PLDI4
2011 AC: composable asynchronous IO for native languages
abstract
This paper introduces AC, a set of language constructs for composable asynchronous IO in native languages such as C/C++. Unlike traditional synchronous IO interfaces, AC lets a thread issue multiple IO requests so that they can be serviced concurrently, and so that long-latency operations can be overlapped with computation. Unlike traditional asynchronous IO interfaces, AC retains a sequential style of programming without requiring code to use multiple threads, and without requiring code to be "stack-ripped" into chains of callbacks. AC provides an "async" statement to identify opportunities for IO operations to be issued concurrently, a "do..finish" block that waits until any enclosed "async" work is complete, and a "cancel" statement that requests cancellation of unfinished IO within an enclosing "do..finish". We give an operational semantics for a core language. We describe and evaluate implementations that are integrated with message passing on the Barrelfish research OS, and integrated with asynchronous file and network IO on Microsoft Windows. We show that AC offers comparable performance to existing C/C++ interfaces for asynchronous IO, while providing a simpler programming model.
Tim Harris 0001, Martín Abadi, Rebecca Isaacs, Ross McIlroy
OOPSLA4
2010 Hera-JVM: a runtime system for heterogeneous multi-core architectures
abstract
Heterogeneous multi-core processors, such as the IBM Cell processor, can deliver high performance. However, these processors are notoriously difficult to program: different cores support different instruction set architectures, and the processor as a whole does not provide coherence between the different cores’ local memories. We present Hera-JVM, an implementation of the Java Virtual Machine which operates over the Cell processor, thereby making this platforms more readily accessible to mainstream developers. Hera-JVM supports the full Java language; threads from an unmodified Java application can be simultaneously executed on both the main PowerPC-based core and on the additional SPE accelerator cores. Migration of threads between these cores is transparent from the point of view of the application, requiring no modification to Java source code or bytecode. Hera-JVM supports the existing Java Memory Model, even though the underlying hardware does not provide cache coherence between the different core types. We examine Hera-JVM’s performance under a series of real-world Java benchmarks from the SpecJVM, Java Grande and Dacapo benchmark suites. These benchmarks show a wide variation in relative performance on the different core types of the Cell processor, depending upon the nature of their workload. Execution of these benchmarks on Hera-JVM can achieve speedups of up to 2.25x by using one of the Cell processor’s SPE accelerator cores, compared to execution on the main PowerPC-based core. When all six SPE cores are exploited, parallel workloads can achieve speedups of up to 13x compared to execution on the single PowerPC core.
Ross McIlroy, Joseph S. Sventek
OOPSLA1
2009 Hera-JVM: Abstracting Processor Heterogeneity Behind a Virtual Machine
Ross McIlroy, Joseph S. Sventek
HotOS1
2009 Helios: heterogeneous multiprocessing with satellite kernels
abstract
Helios is an operating system designed to simplify the task of writing, deploying, and tuning applications for heterogeneous platforms. Helios introduces satellite kernels, which export a single, uniform set of OS abstractions across CPUs of disparate architectures and performance characteristics. Access to I/O services such as file systems are made transparent via remote message passing, which extends a standard microkernel message-passing abstraction to a satellite kernel infrastructure. Helios retargets applications to available ISAs by compiling from an intermediate language. To simplify deploying and tuning application performance, Helios exposes an affinity metric to developers. Affinity provides a hint to the operating system about whether a process would benefit from executing on the same platform as a service it depends upon.
Ed Nightingale, Orion Hodson, Ross McIlroy, Chris Hawblitzel, Galen C. Hunt
SOSP3
2008 Efficient dynamic heap allocation of scratch-pad memory
abstract
An increasing number of processor architectures support scratch-pad memory - software managed on-chip memory. Scratch-pad memory provides low latency data storage, like on-chip caches, but under explicit software control. The simple design and predictable nature of scratchpad memories has seen them incorporated into a number of embedded and real-time system processors. They are also employed by multi-core architectures to isolate processor core local data and act as low latency inter-core shared memory.
Ross McIlroy, Peter Dickman, Joseph S. Sventek
ISMM1