VLDB 2026 Research / reviewers in the wild / expert
John Thomson
dblp:77/6300
· DBLP profile ↗
18ranked-venue papers
0as first author
1since 2021 · last 2021
0000-0002-1905-7539ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 1 since 2021Software engineering, systems software and programming languages · 4Graphics, computer vision, multimedia, augmented reality and games · 2Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Parallel and multicore computing · 40% High-performance computing · 29% Processor architecture and microarchitecture · 19% | |
| Software engineering, system software, and programming languages
2 papers |
Operating systems · 98% Compilers and program optimization · 2% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 87% Machine learning and data management · 13% | |
| Computer graphics and multimedia
1 paper |
Image and video coding · 100% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Operating systems › resource management › process management
CPU scheduling |
0.5 | 1 | 2021 | Collaborative Heterogeneity-Aware OS Scheduler for Asymmetric Multicore Processors · IEEE Trans. Parallel Distributed Syst. 2021 |
Operating systems › resource management › process management › CPU scheduling
multicore scheduling |
0.5 | 1 | 2021 | Collaborative Heterogeneity-Aware OS Scheduler for Asymmetric Multicore Processors · IEEE Trans. Parallel Distributed Syst. 2021 |
Processor architecture and microarchitecture › multicore design › heterogeneous multicore
asymmetric multicore |
0.5 | 1 | 2021 | Collaborative Heterogeneity-Aware OS Scheduler for Asymmetric Multicore Processors · IEEE Trans. Parallel Distributed Syst. 2021 |
Parallel and multicore computing › task scheduling
heterogeneity-aware scheduling |
0.5 | 1 | 2021 | Collaborative Heterogeneity-Aware OS Scheduler for Asymmetric Multicore Processors · IEEE Trans. Parallel Distributed Syst. 2021 |
Data mining
clustering |
0.4 | 1 | 2020 | Large-Scale Automatic K-Means Clustering for Heterogeneous Many-Core Supercomputer · IEEE Trans. Parallel Distributed Syst. 2020 |
Data mining › clustering
k-means clustering |
0.4 | 1 | 2020 | Large-Scale Automatic K-Means Clustering for Heterogeneous Many-Core Supercomputer · IEEE Trans. Parallel Distributed Syst. 2020 |
Parallel and multicore computing
parallel programming models |
0.4 | 1 | 2020 | Large-Scale Automatic K-Means Clustering for Heterogeneous Many-Core Supercomputer · IEEE Trans. Parallel Distributed Syst. 2020 |
High-performance computing
supercomputing |
0.4 | 1 | 2020 | Large-Scale Automatic K-Means Clustering for Heterogeneous Many-Core Supercomputer · IEEE Trans. Parallel Distributed Syst. 2020 |
Image and video coding
image compression |
0.2 | 1 | 2016 | Predicting and Optimizing Image Compression · ACM Multimedia 2016 |
Energy-efficient computing
energy-aware scheduling |
0.1 | 1 | 2021 | Collaborative Heterogeneity-Aware OS Scheduler for Asymmetric Multicore Processors · IEEE Trans. Parallel Distributed Syst. 2021 |
Energy-efficient computing
power management |
0.1 | 1 | 2021 | Collaborative Heterogeneity-Aware OS Scheduler for Asymmetric Multicore Processors · IEEE Trans. Parallel Distributed Syst. 2021 |
Machine learning and data management › automated machine learning
hyperparameter optimization |
0.1 | 1 | 2020 | Large-Scale Automatic K-Means Clustering for Heterogeneous Many-Core Supercomputer · IEEE Trans. Parallel Distributed Syst. 2020 |
Parallel and multicore computing › parallel algorithms › shared-memory parallel algorithms
many-core algorithms |
0.1 | 1 | 2018 | Large-scale hierarchical k-means for heterogeneous many-core supercomputers · SC 2018 |
Mathematical optimization
combinatorial optimization |
0.1 | 1 | 2006 | Predictive search distributions · ICML 2006 |
Mathematical optimization › evolutionary computation
estimation-of-distribution algorithm |
0.1 | 1 | 2006 | Predictive search distributions · ICML 2006 |
Compilers and program optimization
compiler optimization |
0.0 | 1 | 2006 | Predictive search distributions · ICML 2006 |
Methods — techniques the papers use, named apart from their topics
performance and power estimation · 1.0communication pattern detection · 1.0multilevel parallel partition · 0.9automatic hyperparameter determination · 0.9k-means clustering · 0.3hierarchical clustering · 0.3machine learning · 0.2predictive search distribution · 0.2estimation of distribution algorithm · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Collaborative Heterogeneity-Aware OS Scheduler for Asymmetric Multicore ProcessorsabstractAsymmetric multicore processors (AMP) offer multiple types of cores under the same programming interface. Extracting the full potential of AMPs requires intelligent scheduling decisions, matching each thread with the right kind of core, the core that will maximize performance or minimize wasted energy for this thread. Existing OS schedulers are not up to this task. While they may handle certain aspects of asymmetry in the system, none can handle all runtime factors affecting AMPs for the general case of multi-threaded multi-programmed workloads. We address this problem by introducing COLAB, a general purpose asymmetry-aware scheduler targeting multi-threaded multi-programmed workloads. It estimates the performance and power of each thread on each type of core and identifies communication patterns and bottleneck threads. With this information, the scheduler makes coordinated core assignment and thread selection decisions that still provide each application its fair share of the processor's time. We evaluate our approach using both the GEM5 simulator on four distinct big.LITTLE configurations and a development board with ARM Cortex-A73/A53 processors and mixed workloads composed of PARSEC and SPLASH2 benchmarks. Compared to the state-of-the art Linux CFS and AMP-aware schedulers, we demonstrate performance gains of up to 25 and 5 to 15 percent on average, together with an average 5 percent energy saving depending on the hardware setup. Runxin Zhong, Vladimir Janjic, Pavlos Petoumenos, Jidong Zhai, Hugh Leather, John Thomson |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2020 | Modelling VM Latent Characteristics and Predicting Application Performance using Semi-supervised Non-negative Matrix FactorizationabstractSelecting a suitable VM instance type for an application can be a difficult task because of the number of options and the variety of application requirements. Recent research takes a data-driven approach to model VM performance, but this requires carefully choosing a small set of relevant benchmarks as input. We propose a semi-supervised matrix-factorization-based latent variable approach to predict the performance of an unknown new application. This method allows to take a large set of benchmarks as input for VM performance modelling, and it uses the model and the performance measure of the new application on some of the target VMs to predict the performance on the rest of all VMs. We ran experiments with 373 micro-benchmarks from stress-ng and 37 AWS EC2 VMs to predict the scores of Geekbench accurately. Our initial results showed that the RMSE and STD of the predicted scores are 6.7 and 4.5 when sampling Geekbench on 5 VMs, and 10.0 and 2.8 when sampling 10. Yuhui Lin, Adam Barker, John Thomson |
CLOUD | 3 |
| 2020 | COLAB: a collaborative multi-factor scheduler for asymmetric multicore processorsabstractIncreasingly prevalent asymmetric multicore processors (AMP) are necessary for delivering performance in the era of limited power budget and dark silicon. However, the software fails to use them efficiently. OS schedulers, in particular, handle asymmetry only under restricted scenarios. We have efficient symmetric schedulers, efficient asymmetric schedulers for single-threaded workloads, and efficient asymmetric schedulers for single program workloads. What we do not have is a scheduler that can handle all runtime factors affecting AMP for multi-threaded multi-programmed workloads. Pavlos Petoumenos, Vladimir Janjic, Hugh Leather, John Thomson |
CGO | 5 |
| 2020 | Large-Scale Automatic K-Means Clustering for Heterogeneous Many-Core SupercomputerabstractThis article presents an automatic k-means clustering solution targeting the Sunway TaihuLight supercomputer. We first introduce a multilevel parallel partition approach that not only partitions by dataflow and centroid, but also by dimension, which unlocks the potential of the hierarchical parallelism in the heterogeneous many-core processor and the system architecture of the supercomputer. The parallel design is able to process large-scale clustering problems with up to 196,608 dimensions and over 160,000 targeting centroids, while maintaining high performance and high scalability. Furthermore, we propose an automatic hyper-parameter determination process for k-means clustering, by automatically generating and executing the clustering tasks with a set of candidate hyper-parameter, and then determining the optimal hyper-parameter using a proposed evaluation method. The proposed autoclustering solution can not only achieve high performance and scalability for problems with massive high-dimensional data, but also support clustering without sufficient prior knowledge for the number of targeted clusters, which can potentially increase the scope of k-means algorithm to new application areas. Wenlai Zhao, Pan Liu 0002, Vladimir Janjic, Xiaohan Yan, Shicai Wang, Haohuan Fu, Guangwen Yang 0002, John Thomson |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2019 | POSTER: A Collaborative Multi-Factor Scheduler for Asymmetric Multicore ProcessorsabstractAsymmetric multicore processors (AMP) are necessary for extracting performance in an era of limited power budget and dark silicon. We have efficient symmetric schedulers, efficient asymmetric schedulers for single-threaded workloads, and efficient asymmetric schedulers for single program workloads. What we do not have is a scheduler that can handle all three factors affecting AMP scheduling: core affinity, thread criticality, and scheduling fairness. To address this problem, this paper introduces the first general purpose asymmetry-aware scheduler targeting multi-threaded multi-programmed workloads. It estimates the performance of each thread on each type of core and it identifies communication patterns and bottleneck threads. With this information, the scheduler makes coordinated core assignment and thread selection decisions that still provide each application its fair share of the processor's time. We evaluated our approach on GEM5 through four distinct big.LITTLE configurations and multi-threaded multi-programmed workloads composed of PARSEC and SPLASH2 benchmarks. Compared against the Linux CFS scheduler and a state-of-the-art AMP-aware scheduler, we demonstrate performance gains of up to 25% and 5% to 15% on average depending on the hardware setup. Pavlos Petoumenos, Vladimir Janjic, Mingcan Zhu, Hugh Leather, John Thomson |
PACT | 6 |
| 2018 | Lattice-Based Scheduling for Multi-FPGA SystemsabstractAccelerators are becoming increasingly prevalent in distributed computation. FPGAs have been shown to be fast and power efficient for particular tasks, yet scheduling on FPGA-based multi-accelerator systems is challenging when workloads vary significantly in granularity in terms of task size and/or number of computational units required. We present a novel approach for dynamically scheduling tasks on networked multi-FPGA systems which maintains high performance, even in the presence of irregular tasks. Our topological ranking-based scheduling allows realistic irregular workloads to be processed while maintaining a significantly higher level of performance than existing schedulers. Mark Stillwell, Liucheng Guo, Yuchun Ma, John Thomson |
FPT | 6 |
| 2018 | Automated profiling of virtualized media processing functions using telemetry and machine learningabstractMost media streaming services are composed by different virtualized processing functions such as encoding, packaging, encryption, content stitching etc. Deployment of these functions in the cloud is attractive as it enables flexibility in deployment options and resource allocation for the different functions. Yet, most of the time overprovisioning of cloud resources is necessary in order to meet demand variability. This can be costly, especially for large scale deployments. Prior art proposes resource allocation based on analytical models that minimize the costs of cloud deployments under a quality of service (QoS) constraint. However, these models do not sufficiently capture the underlying complexity of services composed of multiple processing functions. Instead, we introduce a novel methodology based on full-stack telemetry and machine learning to profile virtualized or cloud native media processing functions individually. The basis of the approach consists of investigating 4 categories of performance metrics: throughput, anomaly, latency and entropy (TALE) in offline (stress tests) and online setups using cloud telemetry. Machine learning is then used to profile the media processing function in the targeted cloud/NFV environment and to extract the most relevant cloud level Key Performance Indicators (KPIs) that relate to the final perceived quality and known client side performance indicators. The results enable more efficient monitoring, as only KPI related metrics need to be collected, stored and analyzed, reducing the storage and communication footprints by over 85%. In addition a detailed overview of the functions behavior was obtained, enabling optimized initial configuration and deployment, and more fine-grained dynamic online resource allocation reducing overprovisioning and avoiding function collapse. We further highlight the next steps towards cloud native carrier grade virtualized processing functions relevant for future network architectures such as in emerging 5G architectures. Rufael Mekuria, Michael J. McGrath, Vincenzo Riccobene, Victor Bayon-Molino, Christos Tselios, John Thomson, Artem Dobrodub |
MMSys | 6 |
| 2018 | Large-scale hierarchical k-means for heterogeneous many-core supercomputers
Liandeng Li, Wenlai Zhao, Haohuan Fu, Guangwen Yang 0002, John Thomson |
SC | 8 |
| 2017 | ACTiCLOUD: Enabling the Next Generation of Cloud ApplicationsabstractDespite their proliferation as a dominant computing paradigm, cloud computing systems lack effective mechanisms to manage their vast amounts of resources efficiently. Resources are stranded and fragmented, ultimately limiting cloud systems' applicability to large classes of critical applications that pose non-moderate resource demands. Eliminating current technological barriers of actual fluidity and scalability of cloud resources is essential to strengthen cloud computing's role as a critical cornerstone for the digital economy. ACTiCLOUD proposes a novel cloud architecture that breaks the existing scale-up and share-nothing barriers and enables the holistic management of physical resources both at the local cloud site and at distributed levels. Specifically, it makes advancements in the cloud resource management stacks by extending state-of-the-art hypervisor technology beyond the physical server boundary and localized cloud management system to provide a holistic resource management within a rack, within a site, and across distributed cloud sites. On top of this, ACTiCLOUD will adapt and optimize system libraries and runtimes (e.g., JVM) as well as ACTiCLOUD-native applications, which are extremely demanding, and critical classes of applications that currently face severe difficulties in matching their resource requirements to state-of-the-art cloud offerings. Georgios I. Goumas, Konstantinos Nikas, Ewnetu Bayuh Lakew, Christos Kotselidis, Andrew Attwood, Erik Elmroth, Michail Flouris, Nikos Foutris, John Goodacre, Davide Grohmann, Vasileios Karakostas, Panagiotis Koutsourakis, Martin L. Kersten, Mikel Luján, Einar Rustad, John Thomson, Luis Tomás, Atle Vesterkjaer, Jim Webber, Ying Zhang 0027, Nectarios Koziris |
ICDCS | 16 |
| 2016 | EUROSERVER: Share-anything scale-out micro-server design
Manolis Marazakis, John Goodacre, Didier Fuin, Paul M. Carpenter, John Thomson, Emil Matús, Antimo Bruno, Per Stenström, Jérôme Martin, Yves Durand, Isabelle Dor |
DATE | 5 |
| 2016 | Predicting and Optimizing Image CompressionabstractImage compression is a core task for mobile devices, social media and cloud storage backend services. Key evaluation criteria for compression are: the quality of the output, the compression ratio achieved and the computational time (and energy) expended. Predicting the effectiveness of standard compression implementations like libjpeg and WebP on a novel image is challenging, and often leads to non-optimal compression. This paper presents a machine learning-based technique to accurately model the outcome of image compression for arbitrary new images in terms of quality and compression ratio, without requiring significant additional computational time and energy. Using this model, we can actively adapt the aggressiveness of compression on a per image basis to accurately fit user requirements, leading to a more optimal compression. Oleksandr Murashko, John Thomson, Hugh Leather |
ACM Multimedia | 2 |
| 2014 | EUROSERVER: Energy Efficient Node for European Micro-ServersabstractEUROSERVER is a collaborative project that aims to dramatically improve data centre energy-efficiency, cost, and software efficiency. It is addressing these important challenges through the coordinated application of several key recent innovations: 64-bit ARM cores, 3D heterogeneous silicon-on-silicon integration, and fully-depleted silicon-on-insulator (FD SOI) process technology, together with new software techniques for efficient resource management, including resource sharing and workload isolation. We are pioneering a system architecture approach that allows specialized silicon devices to be built even for low-volume markets where NRE costs are currently prohibitive. The EUROSERVER device will embed multiple silicon "chiplets" on an active silicon interposer. Its system architecture is being driven by requirements from three use cases: data centres and cloud computing, telecom infrastructures, and high-end embedded systems. We will build two fully integrated full-system prototypes, based on a common micro-server board, and targeting embedded servers and enterprise servers. Yves Durand, Paul M. Carpenter, Stefano Adami, Angelos Bilas, Denis Dutoit, Alexis Farcy, Georgi Gaydadjiev, John Goodacre, Manolis Katevenis, Manolis Marazakis, Emil Matús, Iakovos Mavroidis, John Thomson |
DSD | 13 |
| 2012 | Preserving Scientific Processes from Design to Publications
Rudolf Mayer, Andreas Rauber, Martin Alexander Neumann, John Thomson, Gonçalo Antunes |
TPDL | 4 |
| 2011 | Automatic OpenCL Device Characterization: Guiding Optimized Kernel Design
Peter Thoman, Klaus Kofler, Heiko Studt, John Thomson, Thomas Fahringer |
Euro-Par (2) | 4 |
| 2010 | Workload characterization supporting the development of domain-specific compiler optimizations using decision trees for data miningabstractEmbedded systems have successfully entered a broad variety of application domains such as automotive and industrial control, telecommunications, networking, digital media, consumer equipment, office automation and many more. In this paper we investigate if there exist any fundamental differences between application domains that justify the development and tuning of domain-specific compilers. We develop an automated approach that is capable of identifying domain-specific workload characterizations and presenting them in a readily interpretable format based on decision trees. The generated workload profiles summarize key resource utilization issues and enable compiler engineers to address the highlighted bottlenecks. We have evaluated our methodology against the industrial EEMBC benchmark suite and three popular embedded processors and have found that workload profiles differ significantly between application domains. We demonstrate that these characteristics can be exploited for the development of domain-specific compiler optimizations. In a case study we show average performance improvements of up to 44% for a class of networking applications. Damon Fenacci, Björn Franke, John Thomson |
SCOPES | 3 |
| 2006 | Using Machine Learning to Focus Iterative OptimizationabstractIterative compiler optimization has been shown to outperform static approaches. This, however, is at the cost of large numbers of evaluations of the program. This paper develops a new methodology to reduce this number and hence speed up iterative optimization. It uses predictive modelling from the domain of machine learning to automatically focus search on those areas likely to give greatest performance. This approach is independent of search algorithm, search space or compiler infrastructure and scales gracefully with the compiler optimization space size. Off-line, a training set of programs is iteratively evaluated and the shape of the spaces and program features are modelled. These models are learnt and used to focus the iterative optimization of a new program. We evaluate two learnt models, an independent and Markov model, and evaluate their worth on two embedded platforms, the Texas Instrument C67I3 and the AMD Au1500. We show that such learnt models can speed up iterative search on large spaces by an order of magnitude. This translates into an average speedup of 1.22 on the TI C6713 and 1.27 on the AMD Au1500 in just 2 evaluations. Felix V. Agakov, Edwin V. Bonilla, John Cavazos, Björn Franke, Grigori Fursin, Michael F. P. O'Boyle, John Thomson, Marc Toussaint, Christopher K. I. Williams |
CGO | 7 |
| 2006 | Predictive search distributionsabstractEstimation of Distribution Algorithms (EDAs) are a popular approach to learn a probability distribution over the "good" solutions to a combinatorial optimization problem. Here we consider the case where there is a collection of such optimization problems with learned distributions, and where each problem can be characterized by some vector of features. Now we can define a machine learning problem to predict the distribution of good solutions q(s|x) for a new problem with features x, where s denotes a solution. This predictive distribution is then used to focus the search. We demonstrate the utility of our method on a compiler optimization task where the goal is to find a sequence of code transformations to make the code run fastest. Results on a set of 12 different benchmarks on two distinct architectures show that our approach consistently leads to significant improvements in performance. Edwin V. Bonilla, Christopher K. I. Williams, Felix V. Agakov, John Cavazos, John Thomson, Michael F. P. O'Boyle |
ICML | 5 |
| 2005 | Probabilistic source-level optimisation of embedded programsabstractEfficient implementation of DSP applications is critical for many embedded systems. Optimising C compilers for embedded processors largely focus on code generation and instruction scheduling which, with their growing maturity, are providing diminishing returns. This paper empirically evaluates another approach, namely source-level transformations and the probabilistic feedback-driven search for “good ” transformation sequences within a large optimisation space. This novel approach combines two selection methods: one based on exploring the optimisation space, the other focused on localised search of good areas. This technique was applied to the UTDSP benchmark suite on two digital signal and multimedia processors (Analog Devices TigerSHARC TS-101, Philips TriMedia TM-1100) and an embedded processor derived from a popular general-purpose processor architecture (Intel Celeron 400). On average, our approach gave a factor of 1.71 times improvement across all platforms equivalent to an average 41 % reduction in execution time, outperforming existing approaches. In certain cases a speedup of up to ≈ 7 was found for individual benchmarks. Björn Franke, Michael F. P. O'Boyle, John Thomson, Grigori Fursin |
LCTES | 3 |