John Thomson

dblp:77/6300 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
1since 2021 · last 2021
0000-0002-1905-7539ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 since 2021Software engineering, systems software and programming languages · 4Graphics, computer vision, multimedia, augmented reality and games · 2Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 40% High-performance computing · 29% Processor architecture and microarchitecture · 19%
Software engineering, system software, and programming languages
2 papers
Operating systems · 98% Compilers and program optimization · 2%
Databases, data mining, and information retrieval
1 paper
Data mining · 87% Machine learning and data management · 13%
Computer graphics and multimedia
1 paper
Image and video coding · 100%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Operating systems › resource management › process management
CPU scheduling
0.512021
Collaborative Heterogeneity-Aware OS Scheduler for Asymmetric Multicore Processors · IEEE Trans. Parallel Distributed Syst. 2021
Operating systems › resource management › process management › CPU scheduling
multicore scheduling
0.512021
Collaborative Heterogeneity-Aware OS Scheduler for Asymmetric Multicore Processors · IEEE Trans. Parallel Distributed Syst. 2021
Processor architecture and microarchitecture › multicore design › heterogeneous multicore
asymmetric multicore
0.512021
Collaborative Heterogeneity-Aware OS Scheduler for Asymmetric Multicore Processors · IEEE Trans. Parallel Distributed Syst. 2021
Parallel and multicore computing › task scheduling
heterogeneity-aware scheduling
0.512021
Collaborative Heterogeneity-Aware OS Scheduler for Asymmetric Multicore Processors · IEEE Trans. Parallel Distributed Syst. 2021
Data mining
clustering
0.412020
Large-Scale Automatic K-Means Clustering for Heterogeneous Many-Core Supercomputer · IEEE Trans. Parallel Distributed Syst. 2020
Data mining › clustering
k-means clustering
0.412020
Large-Scale Automatic K-Means Clustering for Heterogeneous Many-Core Supercomputer · IEEE Trans. Parallel Distributed Syst. 2020
Parallel and multicore computing
parallel programming models
0.412020
Large-Scale Automatic K-Means Clustering for Heterogeneous Many-Core Supercomputer · IEEE Trans. Parallel Distributed Syst. 2020
High-performance computing
supercomputing
0.412020
Large-Scale Automatic K-Means Clustering for Heterogeneous Many-Core Supercomputer · IEEE Trans. Parallel Distributed Syst. 2020
Image and video coding
image compression
0.212016
Predicting and Optimizing Image Compression · ACM Multimedia 2016
Energy-efficient computing
energy-aware scheduling
0.112021
Collaborative Heterogeneity-Aware OS Scheduler for Asymmetric Multicore Processors · IEEE Trans. Parallel Distributed Syst. 2021
Energy-efficient computing
power management
0.112021
Collaborative Heterogeneity-Aware OS Scheduler for Asymmetric Multicore Processors · IEEE Trans. Parallel Distributed Syst. 2021
Machine learning and data management › automated machine learning
hyperparameter optimization
0.112020
Large-Scale Automatic K-Means Clustering for Heterogeneous Many-Core Supercomputer · IEEE Trans. Parallel Distributed Syst. 2020
Parallel and multicore computing › parallel algorithms › shared-memory parallel algorithms
many-core algorithms
0.112018
Large-scale hierarchical k-means for heterogeneous many-core supercomputers · SC 2018
Mathematical optimization
combinatorial optimization
0.112006
Predictive search distributions · ICML 2006
Mathematical optimization › evolutionary computation
estimation-of-distribution algorithm
0.112006
Predictive search distributions · ICML 2006
Compilers and program optimization
compiler optimization
0.012006
Predictive search distributions · ICML 2006

Methods — techniques the papers use, named apart from their topics

performance and power estimation · 1.0communication pattern detection · 1.0multilevel parallel partition · 0.9automatic hyperparameter determination · 0.9k-means clustering · 0.3hierarchical clustering · 0.3machine learning · 0.2predictive search distribution · 0.2estimation of distribution algorithm · 0.2
YearPublicationVenuePosition
2021 Collaborative Heterogeneity-Aware OS Scheduler for Asymmetric Multicore Processors
abstract
Asymmetric multicore processors (AMP) offer multiple types of cores under the same programming interface. Extracting the full potential of AMPs requires intelligent scheduling decisions, matching each thread with the right kind of core, the core that will maximize performance or minimize wasted energy for this thread. Existing OS schedulers are not up to this task. While they may handle certain aspects of asymmetry in the system, none can handle all runtime factors affecting AMPs for the general case of multi-threaded multi-programmed workloads. We address this problem by introducing COLAB, a general purpose asymmetry-aware scheduler targeting multi-threaded multi-programmed workloads. It estimates the performance and power of each thread on each type of core and identifies communication patterns and bottleneck threads. With this information, the scheduler makes coordinated core assignment and thread selection decisions that still provide each application its fair share of the processor's time. We evaluate our approach using both the GEM5 simulator on four distinct big.LITTLE configurations and a development board with ARM Cortex-A73/A53 processors and mixed workloads composed of PARSEC and SPLASH2 benchmarks. Compared to the state-of-the art Linux CFS and AMP-aware schedulers, we demonstrate performance gains of up to 25 and 5 to 15 percent on average, together with an average 5 percent energy saving depending on the hardware setup.
Runxin Zhong, Vladimir Janjic, Pavlos Petoumenos, Jidong Zhai, Hugh Leather, John Thomson
IEEE Trans. Parallel Distributed Syst.7
2020 Modelling VM Latent Characteristics and Predicting Application Performance using Semi-supervised Non-negative Matrix Factorization
abstract
Selecting a suitable VM instance type for an application can be a difficult task because of the number of options and the variety of application requirements. Recent research takes a data-driven approach to model VM performance, but this requires carefully choosing a small set of relevant benchmarks as input. We propose a semi-supervised matrix-factorization-based latent variable approach to predict the performance of an unknown new application. This method allows to take a large set of benchmarks as input for VM performance modelling, and it uses the model and the performance measure of the new application on some of the target VMs to predict the performance on the rest of all VMs. We ran experiments with 373 micro-benchmarks from stress-ng and 37 AWS EC2 VMs to predict the scores of Geekbench accurately. Our initial results showed that the RMSE and STD of the predicted scores are 6.7 and 4.5 when sampling Geekbench on 5 VMs, and 10.0 and 2.8 when sampling 10.
Yuhui Lin, Adam Barker, John Thomson
CLOUD3
2020 COLAB: a collaborative multi-factor scheduler for asymmetric multicore processors
abstract
Increasingly prevalent asymmetric multicore processors (AMP) are necessary for delivering performance in the era of limited power budget and dark silicon. However, the software fails to use them efficiently. OS schedulers, in particular, handle asymmetry only under restricted scenarios. We have efficient symmetric schedulers, efficient asymmetric schedulers for single-threaded workloads, and efficient asymmetric schedulers for single program workloads. What we do not have is a scheduler that can handle all runtime factors affecting AMP for multi-threaded multi-programmed workloads.
Pavlos Petoumenos, Vladimir Janjic, Hugh Leather, John Thomson
CGO5
2020 Large-Scale Automatic K-Means Clustering for Heterogeneous Many-Core Supercomputer
abstract
This article presents an automatic k-means clustering solution targeting the Sunway TaihuLight supercomputer. We first introduce a multilevel parallel partition approach that not only partitions by dataflow and centroid, but also by dimension, which unlocks the potential of the hierarchical parallelism in the heterogeneous many-core processor and the system architecture of the supercomputer. The parallel design is able to process large-scale clustering problems with up to 196,608 dimensions and over 160,000 targeting centroids, while maintaining high performance and high scalability. Furthermore, we propose an automatic hyper-parameter determination process for k-means clustering, by automatically generating and executing the clustering tasks with a set of candidate hyper-parameter, and then determining the optimal hyper-parameter using a proposed evaluation method. The proposed autoclustering solution can not only achieve high performance and scalability for problems with massive high-dimensional data, but also support clustering without sufficient prior knowledge for the number of targeted clusters, which can potentially increase the scope of k-means algorithm to new application areas.
Wenlai Zhao, Pan Liu 0002, Vladimir Janjic, Xiaohan Yan, Shicai Wang, Haohuan Fu, Guangwen Yang 0002, John Thomson
IEEE Trans. Parallel Distributed Syst.9
2019 POSTER: A Collaborative Multi-Factor Scheduler for Asymmetric Multicore Processors
abstract
Asymmetric multicore processors (AMP) are necessary for extracting performance in an era of limited power budget and dark silicon. We have efficient symmetric schedulers, efficient asymmetric schedulers for single-threaded workloads, and efficient asymmetric schedulers for single program workloads. What we do not have is a scheduler that can handle all three factors affecting AMP scheduling: core affinity, thread criticality, and scheduling fairness. To address this problem, this paper introduces the first general purpose asymmetry-aware scheduler targeting multi-threaded multi-programmed workloads. It estimates the performance of each thread on each type of core and it identifies communication patterns and bottleneck threads. With this information, the scheduler makes coordinated core assignment and thread selection decisions that still provide each application its fair share of the processor's time. We evaluated our approach on GEM5 through four distinct big.LITTLE configurations and multi-threaded multi-programmed workloads composed of PARSEC and SPLASH2 benchmarks. Compared against the Linux CFS scheduler and a state-of-the-art AMP-aware scheduler, we demonstrate performance gains of up to 25% and 5% to 15% on average depending on the hardware setup.
Pavlos Petoumenos, Vladimir Janjic, Mingcan Zhu, Hugh Leather, John Thomson
PACT6
2018 Lattice-Based Scheduling for Multi-FPGA Systems
abstract
Accelerators are becoming increasingly prevalent in distributed computation. FPGAs have been shown to be fast and power efficient for particular tasks, yet scheduling on FPGA-based multi-accelerator systems is challenging when workloads vary significantly in granularity in terms of task size and/or number of computational units required. We present a novel approach for dynamically scheduling tasks on networked multi-FPGA systems which maintains high performance, even in the presence of irregular tasks. Our topological ranking-based scheduling allows realistic irregular workloads to be processed while maintaining a significantly higher level of performance than existing schedulers.
Mark Stillwell, Liucheng Guo, Yuchun Ma, John Thomson
FPT6
2018 Automated profiling of virtualized media processing functions using telemetry and machine learning
abstract
Most media streaming services are composed by different virtualized processing functions such as encoding, packaging, encryption, content stitching etc. Deployment of these functions in the cloud is attractive as it enables flexibility in deployment options and resource allocation for the different functions. Yet, most of the time overprovisioning of cloud resources is necessary in order to meet demand variability. This can be costly, especially for large scale deployments. Prior art proposes resource allocation based on analytical models that minimize the costs of cloud deployments under a quality of service (QoS) constraint. However, these models do not sufficiently capture the underlying complexity of services composed of multiple processing functions. Instead, we introduce a novel methodology based on full-stack telemetry and machine learning to profile virtualized or cloud native media processing functions individually. The basis of the approach consists of investigating 4 categories of performance metrics: throughput, anomaly, latency and entropy (TALE) in offline (stress tests) and online setups using cloud telemetry. Machine learning is then used to profile the media processing function in the targeted cloud/NFV environment and to extract the most relevant cloud level Key Performance Indicators (KPIs) that relate to the final perceived quality and known client side performance indicators. The results enable more efficient monitoring, as only KPI related metrics need to be collected, stored and analyzed, reducing the storage and communication footprints by over 85%. In addition a detailed overview of the functions behavior was obtained, enabling optimized initial configuration and deployment, and more fine-grained dynamic online resource allocation reducing overprovisioning and avoiding function collapse. We further highlight the next steps towards cloud native carrier grade virtualized processing functions relevant for future network architectures such as in emerging 5G architectures.
Rufael Mekuria, Michael J. McGrath, Vincenzo Riccobene, Victor Bayon-Molino, Christos Tselios, John Thomson, Artem Dobrodub
MMSys6
2018 Large-scale hierarchical k-means for heterogeneous many-core supercomputers
Liandeng Li, Wenlai Zhao, Haohuan Fu, Guangwen Yang 0002, John Thomson
SC8
2017 ACTiCLOUD: Enabling the Next Generation of Cloud Applications
abstract
Despite their proliferation as a dominant computing paradigm, cloud computing systems lack effective mechanisms to manage their vast amounts of resources efficiently. Resources are stranded and fragmented, ultimately limiting cloud systems' applicability to large classes of critical applications that pose non-moderate resource demands. Eliminating current technological barriers of actual fluidity and scalability of cloud resources is essential to strengthen cloud computing's role as a critical cornerstone for the digital economy. ACTiCLOUD proposes a novel cloud architecture that breaks the existing scale-up and share-nothing barriers and enables the holistic management of physical resources both at the local cloud site and at distributed levels. Specifically, it makes advancements in the cloud resource management stacks by extending state-of-the-art hypervisor technology beyond the physical server boundary and localized cloud management system to provide a holistic resource management within a rack, within a site, and across distributed cloud sites. On top of this, ACTiCLOUD will adapt and optimize system libraries and runtimes (e.g., JVM) as well as ACTiCLOUD-native applications, which are extremely demanding, and critical classes of applications that currently face severe difficulties in matching their resource requirements to state-of-the-art cloud offerings.
Georgios I. Goumas, Konstantinos Nikas, Ewnetu Bayuh Lakew, Christos Kotselidis, Andrew Attwood, Erik Elmroth, Michail Flouris, Nikos Foutris, John Goodacre, Davide Grohmann, Vasileios Karakostas, Panagiotis Koutsourakis, Martin L. Kersten, Mikel Luján, Einar Rustad, John Thomson, Luis Tomás, Atle Vesterkjaer, Jim Webber, Ying Zhang 0027, Nectarios Koziris
ICDCS16
2016 EUROSERVER: Share-anything scale-out micro-server design
Manolis Marazakis, John Goodacre, Didier Fuin, Paul M. Carpenter, John Thomson, Emil Matús, Antimo Bruno, Per Stenström, Jérôme Martin, Yves Durand, Isabelle Dor
DATE5
2016 Predicting and Optimizing Image Compression
abstract
Image compression is a core task for mobile devices, social media and cloud storage backend services. Key evaluation criteria for compression are: the quality of the output, the compression ratio achieved and the computational time (and energy) expended. Predicting the effectiveness of standard compression implementations like libjpeg and WebP on a novel image is challenging, and often leads to non-optimal compression. This paper presents a machine learning-based technique to accurately model the outcome of image compression for arbitrary new images in terms of quality and compression ratio, without requiring significant additional computational time and energy. Using this model, we can actively adapt the aggressiveness of compression on a per image basis to accurately fit user requirements, leading to a more optimal compression.
Oleksandr Murashko, John Thomson, Hugh Leather
ACM Multimedia2
2014 EUROSERVER: Energy Efficient Node for European Micro-Servers
abstract
EUROSERVER is a collaborative project that aims to dramatically improve data centre energy-efficiency, cost, and software efficiency. It is addressing these important challenges through the coordinated application of several key recent innovations: 64-bit ARM cores, 3D heterogeneous silicon-on-silicon integration, and fully-depleted silicon-on-insulator (FD SOI) process technology, together with new software techniques for efficient resource management, including resource sharing and workload isolation. We are pioneering a system architecture approach that allows specialized silicon devices to be built even for low-volume markets where NRE costs are currently prohibitive. The EUROSERVER device will embed multiple silicon "chiplets" on an active silicon interposer. Its system architecture is being driven by requirements from three use cases: data centres and cloud computing, telecom infrastructures, and high-end embedded systems. We will build two fully integrated full-system prototypes, based on a common micro-server board, and targeting embedded servers and enterprise servers.
Yves Durand, Paul M. Carpenter, Stefano Adami, Angelos Bilas, Denis Dutoit, Alexis Farcy, Georgi Gaydadjiev, John Goodacre, Manolis Katevenis, Manolis Marazakis, Emil Matús, Iakovos Mavroidis, John Thomson
DSD13
2012 Preserving Scientific Processes from Design to Publications
Rudolf Mayer, Andreas Rauber, Martin Alexander Neumann, John Thomson, Gonçalo Antunes
TPDL4
2011 Automatic OpenCL Device Characterization: Guiding Optimized Kernel Design
Peter Thoman, Klaus Kofler, Heiko Studt, John Thomson, Thomas Fahringer
Euro-Par (2)4
2010 Workload characterization supporting the development of domain-specific compiler optimizations using decision trees for data mining
abstract
Embedded systems have successfully entered a broad variety of application domains such as automotive and industrial control, telecommunications, networking, digital media, consumer equipment, office automation and many more. In this paper we investigate if there exist any fundamental differences between application domains that justify the development and tuning of domain-specific compilers. We develop an automated approach that is capable of identifying domain-specific workload characterizations and presenting them in a readily interpretable format based on decision trees. The generated workload profiles summarize key resource utilization issues and enable compiler engineers to address the highlighted bottlenecks. We have evaluated our methodology against the industrial EEMBC benchmark suite and three popular embedded processors and have found that workload profiles differ significantly between application domains. We demonstrate that these characteristics can be exploited for the development of domain-specific compiler optimizations. In a case study we show average performance improvements of up to 44% for a class of networking applications.
Damon Fenacci, Björn Franke, John Thomson
SCOPES3
2006 Using Machine Learning to Focus Iterative Optimization
abstract
Iterative compiler optimization has been shown to outperform static approaches. This, however, is at the cost of large numbers of evaluations of the program. This paper develops a new methodology to reduce this number and hence speed up iterative optimization. It uses predictive modelling from the domain of machine learning to automatically focus search on those areas likely to give greatest performance. This approach is independent of search algorithm, search space or compiler infrastructure and scales gracefully with the compiler optimization space size. Off-line, a training set of programs is iteratively evaluated and the shape of the spaces and program features are modelled. These models are learnt and used to focus the iterative optimization of a new program. We evaluate two learnt models, an independent and Markov model, and evaluate their worth on two embedded platforms, the Texas Instrument C67I3 and the AMD Au1500. We show that such learnt models can speed up iterative search on large spaces by an order of magnitude. This translates into an average speedup of 1.22 on the TI C6713 and 1.27 on the AMD Au1500 in just 2 evaluations.
Felix V. Agakov, Edwin V. Bonilla, John Cavazos, Björn Franke, Grigori Fursin, Michael F. P. O'Boyle, John Thomson, Marc Toussaint, Christopher K. I. Williams
CGO7
2006 Predictive search distributions
abstract
Estimation of Distribution Algorithms (EDAs) are a popular approach to learn a probability distribution over the "good" solutions to a combinatorial optimization problem. Here we consider the case where there is a collection of such optimization problems with learned distributions, and where each problem can be characterized by some vector of features. Now we can define a machine learning problem to predict the distribution of good solutions q(s|x) for a new problem with features x, where s denotes a solution. This predictive distribution is then used to focus the search. We demonstrate the utility of our method on a compiler optimization task where the goal is to find a sequence of code transformations to make the code run fastest. Results on a set of 12 different benchmarks on two distinct architectures show that our approach consistently leads to significant improvements in performance.
Edwin V. Bonilla, Christopher K. I. Williams, Felix V. Agakov, John Cavazos, John Thomson, Michael F. P. O'Boyle
ICML5
2005 Probabilistic source-level optimisation of embedded programs
abstract
Efficient implementation of DSP applications is critical for many embedded systems. Optimising C compilers for embedded processors largely focus on code generation and instruction scheduling which, with their growing maturity, are providing diminishing returns. This paper empirically evaluates another approach, namely source-level transformations and the probabilistic feedback-driven search for “good ” transformation sequences within a large optimisation space. This novel approach combines two selection methods: one based on exploring the optimisation space, the other focused on localised search of good areas. This technique was applied to the UTDSP benchmark suite on two digital signal and multimedia processors (Analog Devices TigerSHARC TS-101, Philips TriMedia TM-1100) and an embedded processor derived from a popular general-purpose processor architecture (Intel Celeron 400). On average, our approach gave a factor of 1.71 times improvement across all platforms equivalent to an average 41 % reduction in execution time, outperforming existing approaches. In certain cases a speedup of up to ≈ 7 was found for individual benchmarks.
Björn Franke, Michael F. P. O'Boyle, John Thomson, Grigori Fursin
LCTES3