VLDB 2026 Research / reviewers in the wild / expert
Jeremy Singer
dblp:82/2176
· DBLP profile ↗
38ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0001-9462-6802ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 23 · 6 first-author · 12 since 2021Systems, architecture and hardware · 5 · 2 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Interpreter Memory Safety via Differential Fuzzing with a CHERI on TopabstractMemory safety is a critical issue in embedded systems. Although high-level languages like MicroPython simplify IoT development, their C-based runtimes remain vulnerable to memory errors triggered by Python code or native extensions. The CHERI (Capability Hardware Enhanced RISC Instructions) architecture offers hardware-enforced memory safety, but its effectiveness for exposing latent bugs in real-world interpreters has not yet been fully explored. We present diffCHERI:FruitFly, a novel differential testing framework for systematically uncovering memory defects in MicroPython across conventional (x86/ARM) and CHERI-enabled (Arm Morello) platforms. We mine historic vulnerabilities from diverse Python runtimes to extract recurring stress patterns, then use a large language model to generate new test programs, and apply Concrete Syntax Tree (CST) mutation to diversify inputs. Huanting Wang, Jeremy Singer, Zheng Wang 0001 |
ISMM | 3 |
| 2025 | Secure Scripting with CHERIoT MicroPythonabstractThe lean MicroPython runtime is a widely adopted high level programming framework for embedded microcontroller systems. However, the existing MicroPython codebase has limited security features, rendering it a fundamentally insecure runtime environment. This is a critical problem, given the growing deployment of highly interconnected IoT systems on which society depends. Malicious actors seek to compromise such embedded infrastructure, using sophisticated attack vectors. We have implemented a novel variant of MicroPython, adding support for runtime security features provided in the CHERI RISC-V architecture as instantiated by the CHERIoT-RTOS system. Our new MicroPython port supports hardware-enabled spatial memory safety, mitigating a large set of common runtime memory attacks. We have also compartmentalized the MicroPython runtime, to prevent untrusted code from elevating its permissions and taking control of the entire system. We perform a multi-faceted evaluation of our work, involving a qualitative security-focused case study and a quantitative performance analysis. The case study explores the full set of five publicly reported MicroPython vulnerabilities (CVEs). We demonstrate that the enhanced security provided by CHERIoT MicroPython mitigate two heap buffer overflow CVEs. Our performance analysis shows a geometric mean runtime overhead of 48% for secure execution across a set of ten standard Python benchmarks, although we argue this is indicative of worst-case overhead on our prototype platform and a realistic deployment overhead would be significantly lower. This work opens up a new, secure-by-design approach to IoT application development. Duncan Lowther, Dejice Jacob, Jacob Trevor, Jeremy Singer |
CC | 4 |
| 2025 | SecureMind: A Framework for Benchmarking Large Language Models in Memory Bug Detection and RepairabstractLarge language models (LLMs) hold great promise for automating software vulnerability detection and repair, but ensuring their correctness remains a challenge. While recent work has developed benchmarks for evaluating LLMs in bug detection and repair, existing studies rely on hand-crafted datasets that quickly become outdated. Moreover, systematic evaluation of advanced reasoning-based LLMs using chain-of-thought prompting for software security is lacking. We introduce SecureMind, an open-source framework for evaluating LLMs in vulnerability detection and repair, focusing on memory-related vulnerabilities. SecureMind provides a user-friendly Python interface for defining test plans, which automates data retrieval, preparation, and benchmarking across a wide range of metrics. Using SecureMind, we assess 10 representative LLMs, including 7 state-of-the-art reasoning models, on 16K test samples spanning 8 Common Weakness Enumeration (CWE) types related to memory safety violations. Our findings highlight the strengths and limitations of current LLMs in handling memory-related vulnerabilities. Huanting Wang, Dejice Jacob, David Kelly, Yehia El-khatib, Jeremy Singer, Zheng Wang 0001 |
ISMM | 5 |
| 2025 | OSPtrack: A Labeled Dataset Targeting Simulated Execution of Open-Source SoftwareabstractOpen-source software serves as a foundation for the internet and the cyber supply chain, but its exploitation is becoming increasingly prevalent. While advances in vulnerability detection for OSS have been significant, prior research has largely focused on static code analysis, often neglecting runtime indicators. To address this shortfall, we created a comprehensive dataset spanning five ecosystems, capturing features generated during the execution of packages and libraries in isolated environments. The dataset includes 9,461 package reports, of which 1,962 are identified as malicious, and encompasses both static and dynamic features such as files, sockets, commands, and DNS records. Each report is labeled with verified information and detailed sub-labels for attack types, facilitating the identification of malicious indicators when source code is unavailable. This dataset supports runtime detection, enhances detection model training, and enables efficient comparative analysis across ecosystems, contributing to the strengthening of supply chain security. Zhuoran Tan, Christos Anagnostopoulos 0001, Jeremy Singer |
MSR | 3 |
| 2025 | Advanced Persistent Threats Based on Supply Chain Vulnerabilities: Challenges, Solutions, and Future DirectionsabstractDue to the ever increasing interdependency across a variety of diverse software and hardware components in information and communications technology (ICT) provisioning, supply chain vulnerabilities (SCVs) targeting such dependencies have evolved as a primary choice for malicious actors to stealthy and complex cyber-attacks. The current modus operandi in the cyber threat spectrum is solely correlated with advanced persistent threats (APTs) that have shown to be prevalent across diversified attacks underpinning cyberwarfare and cybercrime. Hence, defense against such threats is undoubtedly considered as a high priority on a global scale. Nonetheless, the reliance on third-party supply chain software and device across diverse ICT ecosystems, combined with the current defense mechanisms’ inability to identify specific compromised entry points, results in an increased risk of APTs. This survey explores the state-of-the-art to stratify and showcase the properties of supply chain-based APTs, elaborate on reported risks from such APTs, and expand on existing defense methods. This study connects academic research with industry practices to highlight a new and growing problem. It examines supply chain compromises, offers unique insight into how these exploitations occur, and equips cybersecurity practitioners with the knowledge required to design next-generation APT defense mechanisms. Zhuoran Tan, Shameem A. Puthiya Parambath, Christos Anagnostopoulos 0001, Jeremy Singer, Angelos K. Marnerides |
IEEE Internet Things J. | 4 |
| 2024 | Online Coding Tutorial Systems: A New Category of Programming Learning PlatformsabstractThis paper presents a new category that has been added to the classification of Kim and Ko (2017) for programming learning systems, namely the Online Coding Tutorial System (OCTS) category. In this current study, firstly, seven popular online coding tutorial systems have been selected to investigate how these systems taught learners and what their characteristics and features were. Secondly, from Kim and Ko's classification, one system has been selected from each category and analyzed across the identified characteristics of online coding tutorial systems to investigate whether any existing category in Kim and Ko's classification shares the same characteristics. As a result, it was found that online coding tutorial systems have adopted many of the features that have been identified in Kim and Ko's first category of interactive platforms, along with some aspects of their creative platforms and MOOCs. Therefore, online coding tutorial systems have been considered a new category of programming learning systems that includes several characteristics from other existing categories. Ohud Abdullah Alasmari, Jeremy Singer, Mireilla Bikanga Ada |
COMPSAC | 2 |
| 2024 | Characterizing Dynamic Memory Behavior in WebAssembly WorkloadsabstractIn this early-stage study, we empirically characterize the runtime behavior of meaningful WebAssembly (Wasm) workloads with respect to memory allocation. We consider a variety of benchmarks and allocators, written in C and compiled to standalone Wasm, to give a broad spectrum of behavior. Yuxin Qin, Dejice Jacob, Jeremy Singer |
ISPASS | 3 |
| 2023 | Improving Robustness Against Adversarial Attacks with Deeply Quantized Neural NetworksabstractReducing the memory footprint of Machine Learning (ML) models, particularly Deep Neural Networks (DNNs), is essential to enable their deployment into resource-constrained tiny devices. However, a disadvantage of DNN models is their vulnerability to adversarial attacks, as they can be fooled by adding slight perturbations to the inputs. Therefore, the challenge is how to create accurate, robust, and tiny DNN models deployable on resource-constrained embedded devices. This paper reports the results of devising a tiny DNN model, robust to adversarial black and white box attacks, trained with an automatic quantization-aware training framework, i.e. QKeras, with deep quantization loss accounted in the learning loop, thereby making the designed DNNs more accurate for deployment on tiny devices. We investigated how QKeras and an adversarial robustness technique, Jacobian Regularization (JR), can provide a co-optimization strategy by exploiting the DNN topology and the per layer JR approach to produce robust yet tiny deeply quantized DNN models. As a result, a new DNN model implementing this co-optimization strategy was conceived, developed and tested on three datasets containing both images and audio inputs, as well as compared its performance with existing benchmarks against various white-box and black-box attacks. Experimental results demonstrated that on average our proposed DNN model resulted in 8.3% and 79.5% higher accuracy than MLCommons/Tiny benchmarks in the presence of white-box and black-box attacks on the CIFAR-10 image dataset and a subset of the Google Speech Commands audio dataset respectively. It was also 6.5% more accurate in the presence of black-box attacks on the SVHN image dataset. Ferheen Ayaz, Idris Zakariyya, José Cano 0001, Sye Loong Keoh, Jeremy Singer, Danilo Pau, Mounia Kharbouche-Harrari |
IJCNN | 5 |
| 2023 | Picking a CHERI Allocator: Security and Performance ConsiderationsabstractSeveral open-source memory allocators have been ported to CHERI, a hardware capability platform. In this paper we examine the security and performance of these allocators when run under CheriBSD on Arm's prototype Morello platform. We introduce a number of security attacks and show that all but one allocator are vulnerable to some of the attacks --- including the default CheriBSD allocator. We then show that while some forms of allocator performance are meaningful, comparing the performance of hybrid and pure capability (i.e. "running in non-CHERI vs. running in CHERI modes") allocators does not currently appear to be meaningful. Although we do not fully understand the reasons for this, it seems to be at least as much due to factors such as immature compiler toolchains and prototype hardware as it is due to the effects of capabilities on performance. Jacob Bramley, Dejice Jacob, Andrei Lascu, Jeremy Singer, Laurence Tratt |
ISMM | 4 |
| 2023 | Towards Secure MicroPython on Morello (WIP)abstractThe Arm Morello platform is a prototype system that supports hardware capabilities for improving runtime security. Although Morello is a server class compute component, there is ongoing work aimed at bringing architectural capabilities to embedded scale devices. For this reason, we are porting the MicroPython framework to Morello. Our intention is to understand the impact of hardware capabilities on lightweight runtime execution environments, like MicroPython, that target embedded devices. In this work-in-progress report, we describe the minimal modifications required to compile the C source code of MicroPython for Morello. We show that this approach gives a working, but not necessarily more secure, version of MicroPython. Our paper proceeds to outline how capabilities could be used to improve runtime system security for MicroPython runtime and hosted applications. Jeremy Singer |
LCTES | 1 |
| 2023 | Capable VMs Project Overview (Poster Abstract)abstractIn this poster, we will outline the scope and contributions of the Capable VMs project, in the framework of the UKRI Digital Security by Design programme. Jacob Bramley, Dejice Jacob, Andrei Lascu, Duncan Lowther, Jeremy Singer, Laurence Tratt |
MPLR | 5 |
| 2023 | Morello MicroPython: A Python Interpreter for CHERIabstractArm Morello is a prototype system that supports CHERI hardware capabilities for improving runtime security. As Morello becomes more widely available, there is a growing effort to port open source code projects to this novel platform. Although high-level applications generally need minimal code refactoring for CHERI compatibility, low-level systems code bases require significant modification to comply with the stringent memory safety constraints that are dynamically enforced by Morello. In this paper, we describe our work on porting the MicroPython interpreter to Morello with the CheriBSD OS. Our key contribution is to present a set of generic lessons for adapting managed runtime execution environments to CHERI, including (1) a characterization of necessary source code changes, (2) an evaluation of runtime performance of the interpreter on Morello, and (3) a demonstration of pragmatic memory safety bug detection. Although MicroPython is a lightweight interpreter, mostly written in C, we believe that the changes we have implemented and the lessons we have learned are more widely applicable. To the best of our knowledge, this is the first published description of meaningful experience for scripting language runtime engineering with CHERI and Morello. Duncan Lowther, Dejice Jacob, Jeremy Singer |
MPLR | 3 |
| 2023 | Could Tierless Languages Reduce IoT Development Grief?abstractInternet of Things (IoT) software is notoriously complex, conventionally comprising multiple tiers. Traditionally an IoT developer must use multiple programming languages and ensure that the components interoperate correctly. A novel alternative is to use a single tierless language with a compiler that generates the code for each component and ensures their correct interoperation. We report a systematic comparative evaluation of two tierless language technologies for IoT stacks: one for resource-rich sensor nodes (Clean with iTask) and one for resource-constrained sensor nodes (Clean with iTask and mTask). The evaluation is based on four implementations of a typical smart campus application: two tierless and two Python-based tiered. (1) We show that tierless languages have the potential to significantly reduce the development effort for IoT systems, requiring 70% less code than the tiered implementations. Careful analysis attributes this code reduction to reduced interoperation (e.g., two embedded domain-specific languages and one paradigm versus seven languages and two paradigms), automatically generated distributed communication, and powerful IoT programming abstractions. (2) We show that tierless languages have the potential to significantly improve the reliability of IoT systems, describing how Clean iTask/mTask maintains type safety, provides higher-order failure management, and simplifies maintainability. (3) We report the first comparison of a tierless IoT codebase for resource-rich sensor nodes with one for resource-constrained sensor nodes. The comparison shows that they have similar code size (within 7%), and functional structure. (4) We present the first comparison of two tierless IoT languages, one for resource-rich sensor nodes and the other for resource-constrained sensor nodes. Mart Lubbers, Pieter W. M. Koopman, Adrian Ramsingh, Jeremy Singer, Philip W. Trinder |
ACM Trans. Internet Things | 4 |
| 2022 | ELSA: A Keyword-based Searchable Encryption for Cloud-edge assisted Industrial Internet of ThingsabstractThe Industrial Internet of Things (IIoT) plays a powerful role in smart manufacturing by performing real-time analysis for large volumes of data. In addition, IIoT systems can monitor several factors, such as data accuracy, network bandwidth and operations latency. To perform these operations securely and in a privacy-preserving manner, one solution is to use cryptographic primitives. However, most cryptographic solutions add performance overhead causing latency. In this paper, we propose an Edge Lightweight Searchable Attribute-based encryption system (ELSA). ELSA leverages the cloud-edge architecture to improve search time beyond the state-of-the-art. The main contributions of this paper are as follows. First, we present an untrusted cloud/trusted edge architecture, which optimises the efficiency of data processing and decision making in the IIoT context. Second, we enhance search performance over current state-of-the-art (LSABE-MA) by an order of magnitude. We achieve this by improving the organisation of the data to provide better than linear search performance. We leverage the edge server to cluster data indices by keyword and introduce a query optimiser. The query optimiser uses k-means clustering to improve the efficiency of range queries, removing the need for linear search. In addition, we achieve this without sacrificing accuracy over the results. Jawhara Aljabri, Anna Lito Michala, Jeremy Singer |
CCGRID | 3 |
| 2022 | Boehm-Demers-Weiser Garbage Collection on MorelloabstractThe Boehm-Demers-Weiser collector is a conservative garbage collection scheme generally used for C/C++ applications. In this demo, we will show the collector running with C benchmark programs on a new CHERI-style hardware capability platform, specifically the Arm Morello system. Dejice Jacob, Jeremy Singer |
MPLR | 2 |
| 2022 | Characterizing WebAssembly BytecodeabstractWebAssembly, known as Wasm, is an interpreted portable bytecode execution format that is growing in popularity. In this work, we perform a simple pair of characterizations of Wasm. Statically, we compare the instruction set to other portable bytecodes like JVM and PCODE, showing that Wasm operates at a lower abstraction level than JVM, similar to PCODE. Dynamically, we study the Wasm instruction mix for a set of common benchmark applications. This investigation reveals that, like JVM, data movement operations occur most frequently in instruction traces. We conclude by discussing possible future directions for optimizing Wasm execution. Yuxin Qin, Dejice Jacob, Jeremy Singer |
MPLR | 3 |
| 2022 | Capability Boehm: challenges and opportunities for garbage collection with capability hardwareabstractThe Boehm-Demers-Weiser Garbage Collector (BDWGC) is a widely used, production-quality memory management framework for C and C++ applications. In this work, we describe our experiences in adapting BDWGC for modern capability hardware, in particular the CHERI system, which provides guarantees about memory safety due to runtime enforcement of fine-grained pointer bounds and permissions. Although many libraries and applications have been ported to CHERI already, to the best of our knowledge this is the first analysis of the complexities of transferring a garbage collector to CHERI. We describe various challenges presented by the CHERI micro-architectural constraints, along with some significant opportunities for runtime optimization. Since we do not yet have access to capability hardware, we present a limited study of software event counts on emulated micro-benchmarks. This experience report should be helpful to other systems implementors as they attempt to support the ongoing CHERI initiative. Dejice Jacob, Jeremy Singer |
VEE | 2 |
| 2022 | Classifying the Reliability of the Microservices ArchitectureabstractMicroservices are popular for web applications as they offer better scalability and reliability than monolithic architectures. Reliability is improved by loose coupling between individual microservices. However in production systems some microservices are tightly coupled, or chained together. We classify the reliability of microservices: if a minor microservice fails then the application continues to operate; if a critical microservice fails, the entire application fails. Combining reliability (minor/critical) with the established classifications of dependence (individual/chained) and state (stateful/stateless) defines a new three dimensional space: the Microservices Dependency State Reliability (MDSR) classification. Using three web application case studies (Hipster-Shop, Jupyter and WordPress) we identify microservice instances that exemplify the six points in MDSR. We present a prototype static analyser that can identify all six classes in Flask web applications, and apply it to s even applications. We explore case study examples that exhibit either a known reliability pattern or a bad smell. We show that our prototype static analyser can identify three of six patterns/bad smells in Flask web applications. Hence MDSR provides a structured classification of microservice software with the potential to improve reliability. Finally, we evaluate the reliability implications of the different MDSR classes by running the case study applications against a fault injector. Adrian Ramsingh, Jeremy Singer, Philip W. Trinder |
WEBIST | 2 |
| 2021 | Optimizing Task Allocation for Edge Micro-Clusters in Smart CitiesabstractCurrent urban technology trends like Internet-of-Things and 5G require ultra low latency compute resource to be distributed liberally at the network edge. We characterize and advocate the need for heterogeneous edge micro-clusters; these are pragmatic, low-power, low-cost, minimal footprint units that can provide sufficient resource for typical edge compute applications in smart cities. However, to make best use of heterogeneous edge micro-clusters, we require resource management techniques that are both efficient and effective. In this paper, we report on an empirical study to demonstrate that mathematical optimization (in particular, mixed integer programming) for resource management is appropriate in terms of overhead, also highly effective for executing batch-arrival workloads in smart city use cases. Yousef Alhaizaey, Jeremy Singer, Anna Lito Michala |
WOWMOM | 2 |
| 2020 | Pricing Python parallelism: a dynamic language cost model for heterogeneous platforms
Dejice Jacob, Philip W. Trinder, Jeremy Singer |
DLS | 3 |
| 2020 | Performance analysis of single board computer clustersabstractThe past few years have seen significant developments in Single Board Computer (SBC) hardware capabilities. These advances in SBCs translate directly into improvements in SBC clusters. In 2018 an individual SBC has more than four times the performance of a 64-node SBC cluster from 2013. This increase in performance has been accompanied by increases in energy efficiency (GFLOPS/W) and value for money (GFLOPS/$). We present systematic analysis of these metrics for three different SBC clusters composed of Raspberry Pi 3 Model B, Raspberry Pi 3 Model B+ and Odroid C2 nodes respectively. A 16-node SBC cluster can achieve up to 60 GFLOPS, running at 80 W. We believe that these improvements open new computational opportunities, whether this derives from a decrease in the physical volume required to provide a fixed amount of computation power for a portable cluster; or the amount of compute power that can be installed given a fixed budget in expendable compute scenarios. We also present a new SBC cluster construction form factor named Pi Stack; this has been designed to support edge compute applications rather than the educational use-cases favoured by previous methods. The improvements in SBC cluster performance and construction techniques mean that these SBC clusters are realising their potential as valuable developmental edge compute devices rather than just educational curiosities. Philip James Basford, Steven J. Ossont, Colin Perkins, Tony Garnock-Jones, Fung Po Tso 0001, Dimitrios P. Pezaros, Robert Mullins 0001, Eiko Yoneki, Jeremy Singer, Simon J. Cox 0001 |
Future Gener. Comput. Syst. | 9 |
| 2019 | Python programmers have GPUs too: automatic Python loop parallelization with staged dependence analysisabstractPython is a popular language for end-user software development in many application domains. End-users want to harness parallel compute resources effectively, by exploiting commodity manycore technology including GPUs. However, existing approaches to parallelism in Python are esoteric, and generally seem too complex for the typical end-user developer. We argue that implicit, or automatic, parallelization is the best way to deliver the benefits of manycore to end-users, since it avoids domain-specific languages, specialist libraries, complex annotations or restrictive language subsets. Auto-parallelization fits the Python philosophy, provides effective performance, and is convenient for non-expert developers. Dejice Jacob, Philip W. Trinder, Jeremy Singer |
DLS | 3 |
| 2019 | Experience Report: Thinkathon - Countering an "I Got It Working" Mentality with Pencil-and-Paper ExercisesabstractGoal-directed problem-solving labs can lead a student to believe that the most important achievement in a first programming course is to get programs working. This is counter to research indicating that code comprehension is an important developmental step for novice programmers. We observed this in our own CS-0 introductory programming course, and furthermore, that students weren't making the connection between code comprehension in labs and a final examination that required solutions to pencil-and-paper comprehension and writing exercises, where sound understanding of programming concepts is essential. Realising these deficiencies late in our course, we put on three 3-hour optional revision evenings just days before the exam. Based on a mastery learning philosophy, students were expected to work through a bank of around 200 pencil-and-paper exercises. By comparison with a machine-based hackathon, we called this a Thinkathon. Students completed a pre and post questionnaire about their experience of the Thinkathon. While we find that Thinkathon attendance positively influences final grades, we believe our reflection on the overall experience is of greater value. We report that: respected methods for developing code comprehension may not be enough on their own; novices must exercise their developing skills away from machines; and there are social learning outcomes in programming courses, currently implicit, that we should make explicit. Quintin I. Cutts, Matthew Barr, Mireilla Bikanga Ada, Peter Donaldson, Stephen W. Draper, Jack Parkinson, Jeremy Singer, Lovisa Sundin |
ITiCSE | 7 |
| 2018 | Peer-to-peer secure updates for heterogeneous edge devicesabstractWe consider the problem of securely distributing software updates to large scale clusters of heterogeneous edge compute nodes. Such nodes are needed to support the Internet of Things and low-latency edge compute scenarios, but are difficult to manage and update because they exist at the edge of the network behind NATs and firewalls that limit connectivity, or because they are mobile and have intermittent network access. We present a prototype secure update architecture for these devices that uses the combination of peer-to-peer protocols and automated NAT traversal techniques. This demonstrates that edge devices can be managed in an environment subject to partial or intermittent network connectivity, where there is not necessarily direct access from a management node to the devices being updated. Herry Herry, Emily Band, Colin Perkins, Jeremy Singer |
NOMS | 4 |
| 2018 | Next generation single board clustersabstractUntil recently, cluster computing was too expensive and too complex for commodity users. However the phenomenal popularity of single board computers like the Raspberry Pi has caused the emergence of the single board computer cluster. This demonstration will present a cheap, practical and portable Raspberry Pi cluster called Pi Stack. We will show pragmatic custom solutions to hardware issues, such as power distribution, and software issues, such as remote updating. We also sketch potential use cases for Pi Stack and other commodity single board computer cluster architectures. Jeremy Singer, Herry Herry, Philip James Basford, Wajdi Hajji, Colin Perkins, Fung Po Tso 0001, Dimitrios P. Pezaros, Robert Mullins 0001, Eiko Yoneki, Simon J. Cox 0001, Steven J. Ossont |
NOMS | 1 |
| 2018 | Commodity single board computer clusters and their applicationsabstractCurrent commodity Single Board Computers (SBCs) are sufficiently powerful to run mainstream operating systems and workloads. Many of these boards may be linked together, to create small, low-cost clusters that replicate some features of large data center clusters. The Raspberry Pi Foundation produces a series of SBCs with a price/performance ratio that makes SBC clusters viable, perhaps even expendable. These clusters are an enabler for Edge/Fog Compute, where processing is pushed out towards data sources, reducing bandwidth requirements and decentralizing the architecture. In this paper we investigate use cases driving the growth of SBC clusters, we examine the trends in future hardware developments, and discuss the potential of SBC clusters as a disruptive technology. Compared to traditional clusters, SBC clusters have a reduced footprint, are low-cost, and have low power requirements. This enables different models of deployment—particularly outside traditional data center environments. We discuss the applicability of existing software and management infrastructure to support exotic deployment scenarios and anticipate the next generation of SBC. We conclude that the SBC cluster is a new and distinct computational deployment paradigm, which is applicable to a wider range of scenarios than current clusters. It facilitates Internet of Things and Smart City systems and is potentially a game changer in pushing application logic out towards the network edge. Steven J. Ossont, Philip James Basford, Colin Perkins, Herry Herry, Fung Po Tso 0001, Dimitrios P. Pezaros, Robert Mullins 0001, Eiko Yoneki, Simon J. Cox 0001, Jeremy Singer |
Future Gener. Comput. Syst. | 10 |
| 2017 | Does CloudSim Accurately Model Micro Datacenters?abstractNovel cloud computing algorithms and techniques are initially evaluated via testbeds, simulators and mathematical models of datacenter infrastructure. However, it can be difficult to perform cross validation of these platforms against realistic scale infrastructures due to the prohibitive costs involved. This paper describes an approach to evaluating a cloud simulator through an empirical study involving a micro datacenter of commodity Raspberry Pi devices. To demonstrate the methodology, we compare performance of real-world workloads on this physical infrastructure against corresponding models of the workloads and infrastructure on the CloudSim simulator. After modelling a Raspberry Pi micro datacenter in CloudSim, we claim that the simulator lacks sufficient accuracy for cloud infrastructure experiments. Dhahi Alshammari, Jeremy Singer, Tim Storer |
CLOUD | 2 |
| 2015 | The judgment of forseti: economic utility for dynamic heap sizing of multiple runtimesabstractWe introduce the FORSETI system, which is a principled approach for holistic memory management. It permits a sysadmin to specify the total physical memory resource that may be shared between all concurrent virtual machines on a physical node. FORSETI models the heap size versus application throughput for each virtual machine, and seeks to maximize the combined throughput of the set of VMs based on concepts from economic utility theory. We evaluate the FORSETI system using a standard Java managed runtime, i.e. OpenJDK. Our results demonstrate that FORSETI enables dramatic reductions (up to 5x) in heap footprint without compromising application execution times. Callum Cameron, Jeremy Singer, David Vengerov |
ISMM | 2 |
| 2015 | Search-Based Refactoring: Metrics Are Not Enough
Chris Simons, Jeremy Singer, David Robert White |
SSBSE | 2 |
| 2014 | SICSA multicore challenge editorial prefaceabstractThis special issue reports on the SICSA multicore challenge, which commenced in 2010 and remains \nan ongoing activity. The aim is to produce a comparative evaluation of a range of standard parallel \nprogramming tools and techniques on a representative set of parallelizable problems, executing on \ncommodity multicore platforms. \nIn this introductory article, we outline contemporary multicore computing trends that give rise to \nthe challenge in Section 2. Then, we describe the specific motivation for the challenge in Section 3. \nWe summarize the parallelizable problem selected for implementation in the second challenge \nphase, the N-body problem, and studied in the papers in this special issue in Section 4. Finally \nin Section 5, we summarize the participation in challenge activities to date and give an overview of \nthe papers that were produced by challenge participants that appear in this particular special issue. \nWe intend this special issue to present the main lessons learnt from the challenge, providing \nmaterial of general relevance for multicore application developers and researchers. Hans-Wolfgang Loidl, Jeremy Singer |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | A Comparative Look at Adaptive Memory Management in Virtual MachinesabstractMemory management plays a vital role in modern virtual machines. Both system- and language-level VMs manage memory to give the illusion of a unbounded allocation space although the underlying physical resources are limited. One of the main challenges for memory management is the range of dynamic characteristics of the workloads. Researchers have developed a large body of work using different mechanisms and dynamic decision making to specialize the memory management system to specific workloads. This design can be considered as a control loop where sensors are monitored, decisions are made and actions are performed by actuators. Nevertheless as is common in systems research, improvement in one property is accomplished at the expense of some other property. In this work we survey different techniques for adaptive memory management expressed as a control loop. We propose to analyse memory management in virtual machines using three seemingly orthogonal characteristics: responsiveness (R), comprehensiveness (C) and intricateness (I). We then present the details of an extensible classification framework which emphasizes the tradeoffs of different approaches. Using this framework, some representative state of the art systems are evaluated showing inherent tensions between R, C and I. José Simão, Jeremy Singer, Luís Veiga |
CloudCom (1) | 2 |
| 2013 | Control theory for principled heap sizingabstractWe propose a new, principled approach to adaptive heap sizing based on control theory. We review current state-of-the-art heap sizing mechanisms, as deployed in Jikes RVM and HotSpot. We then formulate heap sizing as a control problem, apply and tune a standard controller algorithm, and evaluate its performance on a set of well-known benchmarks. We find our controller adapts the heap size more responsively than existing mechanisms. This responsiveness allows tighter virtual machine memory footprints while preserving target application throughput, which is ideal for both embedded and utility computing domains. In short, we argue that formal, systematic approaches to memory management should be replacing ad-hoc heuristics as the discipline matures. Control-theoretic heap sizing is one such systematic approach. David Robert White, Jeremy Singer, Jonathan M. Aitken, Richard E. Jones |
ISMM | 2 |
| 2013 | Cloud engineering is Search Based Software Engineering tooabstractMany of the problems posed by the migration of computation to cloud platforms can be formulated and solved using techniques associated with Search Based Software Engineering (SBSE). Much of cloud software engineering involves problems of optimisation: performance, allocation, assignment and the dynamic balancing of resources to achieve pragmatic trade-offs between many competing technical and business objectives. SBSE is concerned with the application of computational search and optimisation to solve precisely these kinds of software engineering challenges. Interest in both cloud computing and SBSE has grown rapidly in the past five years, yet there has been little work on SBSE as a means of addressing cloud computing challenges. Like many computationally demanding activities, SBSE has the potential to benefit from the cloud; ‘SBSE in the cloud’. However, this paper focuses, instead, of the ways in which SBSE can benefit cloud computing. It thus develops the theme of ‘SBSE for the cloud’, formulating cloud computing challenges in ways that can be addressed using SBSE. Mark Harman, Kiran Lakhotia, Jeremy Singer, David Robert White, Shin Yoo |
J. Syst. Softw. | 3 |
| 2011 | Garbage collection auto-tuning for Java mapreduce on multi-coresabstractMapReduce has been widely accepted as a simple programming pattern that can form the basis for efficient, large-scale, distributed data processing. The success of the MapReduce pattern has led to a variety of implementations for different computational scenarios. In this paper we present MRJ, a MapReduce Java framework for multi-core architectures. We evaluate its scalability on a four-core, hyperthreaded Intel Core i7 processor, using a set of standard MapReduce benchmarks. We investigate the significant impact that Java runtime garbage collection has on the performance and scalability of MRJ. We propose the use of memory management auto-tuning techniques based on machine learning. With our auto-tuning approach, we are able to achieve MRJ performance within 10% of optimal on 75% of our benchmark tests. Jeremy Singer, George Kovoor, Gavin Brown 0001, Mikel Luján |
ISMM | 1 |
| 2010 | The economics of garbage collectionabstractThis paper argues that economic theory can improve our understanding of memory management. We introduce the allocation curve, as an analogue of the demand curve from microeconomics. An allocation curve for a program characterises how the amount of garbage collection activity required during its execution varies in relation to the heap size associated with that program. The standard treatment of microeconomic demand curves (shifts and elasticity) can be applied directly and intuitively to our new allocation curves. As an application of this new theory, we show how allocation elasticity can be used to control the heap growth rate for variable sized heaps in Jikes RVM. Jeremy Singer, Richard E. Jones, Gavin Brown 0001, Mikel Luján |
ISMM | 1 |
| 2008 | Exploiting the Correspondence between Micro Patterns and Class NamesabstractThis paper argues that semantic information encoded in natural language identifiers is a largely neglected resource for program analysis.First we show that words in Java class names relate to class properties, expressed using the recently developed micro patterns language.We analyze a large corpus of Java programs to create a database that links common class name words with micro patterns. Finally we report on prototype tools integrated with the Eclipse development environment. These tools use the database to inform programmers of particular problems or optimization opportunities in their code. Jeremy Singer, Chris C. Kirkham |
SCAM | 1 |
| 2008 | Dynamic analysis of Java program concepts for visualization and profiling
Jeremy Singer, Chris C. Kirkham |
Sci. Comput. Program. | 1 |
| 2007 | Intelligent selection of application-specific garbage collectorsabstractJava program execution times vary greatly with different garbage collection algorithms. Until now, it has not been possible to determine the best GC algorithm for aparticular program without exhaustively profiling that program for all available GC algorithms. This paper presents a new approach. We use machine learning techniques to build a prediction model that, given asingle profile run of a previously unseen Java program,can predict a good GC algorithm for that program. We implement this technique in Jikes RVM and test it onseveral standard benchmark suites. Our techniqueachieves 5% speedup in overall execution time (averagedacross all test programs for all heap sizes) compared with selecting the default GC algorithm in every trial. We present further experiments to show that an oracle predictor could achieve an average 17% speedup on the same experiments. In addition, we provide evidence to suggest that GC behaviour is sometimes independent of program inputs. These observations lead us to propose that intelligent selection of GC algorithms is suitably straight forward, efficient and effective to merit further exploration regarding its potential inclusion in the general Java software deployment process. Jeremy Singer, Gavin Brown 0001, Ian Watson, John Cavazos |
ISMM | 1 |