VLDB 2026 Research / reviewers in the wild / expert
Andy D. Pimentel
dblp:p/AndyDPimentel
· DBLP profile ↗
73ranked-venue papers
6as first author
31since 2021 · last 2026
0000-0002-2043-4469ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 49 · 6 first-author · 18 since 2021Software engineering, systems software and programming languages · 11 · 6 since 2021Artificial intelligence and machine learning · 7 · 4 since 2021Theory of computation · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SwiftSNNI: Optimized Scheduling for Secure Neural Network Inference (SNNI) on Multi-Core SystemsabstractSecure Neural Network Inference (SNNI) enables privacy-preserving inference on encrypted data with strong cryptographic guarantees. However, practical deployments suffer from high preprocessing overhead, significant communication costs, and sequential execution. These limitations lead to low throughput, underutilized system resources, long queueing delays, and poor scalability. This work introduces SwiftSNNI, a unified, resource-aware scheduling framework for SNNI. It implements a hybrid offline–online strategy that orchestrates offline preprocessing (Tpre,i) and online inference (Ton,i) jobs to maximize parallelism. By formulating SNNI scheduling as a constrained optimization problem, SwiftSNNI overlaps Tpre,i phase execution of future requests with active Ton,j, jobs. SwiftSNNI also incorporates optional advance notices to enable proactive Tpre,i, which further reduces average input delay (D). Evaluations using five benchmark neural networks (M1, M2, HiNet, AlexNet, VGG-16) under diverse workloads and stochastic arrival rates confirm substantial performance gains. Compared to a parallelized sequential baseline (MS-SHARK), SwiftSNNI achieves up to 97% lower average input delay (D), a 81% reduction in makespan (≈ 5.4 × speedup), and delivers 5.6 × increase in throughput. Furthermore, SwiftSNNI reduces average waiting time (W) by over 99%, demonstrating robust starvation prevention for high-concurrency workloads. SwiftSNNI supports concurrent execution, scales to larger neural networks, and provides an efficient runtime for SNNI deployments. The SwiftSNNI implementation is available online. Kanwal Batool, Saleem Anwar, Francesco Regazzoni 0001, Andy D. Pimentel, Zoltán Ádám Mann |
ICPE | 4 |
| 2025 | MaCP: Minimal yet Mighty Adaptation via Hierarchical Cosine ProjectionabstractWe present a new adaptation method MaCP, Minimal yet Mighty adaptive Cosine Projection, that achieves exceptional performance while requiring minimal parameters and memory for fine-tuning large foundation models.Its general idea is to exploit the superior energy compaction and decorrelation properties of cosine projection to improve both model efficiency and accuracy.Specifically, it projects the weight change from the low-rank adaptation into the discrete cosine space.Then, the weight change is partitioned over different levels of the discrete cosine spectrum, and each partition's most critical frequency components are selected.Extensive experiments demonstrate the effectiveness of MaCP across a wide range of single-modality tasks, including natural language understanding, natural language generation, text summarization, as well as multimodality tasks such as image classification and video understanding.MaCP consistently delivers superior accuracy, significantly reduced computational complexity, and lower memory requirements compared to existing alternatives. Yixian Shen, Qi Bi, Jia-Hong Huang, Hongyi Zhu 0004, Andy D. Pimentel, Anuj Pathania |
ACL (1) | 5 |
| 2025 | Empowering Sustainability: Energy Labeling of Digital Services Using SimulationabstractThe energy consumption of digital services has become a concern for stakeholders committed to sustainability. Raising awareness of this consumption is essential to improve the energy efficiency of digital services. However, expressing the energy usage of digital services in an easily understandable and actionable way remains a challenge. We address this challenge by proposing a first operational energy labeling method for digital services in the computing continuum. Our approach enables stakeholders, including cloud and network providers, application developers, researchers, and end-users of digital services, to better understand and improve the energy efficiency of their applications. Focusing on video surveillance digital services, and using the enhanced iFogSim framework, we propose an energy labeling scheme, and demonstrate its merits with extensive scenario analysis and simulation. We further discuss how our approach can help reduce energy consumption and/or improve performance, all without modifying the application's functional parameters or system architecture. Saeedeh Baneshi, Anuj Pathania, Benny Akesson, Andy D. Pimentel, Ana Lucia Varbanescu |
CCGrid | 4 |
| 2025 | Unraveling Parallelism in Automated Workload Modeling for Distributed Cyber-Physical SystemsabstractDesigning next generation distributed CyberPhysical Systems (dCPS) requires effective Design Space Exploration (DSE) methods to evaluate system design alternatives and their impact on performance. While existing DSE approaches focus on hardware optimization and software-to-hardware mapping, they often overlook parallel execution opportunities within software tasks. Current application workload models for complex dCPS assume fixed execution orders, limiting the ability to explore and exploit software parallelism. To address this issue, we propose refined workload models derived from execution traces that capture both inter- and intra-process dependencies. Building on these models, we present a method to identify tasks that can be safely reordered or executed in parallel without modifying the existing software implementation. We validate our approach through a case study on the ASML Twinscan lithography machine, demonstrating measurable performance improvements without impacting the system functional correctness. Faezeh Sadat Saadatmand, Todor P. Stefanov, Andy D. Pimentel, Benny Akesson, Ignacio Gonález Alonso |
DSD | 3 |
| 2025 | PEIR: Modeling Performance in Neural Information Retrieval
Pooya Khandel, Andrew Yates, Ana Lucia Varbanescu, Maarten de Rijke, Andy D. Pimentel |
ECIR (2) | 5 |
| 2025 | SSH: Sparse Spectrum Adaptation via Discrete Hartley TransformationabstractYixian Shen, Qi Bi, Jia-hong Huang, Hongyi Zhu, Andy D. Pimentel, Anuj Pathania. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yixian Shen, Qi Bi, Jia-Hong Huang, Hongyi Zhu 0004, Andy D. Pimentel, Anuj Pathania |
NAACL (Long Papers) | 5 |
| 2025 | Model and system robustness in distributed CNN inference at the edgeabstractPrevalent large CNN models pose a significant challenge in terms of computing resources for resource-constrained devices at the Edge. Distributing the computations and coefficients over multiple edge devices collaboratively has been well studied but these works generally do not consider the presence of device failures (e.g., due to temporary connectivity issues, overload, discharged battery of edge devices). Such unpredictable failures can compromise the reliability of edge devices, inhibiting the proper execution of distributed CNN inference. In this paper, we present a novel partitioning method, called RobustDiCE, for robust distribution and inference of CNN models over multiple edge devices. Our method can tolerate intermittent and permanent device failures in a distributed system at the Edge, offering a tunable trade-off between robustness (i.e., retaining model accuracy after failures) and resource utilization . We verify the system’s robustness by validating the overall end-to-end latency under failures. We evaluate RobustDiCE using the ImageNet-1K dataset on several representative CNN models under various device failure scenarios and compare it with several state-of-the-art partitioning methods as well as an optimal robustness approach (i.e., full neuron replication). In addition, we demonstrate RobustDiCE’s advantages in terms of memory usage and energy consumption per device, and system throughput for various system setups with different device counts. Xiaotian Guo, Quan Jiang, Andy D. Pimentel, Todor P. Stefanov |
Integr. | 3 |
| 2024 | RobustDiCE: Robust and Distributed CNN Inference at the EdgeabstractPrevalent large CNN models pose a significant challenge in terms of computing resources for resource-constrained devices at the Edge. Distributing the computations and coefficients over multiple edge devices collaboratively has been well studied but these works generally do not consider the presence of device failures (e.g., due to temporary connectivity issues, overload, discharged battery, etc. of edge devices). Such unpredictable failures can compromise the reliability of edge devices, inhibiting the proper execution of distributed CNN inference. In this paper, we present a novel partitioning method, called RobustDiCE, for robust distribution and inference of CNN models over multiple edge devices. Our method can tolerate intermittent and permanent device failures in a distributed system at the Edge, offering a tunable trade-off between robustness (i.e., retaining model accuracy after failures) and resource utilization. We evaluate RobustDiCE using the ImageNet-1K dataset on several representative CNN models under various device failure scenarios and compare it with several state-of-the-art partitioning methods as well as an optimal robustness approach (i.e., full neuron replication). In addition, we demonstrate RobustDiCE’s advantages in terms of memory usage and energy consumption per device, and system throughput for various system set-ups with different device counts. Xiaotian Guo, Quan Jiang, Andy D. Pimentel, Todor P. Stefanov |
ASPDAC | 3 |
| 2024 | Education Abstract: Design Space Exploration for Deep Learning at the EdgeabstractThe AI revolution, fueled by effective Deep Learning approaches, has seen a recent shift towards processing the AI workloads closer to the user, at the Edge. This paper addresses the instrumental role of system-level design space exploration (DSE) methods for achieving efficient inference of deep-learning models on resource-constrained devices at the Edge. Andy D. Pimentel |
CASES | 1 |
| 2024 | PiQi: Partially Quantized DNN Inference on HMPSoCsabstractDeep Neural Network (DNN) inference is now ubiquitous in embedded applications at the edge. State-of-the-art Heterogeneous Multi-Processors System-on-Chip (HMPSoCs) powering these applications come equipped with powerful Neural Processing Units (NPUs) that significantly outperform other inference-capable HMPSoC components - namely, the CPUs and GPUs - in terms of power consumption and performance. However, CPUs and GPUs can perform full precision inference, whereas NPUs can often only perform a quantized inference. Consequently, low-latency, low-power inference by the NPU comes at an accuracy loss due to the quantization. Ehsan Aghapour, Yixian Shen, Dolly Sapra, Andy D. Pimentel, Anuj Pathania |
ISLPED | 4 |
| 2024 | Using Evolutionary Algorithms to Find Cache-Friendly Generalized Morton Layouts for ArraysabstractThe layout of multi-dimensional data can have a significant impact on the efficacy of hardware caches and, by extension, the performance of applications. Common multi-dimensional layouts include the canonical row-major and column-major layouts as well as the Morton curve layout. In this paper, we describe how the Morton layout can be generalized to a very large family of multi-dimensional data layouts with widely varying performance characteristics. We posit that this design space can be efficiently explored using a combinatorial evolutionary methodology based on genetic algorithms. To this end, we propose a chromosomal representation for such layouts as well as a methodology for estimating the fitness of array layouts using cache simulation. We show that our fitness function correlates to kernel running time in real hardware, and that our evolutionary strategy allows us to find candidates with favorable simulated cache properties in four out of the eight real-world applications under consideration in a small number of generations. Finally, we demonstrate that the array layouts found using our evolutionary method perform well not only in simulated environments but that they can effect significant performance gains---up to a factor ten in extreme cases---in real hardware. Stephen Nicholas Swatman, Ana Lucia Varbanescu, Andy D. Pimentel, Andreas Salzburger, Attila Krasznahorkay |
ICPE | 3 |
| 2024 | EASTER: Learning to Split Transformers at the Edge RobustlyabstractPrevalent large transformer models present significant computational challenges for resource-constrained devices at the Edge. While distributing the workload of deep learning models across multiple edge devices has been extensively studied, these works typically overlook the impact of failures of edge devices. Unpredictable failures, due to, e.g., connectivity issues or discharged batteries, can compromise the reliability of inference serving at the Edge. In this article, we introduce a novel methodology, called EASTER, designed to learn robust distribution strategies for transformer models against device failures that consider the tradeoff between robustness (i.e., maintaining model functionality against failures) and resource utilization (considering memory usage and computations). We evaluate EASTER with three representative transformers—ViT, GPT-2, and Vicuna—under device failures. Our results demonstrate EASTER’s efficiency in memory usage, and possible end-to-end latency improvement for inference across multiple edge devices while preserving model accuracy as much as possible under device failures. Xiaotian Guo, Quan Jiang, Yixian Shen, Andy D. Pimentel, Todor P. Stefanov |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | ARM-CO-UP: ARM COoperative Utilization of ProcessorsabstractHMPSoCs combine different processors on a single chip. They enable powerful embedded devices, which increasingly perform ML inference tasks at the edge. State-of-the-art HMPSoCs can perform on-chip embedded inference using different processors, such as CPUs, GPUs, and NPUs. HMPSoCs can potentially overcome the limitation of low single-processor CNN inference performance and efficiency by cooperative use of multiple processors. However, standard inference frameworks for edge devices typically utilize only a single processor. We present the ARM-CO-UP framework built on the ARM-CL library. The ARM-CO-UP framework supports two modes of operation – Pipeline and Switch. It optimizes inference throughput using pipelined execution of network partitions for consecutive input frames in the Pipeline mode. It improves inference latency through layer-switched inference for a single input frame in the Switch mode. Furthermore, it supports layer-wise CPU/GPU DVFS in both modes for improving power efficiency and energy consumption. ARM-CO-UP is a comprehensive framework for multi-processor CNN inference that automates CNN partitioning and mapping, pipeline synchronization, processor type switching, layer-wise DVFS , and closed-source NPU integration. Ehsan Aghapour, Dolly Sapra, Andy D. Pimentel, Anuj Pathania |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2023 | Thermal Management for S-NUCA Many-Cores via Synchronous Thread RotationsabstractOn-chip thermal management is quintessential to a thermally safe operation of a many-core processor. The presence of a physically distributed logically shared Last-Level Cache (LLC) significantly reduces the performance penalty of migrating threads within the cores of an S-NUCA many-core. This cost reduction allows novel thermal management of these many-cores via synchronous thread migration. Synchronous thread migration provides a viable alternative to Dynamic Voltage and Frequency Scaling (DVFS) and asynchronous thread migration used traditionally to manage thermals of S-NUCA many-cores. We present a theoretical method to compute the peak tem-perature in many-cores with synchronous thread migrations. We use the method to create a thermal management heuristic called HotPotato that maximizes the performance of S-NUCA many-cores under a peak temperature constraint. We implement HotPotato within the state-of-the-art HotSniper simulator. Detailed interval thermal simulations with HotSniper show an average 10.72% improvement in response time of S-NUCA many-cores when scheduling with HotPotato compared to a state-of-the-art thermal-aware S-NUCA scheduler. Yixian Shen, Sobhan Niknam, Anuj Pathania, Andy D. Pimentel |
DATE | 4 |
| 2023 | FLORIA: A Fast and Featherlight Approach for Predicting Cache PerformanceabstractThe cache Miss Ratio Curve (MRC) serves a variety of purposes such as cache partitioning, application profiling and code tuning. In this work, we propose a new metric, called cache miss distribution, that describes cache miss behavior over cache sets, for predicting cache MRCs. Based on this metric, we present FLORIA, a software-based, online approach that approximates cache MRCs on commodity systems. By polluting a tunable number of cache lines in some selected cache sets using our designed microbenchmark, the cache miss distribution for the target workload is obtained via hardware performance counters with the support of precise event based sampling (PEBS). A model is developed to predict the MRC of the target workload based on its cache miss distribution. Jun Xiao 0009, Yaocheng Xiang, Xiaolin Wang 0001, Yingwei Luo, Andy D. Pimentel, Zhenlin Wang 0003 |
ICS | 5 |
| 2023 | PELSI: Power-Efficient Layer-Switched InferenceabstractConvolutional Neural Networks (CNNs) are now quintessential kernels within embedded computer vision applications deployed in edge devices. Heterogeneous Multi-Processor System-on-Chips (HMPSoCs) with Dynamic Voltage and Frequency Scaling (DVFS) capable components (CPUs and GPUs) allow for low-latency, low-power CNN inference on resource-constrained edge devices when employed efficiently. CNNs comprise several heterogeneous layer types that execute with different degrees of power efficiency on different HMPSoC components at different frequencies. We propose the first framework, PELSI, that exploits this layer-wise power efficiency heterogeneity for power-efficient CPU-GPU layer-switched CNN interference on HMPSoCs. PELSI executes each layer of a CNN on an HMPSoC component (CPU or GPU) clocked at just the right frequency for every layer such that the CNN meets its inference latency target with minimal power consumption while still accounting for the power-performance overhead of multiple switching between CPU and GPU mid-inference. PELSI incorporates a Genetic Algorithm (GA) to identify the near-optimal CPU-GPU layer-switched CNN inference configuration from within the large exponential design space that meets the given latency requirement most power efficiently. We evaluate PELSI on Rock-Pi embedded platform. The platform contains an RK3399Pro HMPSoC with DVFS-capable CPU clusters and GPU. Empirical evaluations with five different CNNs show a 44.48% improvement in power efficiency for CNN inference under PELSI over the state-of-the-art. Ehsan Aghapour, Dolly Sapra, Andy D. Pimentel, Anuj Pathania |
RTCSA | 3 |
| 2023 | Analyzing Digital Services Across the Compute Continuum Using iFogSimabstractDigital services enable users to interact with a broad range of applications and, as such, have become an essential part of our daily lives. Although convenient, their ubiquity comes at a significant cost in energy, raising sustainability concerns. We access these services by triggering a computing continuum, spanning from the device to the edge, fog, and cloud. Scheduling decisions made at each layer impact the overall quality of service (QoS) and energy consumption of digital services. Saeedeh Baneshi, Ana Lucia Varbanescu, Anuj Pathania, Benny Akesson, Andy D. Pimentel |
RTCSA | 5 |
| 2023 | Systematically Exploring High-Performance Representations of Vector Fields Through Compile-Time CompositionabstractWe present a novel benchmark suite for implementations of vector fields in high-performance computing environments to aid developers in quantifying and ranking their performance. We decompose the design space of such benchmarks into access patterns and storage backends, the latter of which can be further decomposed into components with different functional and non-functional properties. Through compile-time meta-programming, we generate a large number of benchmarks with minimal effort and ensure the extensibility of our suite. Our empirical analysis, based on real-world applications in high-energy physics, demonstrates the feasibility of our approach on CPU and GPU platforms, and highlights that our suite is able to evaluate performance-critical design choices. Finally, we propose that our work towards composing vector fields from elementary components is not only useful for the purposes of benchmarking, but that it naturally gives rise to a novel library for implementing such fields in domain applications. Stephen Nicholas Swatman, Ana Lucia Varbanescu, Andy D. Pimentel, Andreas Salzburger, Attila Krasznahorkay |
ICPE | 3 |
| 2023 | Automated Exploration and Implementation of Distributed CNN Inference at the EdgeabstractFor model inference of convolutional neural networks (CNNs), we nowadays witness a shift from the Cloud to the Edge. Unfortunately, deploying and inferring large, compute- and memory-intensive CNNs on Internet of Things devices at the Edge is challenging as they typically have limited resources. One approach to address this challenge is to leverage all available resources across multiple edge devices to execute a large CNN by properly partitioning it and running each CNN partition on a separate edge device. However, there currently does not exist a design and programming framework that takes a trained CNN model as input and subsequently allows for efficiently exploring and automatically implementing a range of different CNN partitions on multiple edge devices to facilitate distributed CNN inference. Therefore, in this article, we propose a novel framework that automates the splitting of a CNN model into a set of submodels as well as the code generation needed for the distributed and collaborative execution of these submodels on multiple, possibly heterogeneous, edge devices, while supporting the exploitation of parallelism among and within the edge devices. In addition, since the number of different CNN mapping possibilities on multiple edge devices is vast, our framework also features a multistage and hierarchical design space exploration methodology to efficiently search for (near-)optimal distributed CNN inference implementations. Our experimental results demonstrate that our work allows for rapidly finding and realizing distributed CNN inference implementations with reduced energy consumption and memory usage per edge device, and under certain conditions, with improved system throughput as well. Xiaotian Guo, Andy D. Pimentel, Todor P. Stefanov |
IEEE Internet Things J. | 2 |
| 2023 | Thermal Management for 3D-Stacked Systems via Unified Core-Memory Power Regulationabstract3D-stacked processor-memory systems stack memory (DRAM banks) directly on top of logic (CPU cores) using chiplet-on-chiplet packaging technology to provide the next-level computing performance in embedded platforms. Stacking, however, severely increases the system’s power density without any accompanying increase in the heat dissipation capacity. Consequently, 3D-stacked processor-memory systems suffer more severe thermal issues than their non-stacked counterparts. Nevertheless, 3D-stacked processor-memory systems do inherit power (thermal) management knobs from their non-stacked predecessors - namely Dynamic Voltage and Frequency Scaling (DVFS) for cores and Low Power Mode (LPM) for memory banks. In the context of 3D-stacked processor-memory systems, DVFS and LPM are performance- and power-wise deeply intertwined. Their non-unified independent use on 3D-stacked processor-memory systems results in sub-optimal thermal management. The unified use of DVFS and LPM for thermal management for 3D-stacked processor-memory systems remains unexplored. The lack of implementation of LPM in thermal simulators for 3D-stacked processor-memory systems hinders real-world representative evaluation for a unified approach. We extend the state-of-the-art interval thermal simulator for 3D-stacked processor-memory systems CoMeT with an LPM power management knob for memory banks. We also propose a learning-based thermal management technique for 3D-stacked processor-memory systems that employ DVFS and LPM in a unified manner. Detailed interval thermal simulations with the extended CoMeT framework show a 10.15% average response time improvement with the PARSEC and SPLASH-2 benchmark suites, along with widely-used Deep Neural Network (DNN) workloads against a state-of-the-art thermal management technique for 2.5D processor-memory systems (ported directly to 3D-stacked processor-memory systems) that also proposes unified use of DVFS and LPM. Yixian Shen, Leo Schreuders, Anuj Pathania, Andy D. Pimentel |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2022 | CPU-GPU Layer-Switched Low Latency CNN InferenceabstractConvolutional Neural Networks (CNNs) inference on Heterogeneous Multi-Processor System-on-Chips (HMPSoCs) in edge devices represent cutting-edge embedded machine learning. Embedded CPU and GPU within an HMPSoC can both perform inference using CNNs. However, common practice is to run a CNN on the HMPSoC component (CPU or GPU) provides the best performance (lowest latency) for that CNN. CNNs are not monolithic and are composed of several layers of different types. Some of these layers have lower latency on the CPU, while others execute faster on the GPU. In this work, we investigate the reason behind this observation. We also propose an execution of CNN that switches between CPU and GPU at the layer granularity, wherein a CNN layer executes on the component that provides it with the lowest latency. Switching between the CPU and the GPU back and forth mid-inference introduces additional overhead (delay) in the inference. Regardless of overhead, we show in this work that a CPU-GPU layer switched execution results in, on average, having 4.72% lower CNN inference latency on the Khadas VIM 3 board with Amlogic A311D HMPSoC. Ehsan Aghapour, Dolly Sapra, Andy D. Pimentel, Anuj Pathania |
DSD | 3 |
| 2022 | Design Space Exploration for Distributed Cyber-Physical Systems: State-of-the-art, Challenges, and DirectionsabstractIndustrial Cyber-Physical Systems (CPS) are com-plex heterogeneous and distributed computing systems, typically integrating and interconnecting a large number of subsystems and containing a substantial number of hardware and software components. Producers of these distributed Cyber-Physical Systems (dCPS) face serious challenges with respect to designing the next generations of these machines and require proper support in making (early) design decisions to avoid expensive and time consuming oversights. This calls for efficient and scalable system-level Design Space Exploration (DSE) methods for dCPS. In this position paper, we review the current state of the art in DSE, and argue that efficient and scalable DSE technology for dCPS is more or less non-existing and constitutes a largely unchartered research area. Moreover, we identify several re-search challenges that need to be addressed and discuss possible directions for targeting such DSE technolozy for dCPS. Marius Herget, Faezeh Sadat Saadatmand, Martin Bor, Ignacio Gonzalez Alonso, Todor P. Stefanov, Benny Akesson, Andy D. Pimentel |
DSD | 7 |
| 2022 | Model-Based Testing of Internet of Things Protocols
Xavier Manuel van Dommelen, Machiel van der Bijl, Andy D. Pimentel |
FMICS | 3 |
| 2022 | TCPS: a task and cache-aware partitioned scheduler for hard real-time multi-core systemsabstractShared caches in multi-core processors seriously complicate the timing verification of real-time software tasks due to the task interference occurring in the shared caches. Explicitly calculating the amount of cache interference among tasks and cache partitioning are two major approaches to enhance the schedulability performance in the context of multi-core processors with shared caches. The former approach suffers from pessimistic cache interference estimations that subsequently result in suboptimal schedulability performance, whereas the latter approach may increase the execution time of tasks due to a lower cache usage, also degrading the schedulability performance. Yixian Shen, Jun Xiao 0009, Andy D. Pimentel |
LCTES | 3 |
| 2022 | Modelling Performance Loss due to Thread Imbalance in Stochastic Variable-Length SIMT WorkloadsabstractWhen designing algorithms for single-instruction multiple-thread (SIMT) devices such as general purpose graphics processing units (GPGPUs), thread imbalance is an important performance consideration. Thread imbalance can emerge in iterative applications where workloads are of variable length, because threads processing larger amounts of work will cause threads with less work to idle. This form of thread imbalance influences the design space of algorithms-particularly in terms of processing granularity-but we lack models to quantify its impact on application performance. In this paper, we present a statistical model for quantifying the performance loss due to thread imbalance for iterative SIMT applications with stochastic, variable-length workloads. Our model is designed to operate with minimal knowledge of the implementation details of the algorithm, relying solely on an understanding of the probability distribution of the lengths of the workloads. We validate our model against a synthetic benchmark based on a Monte Carlo simulation of matrix exponentiation, and show that our model achieves nearly perfect accuracy. Compared to empirical data extracted from real hardware, our model maintains a high degree of accuracy, predicting mean performance loss within a margin of 2%. Stephen Nicholas Swatman, Ana Lucia Varbanescu, Attila Krasznahorkay, Andy D. Pimentel |
MASCOTS | 4 |
| 2022 | Designing convolutional neural networks with constrained evolutionary piecemeal trainingabstractAbstract The automated architecture search methodology for neural networks is known as Neural Architecture Search (NAS). In recent times, Convolutional Neural Networks (CNNs) designed through NAS methodologies have achieved very high performance in several fields, for instance image classification and natural language processing. Our work is in the same domain of NAS, where we traverse the search space of neural network architectures with the help of an evolutionary algorithm which has been augmented with a novel approach of piecemeal-training. In contrast to the previously published NAS techniques, wherein the training with given data is considered an isolated task to estimate the performance of neural networks, our work demonstrates that a neural network architecture and the related weights can be jointly learned by combining concepts of the traditional training process and evolutionary architecture search in a single algorithm. The consolidation has been realised by breaking down the conventional training technique into smaller slices and collating them together with an integrated evolutionary architecture search algorithm. The constraints on architecture search space are placed by limiting its various parameters within a specified range of values, consequently regulating the neural network’s size and memory requirements. We validate this concept on two vastly different datasets, namely, the CIFAR-10 dataset in the domain of image classification, and PAMAP2 dataset in the Human Activity Recognition (HAR) domain. Starting from randomly initialized and untrained CNNs, the algorithm discovers models with competent architectures, which after complete training, reach an accuracy of of 92.5% for CIFAR-10 and 94.36% PAMAP2. We further extend the algorithm to include an additional conflicting search objective: the number of parameters of the neural network. Our multi-objective algorithm produces a Pareto optimal set of neural networks, by optimizing the search for both the accuracy and the parameter count, thus emphasizing the versatility of our approach. Dolly Sapra, Andy D. Pimentel |
Appl. Intell. | 2 |
| 2022 | Improving the robustness of industrial Cyber-Physical Systems through machine learning-based performance anomaly identificationabstractWe propose a versatile and fully data-centric methodology towards anomaly detection and identification in modern industrial Cyber–Physical Systems (CPS). Our motivation behind this move is the ever-growing computerisation in these systems, in the form of complex distributed computing nodes, running complex distributed software. Industrial CPS also demonstrate heavy deployment of hardware sensors, as well as an increasing role for software. We observe the insufficiency and costliness of design-time measures in prevention of anomalies. As our main contribution, our methodology is taking advantage of this data-rich environment by means of Extra-Functional Behaviour (EFB) monitoring, analytics pipelines and Artificial Intelligence (AI). Specifically, we demonstrate the use of compartmentalisation of execution timelines into distinct units, i.e., execution phases. We introduce the generation of representations for these phases, i.e., behavioural signatures and behavioural passports, as our way of behavioural fingerprinting. Composed using regression modelling techniques, signatures as the representation of ongoing behaviour, are compared to passports, representing reference behaviour. The comparison is done by means of goodness-of-fit scores, creating quantifiable measures of deviation between different recorded behaviours. We have used both partially synthetic and real-world traces in our experiments, depending on the use-case. We have also followed both white box and black box approaches for our use-cases, with discussions on the pros and cons of each. The effectiveness of our data-centric methodology is demonstrated by two proofs-of-concept from the industry, to represent the two ends of the industrial CPS complexity spectrum, with one being a large semiconductor photolithography machine, while the other is an image analysis platform. Each use-case comes with its own characteristics and limitations, confirming the flexibility of our methodology and the relevance of its integral steps in the approach towards the initial analysis and data transformations. The results of anomaly classification show overall high accuracies, as high as 99% in certain set-ups. These results show the capability of our data-centric methodology, suiting the presented modern industrial CPS designs. Uraz Odyurt, Andy D. Pimentel, Ignacio Gonzalez Alonso |
J. Syst. Archit. | 2 |
| 2022 | Scenario Based Run-Time Switching for Adaptive CNN-Based Applications at the EdgeabstractConvolutional Neural Networks (CNNs) are biologically inspired computational models that are at the heart of many modern computer vision and natural language processing applications. Some of the CNN-based applications are executed on mobile and embedded devices. Execution of CNNs on such devices places numerous demands on the CNNs, such as high accuracy, high throughput, low memory cost, and low energy consumption. These requirements are very difficult to satisfy at the same time, so CNN execution at the edge typically involves trade-offs (e.g., high CNN throughput is achieved at the cost of decreased CNN accuracy). In existing methodologies, such trade-offs are either chosen once and remain unchanged during a CNN-based application execution, or are adapted to the properties of the CNN input data. However, the application needs can also be significantly affected by the changes in the application environment, such as a change of the battery level in the edge device. Thus, CNN-based applications need a mechanism that allows to dynamically adapt their characteristics to the changes in the application environment at run-time. Therefore, in this article, we propose a scenario-based run-time switching (SBRS) methodology, that implements such a mechanism. Svetlana Minakova, Dolly Sapra, Todor P. Stefanov, Andy D. Pimentel |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2022 | Cache Interference-aware Task Partitioning for Non-preemptive Real-time Multi-core SystemsabstractShared caches in multi-core processors introduce serious difficulties in providing guarantees on the real-time properties of embedded software due to the interaction and the resulting contention in the shared caches. Prior work has studied the schedulability analysis of global scheduling for real-time multi-core systems with shared caches. This article considers another common scheduling paradigm: partitioned scheduling in the presence of shared cache interference. To achieve this, we propose CITTA, a cache interference-aware task partitioning algorithm. We first analyze the shared cache interference between two programs for set-associative instruction and data caches. Then, an integer programming formulation is constructed to calculate the upper bound on cache interference exhibited by a task, which is required by CITTA. We conduct schedulability analysis of CITTA and formally prove its correctness. A set of experiments is performed to evaluate the schedulability performance of CITTA against global EDF scheduling and other greedy partition approaches such as First-fit and Worst-fit over randomly generated tasksets and realistic workloads in embedded systems. Our empirical evaluations show that CITTA outperforms global EDF scheduling and greedy partition approaches in terms of task sets deemed schedulable. Jun Xiao 0009, Yixian Shen, Andy D. Pimentel |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2021 | T-TSP: Transient-Temperature Based Safe Power Budgeting in Multi-/Many-Core ProcessorsabstractPower budgeting techniques allow thermally safe operation in multi-/many-core processors while still allowing for efficient exploitation of available thermal headroom. Core-level power budgeting techniques like Thermal Safe Power (TSP) have allowed for more efficient operations than chip-level power budgeting techniques like Thermal Design Power (TDP) since the liner granularity permits operations closer to the threshold temperature without thermal violations.State-of-the-art TSP bases its power budgeting calculations on the long-term steady-state temperature of cores while ignoring trends in their short-term transient temperature. In this paper, we propose a new power budgeting technique called T-TSP (Transient-Temperature-based Safe Power) that bases its calculation on the current temperature of the core, a detail ignored by TSP. T-TSP provides a dynamic power budget to a core, which inversely correlates with the core’s thermal headroom. Dynamic power budgeting with T-TSP allows cores to reach the threshold temperature faster than TSP and operate safely close to it in perpetuity. Therefore, it provides the same thermal guarantees as TSP but enables even more efficient exploitation of thermal headroom.We integrate T-TSP with a state-of-the-art thermal interval simulation toolchain. Our detailed evaluations show that benchmarks execute faster by up to 17.94% and 8.37% on average when we do power budgeting with T-TSP instead of the state-of- the-art TSP. Finally, we make T-TSP publicly available in both its integrated and stand-alone forms. Sobhan Niknam, Anuj Pathania, Andy D. Pimentel |
ICCD | 3 |
| 2021 | The Choice of AI Matters: Alternative Machine Learning Approaches for CPS Anomalies
Uraz Odyurt, Dolly Sapra, Andy D. Pimentel |
IEA/AIE (2) | 3 |
| 2020 | An evolutionary optimization algorithm for gradually saturating objective functionsabstractEvolutionary algorithms have been actively studied for dynamic optimization problems in the last two decades, however the research is mainly focused on problems with large, periodical or abrupt changes during the optimization. In contrast, this paper concentrates on gradually changing environments with an additional imposition of a saturating objective function. This work is motivated by an evolutionary neural architecture search methodology where a population of Convolutional Neural Networks (CNNs) is evaluated and iteratively modified using genetic operators during the training process. The objective of the search, namely the prediction accuracy of a CNN, is a continuous and slow moving target, increasing with each training epoch and eventually saturating when the training is nearly complete. Population diversity is an important consideration in dynamic environments wherein a large diversity restricts the algorithm from converging to a small area of the search space while the environment is still transforming. Our proposed algorithm adaptively influences the population diversity, depending on the rate of change of the objective function, using disruptive crossovers and non-elitist population replacements. We compare the results of our algorithm with a traditional evolutionary algorithm and demonstrate that the proposed modifications improve the algorithm performance in gradually saturating dynamic environments. Dolly Sapra, Andy D. Pimentel |
GECCO | 2 |
| 2020 | Deep Learning Model Reuse and Composition in Knowledge Centric NetworkingabstractMachine learning has inadvertently pioneered the transition of big data into big knowledge. Machine learning models absorb and incorporate knowledge from large scale data through training and can be regarded as a representation of the knowledge learnt. There are multitude of use cases where this acquired knowledge can be used to enhance future applications or speed up the training of new models. Yet, the efficient sharing, exploitation and reusability of this knowledge remains a challenge. In this paper we propose a framework for deep learning models that facilitates the reuse of model architectures, transfer coefficients between models for knowledge composition and updates, and apply compression and pruning techniques for efficient storage and communication. We discuss the framework and its application in the context of Knowledge Centric Networking (KCN) and demonstrate the framework potential through various experiments, i.e. when knowledge has to be updated to accommodate new (raw) data or to reduce complexity. Dolly Sapra, Andy D. Pimentel |
ICCCN | 2 |
| 2020 | Constrained Evolutionary Piecemeal Training to Design Convolutional Neural Networks
Dolly Sapra, Andy D. Pimentel |
IEA/AIE | 2 |
| 2020 | CITTA: Cache Interference-aware Task Partitioning for Real-time Multi-core SystemsabstractShared caches in multi-core processors introduce serious difficulties in providing guarantees on the real-time properties of embedded software due to the interaction and the resulting contention in the shared caches. Prior work has studied the schedulability analysis of global scheduling for real-time multi-core systems with shared caches. This paper considers another common scheduling paradigm: partitioned scheduling in the presence of shared cache interference. To achieve this, we propose CITTA, a cache-interference aware task partitioning algorithm. An integer programming formulation is constructed to calculate the upper bound on cache interference exhibited by a task, which is required by CITTA. We conduct schedulability analysis of CITTA and formally prove its correctness. A set of experiments is performed to evaluate the schedulability performance of CITTA against global EDF scheduling over randomly generated tasksets. Our empirical evaluations show that CITTA outperforms global EDF scheduling in terms of task sets deemed schedulable. Jun Xiao 0009, Andy D. Pimentel |
LCTES | 2 |
| 2020 | Schedulability Analysis of Global Scheduling for Multicore Systems With Shared CachesabstractShared caches in multicore processors introduce serious difficulties in providing guarantees on the real-time properties of embedded software due to the interaction and the resulting contention in the shared caches. To address this problem, we develop a new schedulability analysis for real-time multicore systems with shared caches, globally scheduled by Earliest Deadline First (EDF) and Fixed Priority (FP) algorithms. We construct an integer programming formulation, which can be transformed to an integer linear programming formulation, to calculate an upper bound on cache interference exhibited by a task within a given execution window. Using the integer programming formulation, an iterative algorithm is presented to obtain the upper bound on cache interference a task may exhibit during one job execution. The upper bound on cache interference is subsequently integrated into the schedulability analysis to derive a new schedulability condition. A range of experiments is performed to investigate how the schedulability is degraded by shared cache interference. We also evaluate the schedulability performance of EDF against FP scheduling over randomly generated tasksets. Our empirical evaluations show that EDF is better than FP scheduling in terms of the number of task sets deemed schedulable. Jun Xiao 0009, Sebastian Altmeyer, Andy D. Pimentel |
IEEE Trans. Computers | 3 |
| 2019 | Optimization and deployment of CNNs at the edge: the ALOHA experienceabstractDeep learning (DL) algorithms have already proved their effectiveness on a wide variety of application domains, including speech recognition, natural language processing, and image classification. To foster their pervasive adoption in applications where low latency, privacy issues and data bandwidth are paramount, the current trend is to perform inference tasks at the edge. This requires deployment of DL algorithms on low-energy and resource-constrained computing nodes, often heterogenous and parallel, that are usually more complex to program and to manage without adequate support and experience. In this paper, we present ALOHA, an integrated tool flow that tries to facilitate the design of DL applications and their porting on embedded heterogenous architectures. The proposed tool flow aims at automating different design steps and reducing development costs. ALOHA considers hardware-related variables and security, power efficiency, and adaptivity aspects during the whole development process, from pre-training hyperparameter optimization and algorithm configuration to deployment. Paolo Meloni, Daniela Loi, Paola Busia, Gianfranco Deriu, Andy D. Pimentel, Dolly Sapra, Todor P. Stefanov, Svetlana Minakova, Francesco Conti 0001, Luca Benini, Maura Pintor, Battista Biggio, Bernhard Moser 0001, Natalia Shepeleva, Nikos Fragoulis, Ilias Theodorakopoulos, Michael Masin, Francesca Palumbo |
CF | 5 |
| 2019 | CPpf: a prefetch aware LLC partitioning approachabstractHardware cache prefetching is deployed in modern multicore processors to reduce memory latencies, addressing the memory wall problem. However, it tends to increase the Last Level Cache (LLC) contention among applications in multiprogrammed workloads, leading to a performance degradation for the overall system. To study the interaction between hardware prefetching and LLC cache management, we first analyze the variation of application performance when varying the effective LLC space in the presence and absence of hardware prefetching. We observe that hardware prefetching can compensate the application performance loss due to the reduced effective cache space. Motivated by this observation, we classify applications into two categories, prefetching sensitive (PS) and non prefetching sensitive (NPS) applications, by the degree of performance benefit they experience from hardware prefetchers. To address the cache contention and also to mitigate the potential prefetch-related cache interference, we propose CPpf, a cache partitioning approach for improving the shared cache management in the presence of hardware prefetching. CPpf consists of a method using Precise Event-Based Sampling techniques for the online classification of PS and NPS applications and a cache partitioning scheme using Cache Allocation technology to distribute the cache space among PS and NPS applications. We implemented CPpf as a user-level runtime system on Linux. Compared with a non-partitioning approach, CPpf achieves speedups of up to 1.20, 1.08 and 1.06 for workloads with 2, 4 and 8 single-threaded applications, respectively. Moreover, it achieves speedups of up to 1.22 and 1.11 for workloads composed of two applications with 4 threads and 8 threads, respectively. Jun Xiao 0009, Andy D. Pimentel, Xu Liu 0001 |
ICPP | 2 |
| 2018 | Communication-centric analysis of complex embedded computing systems: work-in-progressabstractWe show experimental evidence and argue that communication-centric modelling of complex embedded computing systems provides predictive power over the workload dependent behaviour of these systems. System and external observables included in this behaviour can be utilised in the system's analysis. We provide the preliminary results from our detection (monitoring) and imitation (simulation) phases, both part of a larger workflow in development. Uraz Odyurt, Hugo Meyer, Simon Polstra, Evangelos Paradas, Ignacio Gonzalez Alonso, Andy D. Pimentel |
EMSOFT | 6 |
| 2017 | EDiFy: An Execution time Distribution FinderabstractEmbedded real-time systems are subjected to stringent timing constraints. Analysing their timing behaviour is therefore of great significance. So far, research on the timing behaviour of real-time systems has been primarily focused on finding out what happens in the worst-case (i.e., finding the worst case execution time, or WCET). Boudewijn Braams, Sebastian Altmeyer, Andy D. Pimentel |
DAC | 3 |
| 2017 | Schedulability Analysis of Non-preemptive Real-Time Scheduling for Multicore Processors with Shared CachesabstractShared caches in multicore processors introduce serious difficulties in providing guarantees on the real-time properties of embedded software due to the interaction and the resulting contention in the shared caches. To address this problem, we develop a new schedulability analysis for real-time multicore systems with shared caches. To the best of our knowledge, this is the first work that addresses the schedulability problem with inter-core cache interference. We construct an integer programming formulation, which can be transformed to an integer linear programming formulation, to calculate an upper bound on cache interference exhibited by a task within a given execution window. Using the integer programming formulation, an iterative algorithm is presented to obtain the upper bound on cache interference a task may exhibit during one job execution. The upper bound on cache interference is subsequently integrated into the schedulability analysis to derive a new schedulability condition. A range of experiments is performed to investigate how the schedulability is degraded by shared cache interference. Jun Xiao 0009, Sebastian Altmeyer, Andy D. Pimentel |
RTSS | 3 |
| 2016 | Scenario-based run-time adaptive MPSoC systems
Andy D. Pimentel |
J. Syst. Archit. | 2 |
| 2015 | Fast and precise cache performance estimation for out-of-order execution
Roeland Douma, Sebastian Altmeyer, Andy D. Pimentel |
DATE | 3 |
| 2015 | A Hybrid Task Mapping Algorithm for Heterogeneous MPSoCsabstractThe application workloads in modern MPSoC-based embedded systems are becoming increasingly dynamic. Different applications concurrently execute and contend for resources in such systems, which could cause serious changes in the intensity and nature of the workload demands over time. To cope with the dynamism of application workloads at runtime and improve the efficiency of the underlying system architecture, this article presents a hybrid task mapping algorithm that combines a static mapping exploration and a dynamic mapping optimization to achieve an overall improvement of system efficiency. We evaluate our algorithm using a heterogeneous MPSoC system with three real applications. Experimental results reveal the effectiveness of our proposed algorithm by comparing derived solutions to the ones obtained from several other runtime mapping algorithms. In test cases with three simultaneously active applications, the mapping solutions derived by our approach have average performance improvements ranging from 45.9% to 105.9% and average energy savings ranging from 14.6% to 23.5%. Andy D. Pimentel |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2014 | A system-level simulation framework for evaluating task migration in MPSoCsabstractTask migration is the transfer of the execution of a process (task) from one processing element to another. It originates from the massive deployment of distributed systems in the parallel computing field to enable dynamic load distribution, fault resilience and to enhance data access locality. With the development of MultiProcessor System-on-Chip (MPSoC) architectures, the topic of task migration has recently regained research interest in the embedded systems domain. In this paper, we present a high-level simulation framework to study task migration for MPSoC systems. With this framework, different migration methodologies on different underlying hardware systems can be easily and rapidly modeled, simulated and evaluated during the early stages of design. By using this high-level simulation framework, a designer can study the migration impact on the overall performance of the system by exploring different task migration mechanisms (determining what and how to migrate) or using different migration policies (determining when to migrate which tasks whereto) in a specific task migration mechanism. Using a number of experiments, we demonstrate the capabilities of our simulation framework. Andy D. Pimentel |
CASES | 2 |
| 2014 | Towards Exploring Vast MPSoC Mapping Design Spaces Using a Bias-Elitist Evolutionary ApproachabstractThe problem of optimally mapping a set of tasks onto a set of given heterogeneous processors for maximal throughput has been known, in general, to be NP-complete. Previous research has shown that Genetic Algorithms (GA) typically are a good choice to solve this problem when the solution space is relatively small. However, when the size of the problem space increases, classic genetic algorithms still suffer from the problem of long evolution times. To address this problem, this paper proposes a novel bias-elitist genetic algorithm that is guided by domain-specific heuristics to speed up the evolution process. Experimental results reveal that our proposed algorithm is able to handle large scale task mapping problems and produces high-quality mapping solutions in only a short time period. Andy D. Pimentel |
DSD | 2 |
| 2014 | Emulating Asymmetric MPSoCs on the Intel SCC Many-Core ProcessorabstractThe Single-chip Cloud Computer (SCC) is a 48-core experimental processor created by Intel Labs targeting the manycore research community. It has extensive frequency and voltage scaling support as well as on board power monitors. In this paper we present a detailed study of the power properties of the SCC. Then, we show how the SCC can be used as a substrate to emulate asymmetric multi processor systems on chip (MPSoCs) to be used for studying power/performance trade-offs. Roy Bakker, Michiel W. van Tol, Andy D. Pimentel |
PDP | 3 |
| 2013 | A scenario-based run-time task mapping algorithm for MPSoCsabstractThe application workloads in modern MPSoC-based embedded systems are becoming increasingly dynamic. Different applications concurrently execute and contend for resources in such systems which could cause serious changes in the intensity and nature of the workload demands over time. To cope with the dynamism of application workloads at run time and improve the efficiency of the underlying system architecture, this paper presents a novel scenario-based run-time task mapping algorithm. This algorithm combines a static mapping strategy based on workload scenarios and a dynamic mapping strategy to achieve an overall improvement of system efficiency. We evaluated our algorithm using a homogeneous MPSoC system with three real applications. From the results, we found that our algorithm achieves an 11.3% performance improvement and a 13.9% energy saving compared to running the applications without using any run-time mapping algorithm. When comparing our algorithm to three other, well-known run-time mapping algorithms, it is superior to these algorithms in terms of quality of the mappings found while also reducing the overheads compared to most of these algorithms. Andy D. Pimentel |
DAC | 2 |
| 2013 | Exploiting domain knowledge in system-level MPSoC design space exploration
Mark Thompson 0001, Andy D. Pimentel |
J. Syst. Archit. | 2 |
| 2013 | Fitness Prediction Techniques for Scenario-Based Design Space ExplorationabstractModern embedded systems are becoming increasingly multifunctional. The dynamism in multifunctional embedded systems manifests itself with more dynamic applications and the presence of multiple applications executing on a single embedded system. This dynamism in the application workload must be taken into account during the early system-level design space exploration (DSE) of multiprocessor system-on-a-chip (MPSoC)-based embedded systems. Scenario-based DSE utilizes the concept of application scenarios to search for optimal mappings of a multi-application workload onto an MPSoC. The scenario-based DSE uses a multi-objective genetic algorithm (GA) to identifying the mapping with the best average quality for all the application scenarios in the workload. In order to keep the exploration of the scenario-based DSE efficient, fitness prediction is used to obtain the quality of a mapping. This fitness prediction is performed using a representative subset of application scenarios that is obtained using co-exploration of the scenario subset space. In this paper, multiple fitness prediction techniques are presented: stochastic, deterministic, and a hybrid combination. Results show that, for our test cases, accurate fitness prediction is already provided for subsets containing only 1-4% of the application scenarios. Larger subsets will obtain a similar accuracy, but the DSE will require more time to identify promising mappings that meet the requirements of multifunctional embedded systems. Peter van Stralen, Andy D. Pimentel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2013 | A system-level infrastructure for multidimensional MP-SoC design space co-explorationabstractIn this article, we present a flexible and extensible system-level MP-SoC design space exploration (DSE) infrastructure, called NASA. This highly modular framework uses well-defined interfaces to easily integrate different system-level simulation tools as well as different combinations of search strategies in a simple plug-and-play fashion. Moreover, NASA deploys a so-called dimension-oriented DSE approach, allowing designers to configure the appropriate number of, well-tuned and possibly different, search algorithms to simultaneously co-explore the various design space dimensions. As a result, NASA provides a flexible and re-usable framework for the systematic exploration of the multidimensional MP-SoC design space, starting from a set of relatively simple user specifications. To demonstrate the capabilities of the NASA framework and to illustrate its distinct aspects, we also present several DSE experiments in which, for example, we compare NASA configurations using a single search algorithm for all design space dimensions to configurations using a separate search algorithm per dimension. These proof-of-concept experiments indicate that the latter multidimensional co-exploration can find better design points and evaluates a higher diversity of design alternatives as compared to the more traditional approach of using a single search algorithm for all dimensions. Zai Jian Jia, Tomás Bautista, Antonio Núñez, Andy D. Pimentel, Mark Thompson 0001 |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2012 | Design space pruning through hybrid analysis in system-level design space explorationabstractSystem-level design space exploration (DSE), which is performed early in the design process, is of eminent importance to the design of complex multi-processor embedded system architectures. During system-level DSE, system parameters like, e.g., the number and type of processors, the type and size of memories, or the mapping of application tasks to architectural resources, are considered. Simulation-based DSE, in which different design instances are evaluated using system-level simulations, typically are computationally costly. Even using high-level simulations and efficient exploration algorithms, the simulation time to evaluate design points forms a real bottleneck in such DSE. Therefore, the vast design space that needs to be searched requires effective design space pruning techniques. This paper presents a technique to reduce the number of simulations needed during system-level DSE. More specifically, we propose an iterative design space pruning methodology based on static throughput analysis of different application mappings. By interleaving these analytical throughput estimations with simulations, our hybrid approach can significantly reduce the number of simulations that are needed during the process of DSE. Roberta Piscitelli, Andy D. Pimentel |
DATE | 2 |
| 2012 | Introduction to the Special Section on ESTIMedia'08abstractNo abstract available. Mladen Berekovic, Samarjit Chakraborty, Petru Eles, Andy D. Pimentel |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2012 | Introduction to special section ESTIMedia'09abstractNo abstract available. Andy D. Pimentel, Naehyuck Chang, Mladen Berekovic |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2010 | Evaluation of runtime task mapping heuristics with rSesame - a case studyabstractrSesame is a generic modeling and simulation framework which can explore and evaluate reconfigurable systems at the early design stages. The framework can be used to explore different HW/SW partitionings, task mappings and scheduling strategies at both design time and runtime. The framework strives for a high degree of flexibility, ease of use, fast performance and applicability. In this paper, we want to evaluate the framework's characteristics by showing that it can easily and quickly model, simulate and compare a wide range of runtime mapping heuristics from various domains. A case study with a Motion-JPEG (MJPEG) application demonstrates that the presented model can be efficiently used to model and simulate a wide variety of mapping heuristics as well as to perform runtime exploration of various non-functional design parameters such as execution time, number of reconfigurations, area usage, etc. Kamana Sigdel, Mark Thompson 0001, Carlo Galuzzi, Andy D. Pimentel, Koen Bertels |
DATE | 4 |
| 2010 | Visualization of Multi-objective Design Space Exploration for Embedded SystemsabstractModern embedded systems come with contradictory design constraints. On one hand, these systems often target mass production and battery-based devices, and therefore should be cheap and power efficient. On the other hand, they need to achieve high (real-time) performance. This wide spectrum of design requirements leads to complex heterogeneous system-on-chip (SoC) architectures. The complexity of embedded systems forces designers to model and simulate systems and their components to explore the wide range of design choices. Such design space exploration is especially needed during the early design stages, where the design space is at its largest. Due to the exponential design space in real problems and multiple criteria to be considered, multi-objective evolutionary algorithms (MOEAs) are often used to trim down a large design space into a finite set of points and provide the designer a set of tradable solutions with respect to the design criteria. Interpreting the search results (e.g., where are the Pareto points located), understanding their relations and analyzing how the design space was searched by such searching algorithms is of invaluable importance to the designer. To this end, this paper presents a novel interactive visualization tool, based on tree visualization, to understand the search dynamics of a MOEA and to visualize where the optimum design points are located in the design space and what objective values they have. Toktam Taghavi, Andy D. Pimentel |
DSD | 2 |
| 2010 | VMODEX: A visualization tool for multi-objective Design Space ExplorationabstractVMODEX is an interactive visualization tool to support system-level Design Space Exploration (DSE). It provides insight into the search process of Multi-Objective Evolutionary Algorithms (MOEAs) that are typically used in the DSE process, and therefore it facilitates the analysis of the DSE results. In our tool, we provide several capabilities to be able to handle large design spaces and filter design points according to their objective values to see only preferred solutions. Toktam Taghavi, Andy D. Pimentel |
FPT | 2 |
| 2010 | Scenario-based design space exploration of MPSoCsabstractEarly design space exploration (DSE) is a key ingredient in system-level design of MPSoC-based embedded systems. The state of the art in this field typically still explores systems under a single, fixed application workload. In reality, however, the applications are concurrently executing and contending for system resources in such systems. As a result, the intensity and nature of application demands can change dramatically over time. This paper therefore introduces the concept of workload scenarios in the DSE process, capturing dynamic behavior both within and between applications. More specifically, we present and evaluate a novel, scenario-based DSE approach based on a coevolutionary genetic algorithm. Peter van Stralen, Andy D. Pimentel |
ICCD | 2 |
| 2009 | Introduction
Felix Wolf 0001, Andy D. Pimentel, Luiz De Rose, Soonhoi Ha, Thilo Kielmann, Anna Sikora |
Euro-Par | 2 |
| 2009 | System-level runtime mapping exploration of reconfigurable architecturesabstractDynamic reconfigurable systems can evolve under various conditions due to changes imposed either by the architecture, or by the applications, or by the environment. In such systems, the design process becomes more sophisticated as all the design decisions have to be optimized in terms of runtime behaviors and values. Runtime mapping exploration allows to explore reconfigurable systems at runtime to optimize task mappings in order to adapt to the changing behavior of the application(s), the architecture, or the environment. Performing such explorations at runtime enables a system to be more efficient in terms of various design constraints such as performance, chip area, power consumption, etc. Towards this goal, in this paper, we present a model that facilitates runtime mapping exploration of reconfigurable architectures. A case study of an MJPEG application shows that the presented model can be used to perform runtime exploration of various functional and non-functional design parameters. Kamana Sigdel, Mark Thompson 0001, Andy D. Pimentel, Carlo Galuzzi, Koen Bertels |
IPDPS | 3 |
| 2009 | Electronic System-Level Synthesis MethodologiesabstractWith ever-increasing system complexities, all major semiconductor roadmaps have identified the need for moving to higher levels of abstraction in order to increase productivity in electronic system design. Most recently, many approaches and tools that claim to realize and support a design process at the so-called electronic system level (ESL) have emerged. However, faced with the vast complexity challenges, in most cases at best, only partial solutions are available. In this paper, we develop and propose a novel classification for ESL synthesis tools, and we will present six different academic approaches in this context. Based on these observations, we can identify such common principles and needs as they are leading toward and are ultimately required for a true ESL synthesis solution, covering the whole design process from specification to implementation for complete systems across hardware and software boundaries. Andreas Gerstlauer, Christian Haubelt, Andy D. Pimentel, Todor P. Stefanov, Daniel Gajski, Jürgen Teich |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2008 | Daedalus: toward composable multimedia MP-SoC designabstractDaedalus is a system-level design flow for the design of multiprocessor system-on-chip (MP-SoC) based embedded multimedia systems. It offers a fully integrated tool-flow in which design space exploration (DSE), system-level synthesis, application mapping, and system prototyping of MP-SoCs are highly automated. In this paper, we describe our first industrial deployment experiences with the Daedalus framework. Daedalus is currently being deployed in the early stages of the design of an image compression system for very high resolution cameras targeting medical appliances. In this context, we performed a DSE study with a JPEG encoder application, which exploits both task and data parallelism. This application was mapped onto a range of different MP-SoC architectures. We achieved a performance speed-up of up to 20x compared to a single processor system. In addition, the results show that the Daedalus high-level MP-SoC models accurately predict the overall system performance, i.e., the performance error is around 5%. Hristo Nikolov, Mark Thompson 0001, Todor P. Stefanov, Andy D. Pimentel, Simon Polstra, Raj Bose, Claudiu Zissulescu, Ed F. Deprettere |
DAC | 4 |
| 2008 | Editorial
Mladen Berekovic, Andy D. Pimentel, Timo Hämäläinen 0001 |
J. Syst. Archit. | 2 |
| 2007 | Static priority scheduling of event-triggered real-time embedded systems
Cagkan Erbas, Andy D. Pimentel, Selin Cerav-Erbas |
Formal Methods Syst. Des. | 2 |
| 2007 | Static priority scheduling of event-triggered real-time embedded systems
Cagkan Erbas, Andy D. Pimentel, Selin Cerav-Erbas |
Formal Methods Syst. Des. | 2 |
| 2007 | Editorial
Jarmo Takala, Timo Hämäläinen 0001, Andy D. Pimentel, Stamatis Vassiliadis |
J. Syst. Archit. | 3 |
| 2006 | A Systematic Approach to Exploring Embedded System Architectures at Multiple Abstraction LevelsabstractThe sheer complexity of today's embedded systems forces designers to start with modeling and simulating system components and their interactions in the very early design stages. It is therefore imperative to have good tools for exploring a wide range of design choices, especially during the early design stages, where the design space is at its largest. This paper presents an overview of the Sesame framework, which provides high-level modeling and simulation methods and tools for system-level performance evaluation and exploration of heterogeneous embedded systems. More specifically, we describe Sesame's modeling methodology and trajectory. It takes a designer systematically along the path from selecting candidate architectures, using analytical modeling and multiobjective optimization, to simulating these candidate architectures with our system-level simulation environment. This simulation environment subsequently allows for architectural exploration at different levels of abstraction while maintaining high-level and architecture-independent application specifications. We illustrate all these aspects using a case study in which we traverse Sesame's exploration trajectory for a motion-JPEG encoder application. Andy D. Pimentel, Cagkan Erbas, Simon Polstra |
IEEE Trans. Computers | 1 |
| 2006 | Multiobjective optimization and evolutionary algorithms for the application mapping problem in multiprocessor system-on-chip designabstractSesame is a software framework that aims at developing a modeling and simulation environment for the efficient design space exploration of heterogeneous embedded systems. Since Sesame recognizes separate application and architecture models within a single system simulation, it needs an explicit mapping step to relate these models for cosimulation. The design tradeoffs during the mapping stage, namely, the processing time, power consumption, and architecture cost, are captured by a multiobjective nonlinear mixed integer program. This paper aims at investigating the performance of multiobjective evolutionary algorithms (MOEAs) on solving large instances of the mapping problem. With two comparative case studies, it is shown that MOEAs provide the designer with a highly accurate set of solutions in a reasonable amount of time. Additionally, analyses for different crossover types, mutation usage, and repair strategies for the purpose of constraints handling are carried out. Finally, a number of multiobjective optimization results are simulated for verification. Cagkan Erbas, Selin Cerav-Erbas, Andy D. Pimentel |
IEEE Trans. Evol. Comput. | 3 |
| 2004 | Static priority scheduling of event triggered real time embedded systemsabstractReal-time embedded systems are often specified as a collection of independent tasks, each generating a sequence of event-triggered code blocks, and the scheduling in this domain tries to find an execution order which satisfies all real-time constraints. Within the context of recurring real-time tasks, all previous work either allowed preemptions, or only considered dynamic scheduling, and generally had exponential complexity. However for many embedded systems running on limited resources, preemptive scheduling may be very costly due to high context switching and memory overheads, and dynamic scheduling can be less desirable due to high CPU overhead. In this paper we study static priority scheduling of recurring real-time tasks. We focus on the non-preemptive uniprocessor case and obtain schedule- theoretic results for this case. To this end, we derive a sufficient (albeit not necessary) condition for schedulability under static priority scheduling and show that this condition can be efficiently tested in practice. The latter is demonstrated with examples, where in each case, an optimal solution for a given problem specification is obtained within reasonable time, by first detecting good candidates using meta-heuristics, and then by testing them for schedulability. Cagkan Erbas, Selin C. Erbas, Andy D. Pimentel |
MEMOCODE | 3 |
| 2003 | An IDF-based trace transformation method for communication refinementabstractIn the Artemis project, design space exploration of embedded systems is provided by modeling application behavior and architectural performance constraints separately. Mapping an application model onto an architecture model is performed using trace-driven co-simulation, where event traces generated by an application model drive the underlying architecture model. The abstract communication events from the application model may, however, not match the architecture-level communication primitives. This paper presents a trace transformation method, which is based on integer-controlled data-flow models, to perform communication refinement of application-level events. We discuss the proposed method in the context of our prototype modeling and simulation environment. Moreover, using several examples and a case study, we demonstrate that our method allows for efficient exploration of different communication behaviors at architecture level without affecting the application model. Andy D. Pimentel, Cagkan Erbas |
DAC | 1 |
| 1999 | Evaluation of LH*LH for a Multicomputer Architecture
Andy D. Pimentel, Louis O. Hertzberger |
Euro-Par | 1 |
| 1999 | TriMedia CPU64 ArchitectureabstractWe present a new VLIW core as a successor to the TriMedia TM1000. The processor is targeted for embedded use in media-processing devices like DTVs and set-top boxes. Intended as a core, its design must be supplemented with on-chip co-processors to obtain a cost-effective system. Good performance is obtained through a uniform 64-bit 5 issue-slot VLIW design, supporting subword parallelism with an extensive instruction set optimized with respect to media-processing. Multi-slot 'super-ops' allow powerful multi-argument and multi-result operations. As an example, the IDCT algorithm shows a very low instruction count in comparison with other processors. To achieve good performance, critical sections in the application program source code need to be rewritten with vector data types and function calls for media operations. Benchmarking with several media applications was used to tune the instruction set and study cache behaviour. This resulted in a VLIW architecture with wide data paths and relatively simple CPU control. Jos T. J. van Eijndhoven, Kees A. Vissers, Evert-Jan D. Pol, P. Struik, R. H. J. Bloks, Pieter van der Wolf, Harald P. E. Vranken, Frans Sijstermans, M. J. A. Tromp, Andy D. Pimentel |
ICCD | 10 |
| 1996 | Evaluation of a Mesh of Clos wormhole networkabstractThe pursuit of high connectivity in network design for multicomputers is often complicated by wiring constraints, resulting in a trade-off between efficiency and realizability. The Mesh of Clos topology addresses this trade-off by combining a multistage network with a mesh network. In this paper, a simulation study is presented in order to evaluate wormhole-routed Mesh of Clos communication networks. It is shown that this type of network can substantially reduce contention in comparison with flatter mesh networks. Furthermore, we found that increasing the number of flit-buffers on router devices does not necessarily lead to improved communication performance. For some application loads it may even result in a degradation of performance. Andy D. Pimentel, Louis O. Hertzberger |
HiPC | 1 |