EDBT 2026 Demo / reviewers in the wild / expert
Tiago Carvalho 0001
dblp:07/10699-1
· DBLP profile ↗
16ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0002-5826-7643ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Lightweight Performance Monitoring of Real-Time Applications in RISC-V PlatformsabstractAs RISC-V platforms become a target for real-time systems it is crucial to ensure and effect performance analysis to make sure that these systems meet the respective time constraints while also perform reliably. To achieve these goals, performance monitoring becomes a critical aspect, especially when considering resource-constrained environments where efficient resource usage is required. This paper focuses on the study and development of a solution to simplify the interaction with machine-level privileged counters and registers, considering two essential non-functional requirements (NFRs): low-overhead access to performance metrics and low memory API usage. The provided solution allows developers to retrieve and analyse performance data directly from user-level space with a simplified interface, while providing feedback for application optimization, isolation, and improved system reliability. The demonstrated results showcase how our approach meets the two NFRs and its potential in terms of customization for the target platform. Nuno Soares, Tiago Carvalho 0001, Luís Miguel Pinho |
DSD | 2 |
| 2025 | Supporting Soft Real-Time Tasks in Zephyr With Constant Bandwidth ServersabstractThe Constant Bandwidth Server (CBS) is a mechanism used in real-time systems to enable aperiodic soft realtime tasks with unknown execution parameters to run under a dynamic scheduling policy such as Earliest Deadline First (EDF), while still ensuring schedulability by using a bandwidth reservation strategy. This paper proposes an approach to extend the Zephyr open-source real-time operating system, currently maintained by the Linux Foundation, to support aperiodic tasks with CBS. The paper provides the proposed architecture and the design and implementation of the CBS mechanisms in the operating system, which are then evaluated in two test cases in an embedded platform. Alexander Paschoaletto, Paulo Baltarejo Sousa, Luís Miguel Pinho, Tiago Carvalho 0001 |
ISORC | 4 |
| 2025 | Energy Monitoring Systems Analysis and Development: A Case Study for Graph-Based Modelling
Tiago Carvalho 0001, Sebastian Reiter 0003, Luís Miguel Pinho |
MODELSWARD | 1 |
| 2024 | Evaluation of Heuristic Task-to-Thread Mapping Using Static and Dynamic Approaches
Tiago Carvalho 0001, Luís Miguel Pinho, Sara Royuela |
JSSPP | 2 |
| 2024 | Foundations for a Rust-Like Borrow Checker for CabstractMemory safety issues in C are the origin of various vulnerabilities that can compromise a program's correctness or safety from attacks. We propose a different approach to tackle memory safety, the replication of Rust's Mid-level Intermediate Representation (MIR) Borrow Checker, through the usage of static analysis and successive source-to-source code transformations, to be composed upstream of the compiler, thus ensuring maximal compatibility with most build systems. This allows us to approximate a subset of C to Rust's core concepts, applying the memory safety guarantees of the rustc compiler to C. In this work, we present a survey of Rust's efforts towards ensuring memory safety, and describe the theoretical basis for a C borrow checker, alongside a proof-of-concept that was developed to demonstrate its potential. This prototype correctly identified violations of the ownership and aliasing rules, and accurately reported each error with a level of detail comparable to that of the rustc compiler. João Bispo, Tiago Carvalho 0001 |
LCTES | 3 |
| 2024 | Time-predictable task-to-thread mapping in multi-core processorsabstractThe performance of time-predictable systems can be improved in multi-core processors using parallel programming models (e.g., OpenMP). However, schedulability analysis of parallel applications is a big challenge due to their sophisticated structure. The common drawbacks of current task-to-thread mapping approaches in OpenMP are that they (i) utilize a global queue in the mapping process, which may increase contention, (ii) do not apply heuristic techniques, which may reduce the predictability and performance of the system, and (iii) use basic analytical techniques, which may cause notable pessimism in the temporal conditions. Accordingly, this paper proposes a task-to-thread mapping method in multi-core processors based on the OpenMP framework. The mapping process is carried out through two phases: allocation and dispatching. Each thread has an allocation queue in order to minimize contention, and the allocation and dispatching processes are performed using several heuristic algorithms to enhance predictability. In the allocation phase, each task-part from the OpenMP DAG is allocated to one of the allocation queues, which includes both sibling and child task-parts. A suitable thread (i.e., allocation queue) is selected using one of the suggested heuristic allocation algorithms. In the dispatching phase, when a thread is idle, a task-part is selected from its allocation queue using one of the suggested heuristic dispatching algorithms and then dispatched to and executed by the thread. The performance of the proposed method is evaluated under different conditions (e.g., varying the number of tasks and the number of threads) in terms of application response time and overhead of the mapping process. The simulation results show that the proposed method surpasses the other methods, especially in the scenario that includes overhead of the mapping. In addition, a prototype implementation of the main heuristics is evaluated using two kernels from real-world applications, showing that the methods work better than LLVM's default scheduler in most of the configurations. Sara Royuela, Luís Miguel Pinho, Tiago Carvalho 0001, Eduardo Quiñones |
J. Syst. Archit. | 4 |
| 2023 | Framework for the Analysis and Configuration of Real-Time OpenMP ApplicationsabstractHigh-performance cyber-physical applications impose several requirements with respect to performance, functional correctness and non-functional aspects. Nowadays, the design of these systems usually follows a model-driven approach, where models generate executable applications, usually with an automated approach. As these applications might execute in different parallel environments, their behavior becomes very hard to predict, and making the verification of non-functional requirements complicated. In this regard, it is crucial to analyse and understand the impact that the mapping and scheduling of computation have on the real-time response of the applications. In fact, different strategies in these steps of the parallel orchestration may produce significantly different interference, leading to different timing behaviour.Tuning the application parameters and the system configuration proves to be one of the most fitting solutions. The design space can however be very cumbersome for a developer to test manually all combinations of application and system configurations. This paper presents a methodology and a toolset to profile, analyse, and configure the timing behaviour of high-performance cyber-physical applications and the target platforms. The methodology leverages on the possibility of generating a task dependency graph representing the parallel computation to evaluate, through measurements, different mapping configurations and select the one that minimizes response time. Tiago Carvalho 0001, Luís Miguel Pinho, Sara Royuela, Adrian Munera, Eduardo Quiñones |
INDIN | 1 |
| 2022 | Heuristic-based Task-to-Thread Mapping in Multi-Core ProcessorsabstractOpenMP can be used in real-time applications to enhance system performance. However, predictability of OpenMP applications is still a challenge. This paper investigates heuristics for the mapping of OpenMP task graphs in underlying threads, for the development of time-predictable OpenMP programs. These approaches are based on a global scheduling queue, as well as per-thread allocation queues. The proposed method is divided into scheduling and allocation phases. In the former phase, OpenMP task-parts are discovered from OpenMP graph and placed in the scheduling queue. Afterwards, an appropriate allocation queue is selected for each task-part using four heuristic algorithms. In the latter phase, the best task-part is selected from the allocation queue to be allocated to and executed by an idle thread. Preliminary simulation results show that the new method overcomes BFS and WFS in terms of scheduling time and idle time. Sara Royuela, Luís Miguel Pinho, Tiago Carvalho 0001, Eduardo Quiñones |
ETFA | 4 |
| 2022 | Configuration of Parallel Real-Time Applications on Multi-Core ProcessorsabstractParallel programming models (e.g., OpenMP) are more and more used to improve the performance of real-time applications in modern processors. Nevertheless, these processors have complex architectures, being very difficult to understand their timing behavior. The main challenge with most of existing works is that they apply static timing analysis for simpler models or measurement-based analysis using traditional platforms (e.g., single core) or considering only sequential algorithms. How to provide an efficient configuration for the allocation of the parallel program in the computing units of the processor is still an open challenge. This paper studies the problem of performing timing analysis on complex multi-core platforms, pointing out a methodology to understand the applications’ timing behavior, and guide the configuration of the platform. As an example, the paper uses an OpenMP-based program of the Heat benchmark on a NVIDIA Jetson AGX Xavier. The main objectives are to analyze the execution time of OpenMP tasks, specify the best configuration of OpenMP directives, identify critical tasks, and discuss the predictability of the system/application. A Linux perf based measurement tool, which has been extended by our team, is applied to measure each task across multiple executions in terms of total CPU cycles, the number of cache accesses, and the number of cache misses at different cache levels, including L1, L2 and L3. The evaluation process is performed using the measurement of the performance metrics by our tool to study the predictability of the system/application. Tiago Carvalho 0001, Luís Miguel Pinho |
INDIN | 2 |
| 2021 | An ensemble of autonomous auto-encoders for human activity recognitionabstractHuman Activity Recognition is focused on the use of sensing technology to classify human activities and to infer human behavior. While traditional machine learning approaches use hand-crafted features to train their models, recent advancements in neural networks allow for automatic feature extraction. Auto-encoders are a type of neural network that can learn complex representations of the data and are commonly used for anomaly detection. In this work we propose a novel multi-class algorithm which consists of an ensemble of auto-encoders where each auto-encoder is associated with a unique class. We compared the proposed approach with other state-of-the-art approaches in the context of human activity recognition. Experimental results show that ensembles of auto-encoders can be efficient, robust and competitive. Moreover, this modular classifier structure allows for more flexible models. For example, the extension of the number of classes, by the inclusion of new auto-encoders, without the necessity to retrain the whole model. Kemilly Dearo Garcia, Cláudio Rebelo de Sá, Mannes Poel, Tiago Carvalho 0001, João Mendes-Moreira 0001, João M. P. Cardoso, André C. P. L. F. de Carvalho, Joost N. Kok |
Neurocomputing | 4 |
| 2018 | Aspect composition for multiple target languages using LARA
Pedro Pinto 0002, Tiago Carvalho 0001, João Bispo, Miguel António Ramalho, João M. P. Cardoso |
Comput. Lang. Syst. Struct. | 2 |
| 2016 | Performance-driven instrumentation and mapping strategies using the LARA aspect-oriented programming approachabstractSummary The development of applications for high‐performance embedded systems is a long and error‐prone process because in addition to the required functionality, developers must consider various and often conflicting nonfunctional requirements such as performance and/or energy efficiency. The complexity of this process is further exacerbated by the multitude of target architectures and mapping tools. This article describes LARA, an aspect‐oriented programming language that allows programmers to convey domain‐specific knowledge and nonfunctional requirements to a toolchain composed of source‐to‐source transformers, compiler optimizers, and mapping/synthesis tools. LARA is sufficiently flexible to target different tools and host languages while also allowing the specification of compilation strategies to enable efficient generation of software code and hardware cores (using hardware description languages) for hybrid target architectures – a unique feature to the best of our knowledge not found in any other aspect‐oriented programming language. A key feature of LARA is its ability to deal with different models of join points, actions, and attributes. In this article, we describe the LARA approach and evaluate its impact on code instrumentation and analysis and on selecting critical code sections to be migrated to hardware accelerators for two embedded applications from industry. Copyright © 2014 John Wiley & Sons, Ltd. João M. P. Cardoso, José Gabriel F. Coutinho, Tiago Carvalho 0001, Pedro C. Diniz, Zlatko Petrov, Wayne Luk, Fernando M. Gonçalves |
Softw. Pract. Exp. | 3 |
| 2015 | Programming Strategies for Contextual Runtime SpecializationabstractRuntime adaptability is expected to adjust the application and the mapping of computations according to usage contexts, operating environments, resources availability, etc. However, extending applications with adaptive features can be a complex task, especially due to the current lack of programming models and compiler support. One of the run-time adaptability possibilities is the use of specialized code according to data workloads and environments. Traditional approaches use multiple code versions generated offline and, during runtime, a strategy is responsible to select a code version. Moving code generation to runtime can achieve important improvements but may impose unacceptable overhead. This paper presents an aspect-oriented programming approach for runtime adaptability. We focus on a separation of concerns (strategies vs. application) promoted by a domain-specific language for programming runtime strategies. Our strategies allow runtime specialization based on contextual information. We use a template-based runtime code generation approach to achieve program specialization. We demonstrate our approach with examples from image processing, which depict the benefits of runtime specialization and illustrate how several factors need to be considered to efficiently adapt the application. Tiago Carvalho 0001, Pedro Pinto 0002, João M. P. Cardoso |
SCOPES | 1 |
| 2013 | The MATISSE MATLAB compilerabstractThis paper describes MATISSE, a MATLAB to C compiler targeting embedded systems that is based on Strategic and Aspect-Oriented Programming concepts. MATISSE takes as input: (1) MATLAB code and (2) LARA aspects related to types and shapes, code insertion/removal, and specialization based directives defining default variable values. In this paper we also illustrate the use of MATISSE in leveraging data types and shapes to generate customized C code suitable for high-level hardware synthesis tools. The preliminary experimental results presented here reveal the described approach to yield performance results for the resulting hardware and software references implementations that are comparable in terms of performance with hand-crafted solutions but derived automatically at a fraction of the cost. João Bispo, Pedro Pinto 0002, Ricardo Nobre, Tiago Carvalho 0001, João M. P. Cardoso, Pedro C. Diniz |
INDIN | 4 |
| 2013 | Enriching MATLAB with aspect-oriented features for developing embedded systems
João M. P. Cardoso, João M. Fernandes 0001, Miguel P. Monteiro 0001, Tiago Carvalho 0001, Ricardo Nobre |
J. Syst. Archit. | 4 |
| 2012 | Controlling Hardware Synthesis with AspectsabstractThe synthesis and mapping of applications to configurable embedded systems is a notoriously hard process. Tools have a wide range of parameters, which interact in very unpredictable ways, thus creating a large and complex design space. When exploring this space, designers must understand the interfaces to the various tools and apply, often manually, a sequence of tool-specific transformations making this an extremely cumbersome and error-prone process. This paper describes the use of aspect-oriented techniques for capturing synthesis strategies for tuning the performance of applications' kernels. We illustrate the use of this approach when designing application-specific architectures generated by a high-level synthesis tool. The results highlight the impact of the various strategies when targeting custom hardware and expose the difficulties in devising these strategies. João M. P. Cardoso, Tiago Carvalho 0001, José Gabriel F. Coutinho, Pedro C. Diniz, Zlatko Petrov, Wayne Luk |
DSD | 2 |