Sandro Pinto 0001

dblp:170/0331 · DBLP profile ↗
← Back
46ranked-venue papers
7as first author
31since 2021 · last 2026
0000-0003-4580-7484ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Security and privacy · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DAAT‑MCS: A Framework to Deploy, Analyse, and Auto‑Tune Interference Mitigation Techniques in Mixed‑Criticality Systems
abstract
The consolidation of workloads with different criticality levels onto shared multi-core platforms is widely adopted to meet stringent SWaP-C constraints, giving rise to mixed-criticality systems. This consolidation is often enabled by static partitioning hypervisors, which provide strong spatial isolation via static assignment of resources such as cores, memory regions, and devices. However, temporal isolation remains challenging because key microarchitectural resources (e.g., the last-level cache and the memory subsystem) are still shared and can induce contention-driven slowdowns and loss of predictability. Cache coloring and related mitigation techniques can reduce this interference, but their configuration is still typically performed offline via manual profiling and trial-and-error, which is time-consuming, error-prone, and highly workload- and platform-dependent. The remaining gap is therefore not the lack of mitigation mechanisms, but the lack of a systematic way to derive a deployable partition for a given workload mix under explicit timing requirements. This paper presents DAAT-MCS, an end-to-end framework that automates cache-color assignment for static partitioning hypervisors. It consists of two components: (i) the DAAT-MCS profiler, which systematically deploys and measures a bounded set of cache partitions on real hardware to quantify interference effects in mixed-criticality consolidations; and (ii) the DAAT-MCS tuner, which uses these measurements to automatically derive the set of system-wide cache partitions that satisfy user-defined per-core performance-degradation tolerances. Together, these components turn cache-color tuning from ad-hoc trial-and-error into a repeatable, measurement-driven integration step with explicit acceptance criteria. We evaluate DAAT-MCS on an ARMv8-A Xilinx UltraScale+ ZCU104 platform across 264 workload combinations and 7 cache-partitioning configurations (1848 setups, >60 h of automated profiling). For 28 tuner settings on a subset of 154 out of the 1848 consolidations (one memory-intensive benchmark co-scheduled with 22 co-runners across 7 cache partitions), totaling 4,312 tuner runs, DAAT-MCS finds no acceptable cache partition in 84.53% of cases and returns at least one deployable configuration in the remaining 15.47%, reducing the integration search space whenever feasibility exists and avoiding manual trial-and-error. All artifacts, including code, raw traces, and reproduction infrastructure, are publicly available.
Sandro Pinto 0001
ECRTS2
2026 Classifier-Prototype Alignment with Dual-View Consistency for Long-Tailed Recognition
Ziyao Meng 0001, Jian Li 0080, Xue Gu, Adriano Tavares, Sandro Pinto 0001, Hao Xu 0012
ICIC (20)5
2026 ISense: An assessment instrument that predicts self-regulated learning using Mobile sensing
abstract
Self-regulated learning significantly influences students’ academic behavior and performance. Traditional SRL assessment heavily relies on subjective and self-evaluation methods, which are susceptible to personal biases and cognitive limitations. In this study, we leverage mobile devices with GPS and sensors to collect self-reported and passive mobile sensing data from 211 college students over a year. Our aim is to conduct a passive assessment of self-regulated learning. To achieve this, we apply four deep learning models to analyze behavioral features associated with self-regulated learning using students’ life record data. Our findings demonstrate a precise assessment of student self-regulated learning, encompassing the following subscales. Environment structuring (MAE = 2.88, r = 0.54) represents participation in the organization and construction of the learning environment. Time management (MAE = 2.82, r = 0.57) represents planning study time and balancing study with other activities; Help seeking (MAE = 4.86, r = 0.53) represents the tendency and frequency of students to seek help during the learning process. Our study helps inspire new forms of education to assess self-regulated learning and paves the way for individualized interventions in future studies. • We develop and implement the iSense system for SRL assessment. • Collecting mobile sensing and self-reported data from 211 students over a year. • Establishing an SRL assessment model using mobile sensing data and deep learning. • We observe that SRL highly correlates with app usage patterns and sensing data. • Our study can achieve longitudinal monitoring and early intervention of SRL.
Tongyu Zhao, Jiaying Gao, Yatong Zu, Adriano Tavares, Tiago Gomes 0001, Sandro Pinto 0001, Hao Xu 0012
Int. J. Hum. Comput. Stud.7
2026 Beneath the fabric: A survey on threats and opportunities in reconfigurable systems
Sérgio Pereira, Tiago Gomes 0001, Jorge Cabral 0001, Sandro Pinto 0001
J. Syst. Archit.4
2025 Passive Behavioral Sensing: Using Within-Person Variability Features from Mobile Sensing to Assess Self-Regulated Learning
Tongyu Zhao, Jiaying Gao, Yatong Zu, Adriano Tavares, Tiago Gomes 0001, Sandro Pinto 0001, Hao Xu 0012
CogSci7
2025 A preliminary study to Assess Self-regulated learning and Academic Emotional Regulation of College Students Using Smartphones
Tongyu Zhao, Jiaying Gao, Yatong Zu, Adriano Tavares, Tiago Gomes 0001, Sandro Pinto 0001, Hao Xu 0012
CogSci7
2025 Enhancing Educator Support in MOOC Forums: A Multi-Task Model for Detecting Learning Confusion
Tongyu Zhao, Jiaying Gao, Yatong Zu, Adriano Tavares, Tiago Gomes 0001, Sandro Pinto 0001, Hao Xu 0012
CogSci6
2025 RAOCSL: A BERT-Based Strategy for Identifying Learner Confusion under Class Imbalance
abstract
Understanding and identifying the nature of learner confusion is important for online learning platforms. In this study, we address this problem by analyzing forum posts from large-scale online courses. However, due to the large volume of comments and frequent interactions, confusion posts are often overlooked. Existing methods and models, while capable of detecting confusion, typically rely on linguistic features of posts and community factors (e.g. votes, views) but ignore personalized contexts, such as the specific causes and types of confusion. To address this problem, we create the first deep learning dataset focused on confusion types and develop a BERT-based network to model personalized features and identify confusion types. Considering the highly imbalanced distribution of different types of confusion, we further design a novel loss function that adaptively optimizes the training weights for each type. Our method’s effectiveness is confirmed through extensive experimentation.
Tongyu Zhao, Jiaying Gao, Yatong Zu, Adriano Tavares, Tiago Gomes 0001, Sandro Pinto 0001, Hao Xu 0012
ICASSP7
2025 Bridging the Interoperability Gaps Among Trusted Architectures in MCUs
Sandro Pinto 0001, Daniel Oliveira 0003, Michele Grisafi, Emanuele Beozzo, Bruno Crispo
ICICS (2)1
2025 Beyond the Bermuda Triangle of Contention: IOMMU Interference in Mixed Criticality Systems
José Martins 0004, Sandro Pinto 0001
RTCSA3
2025 SecureQNN: Shielding the Intellectual Property of QNNs in TinyML Systems
abstract
Building accurate Machine Learning (ML) models requires substantial expertise and large-scale datasets typically only available in big data companies. These companies have been selling their models as Machine Learning as a Service. However, concerns about data privacy and the appliance of ML in mission-critical scenarios are forcing ML computation to move from the cloud to the deep edge, near sensor data. If edge ML increases user data privacy and makes decision latency predictable, deploying proprietary ML models on untrusted edge devices may harm the intellectual property of Service Providers. A natural response to this issue comes from Trusted Execution Environments (TEEs), which provide hardware-based security. However, adapting ML computation to the constraints of TEEs remains an open challenge. In this article, we propose SecureQNN, a framework that leverages state-of-the-art TEE technology (TrustZone-M) available in Arm Cortex-M microcontrollers to increase the protection of Quantized Neural Networks (QNNs) against unauthorized replication. SecureQNN evaluates which layers lessen the effort of an attacker when building a surrogate QNN with the same accuracy. The layers making the attacker spend less training epochs than the owner training from scratch are stored and executed in the secure-world of the TEE. Experiments demonstrate that SecureQNN undermines the cost-benefit of unauthorized QNN replication by isolating 51%–65% of the model size while incurring a worst-case decision latency overhead of only 0.06%. Although SecureQNN has broader applicability, the negligible impact on decision latency and its deterministic behavior highlight the suitability of SecureQNN for mission-critical and real-time applications.
Tiago Gomes 0001, Sandro Pinto 0001
IEEE Internet Things J.3
2024 MCU-Wide Timing Side Channels and Their Detection
abstract
Microarchitectural timing side channels have been thoroughly investigated as a security threat in hardware designs featuring shared buffers (e.g., caches) and/or parallelism between attacker and victim task execution. However, contradicting common intuitions, recent activities demonstrate that this threat is real even in microcontroller SoCs without such features. In this paper, we describe SoC-wide timing side channels previously neglected by security analysis and present a new formal method to close this gap. In a case study on the RISC-V Pulpissimo SoC, our method detected a vulnerability to a previously unknown attack variant that allows an attacker to obtain information about a victim's memory access behavior. After implementing a conservative fix, we were able to verify that the SoC is now secure w.r.t. the considered class of timing side channels.
Johannes Müller 0006, Anna Lena Duque Antón, Lucas Deutschmann, Dino Mehmedagic, Cristiano Rodrigues, Daniel Oliveira 0003, Mohammad Rahmani Fadiheh, Keerthikumara Devarajegowda, Sandro Pinto 0001, Dominik Stoffel, Wolfgang Kunz
DAC9
2024 Shared Resource Contention in MCUs: A Reality Check and the Quest for Timeliness
abstract
Arm Cortex-M processors are the most widely used 32-bit microcontrollers among embedded and Internet-of-Things devices. Despite the widespread usage, there has been little effort in summarizing their hardware security features, characterizing the limitations and vulnerabilities of their hardware and software stack, and systematizing the research on securing these systems. The goals and contributions of this paper are multi-fold. First, we analyze the hardware security limitations and issues of Cortex-M systems. Second, we conducted a deep study of the software stack designed for Cortex-M and revealed its limitations, which is accompanied by an empirical analysis of 1,797 real-world firmware. Third, we categorize the reported bugs in Cortex-M software systems. Finally, we systematize the efforts that aim at securing Cortex-M systems and evaluate them in terms of the protections they offer, runtime performance, required hardware features, etc. Based on the insights, we develop a set of recommendations for the research community and MCU software developers.
Daniel Oliveira 0003, Weifan Chen 0003, Sandro Pinto 0001, Renato Mancuso 0001
ECRTS3
2024 David and Goliath: An Empirical Evaluation of Attacks and Defenses for QNNs at the Deep Edge
abstract
Machine learning (ML) is shifting from the cloud to the edge. Edge computing reduces the surface exposing private data and enables reliable throughput guarantees in real-time applications. Of the panoply of devices deployed at the edge, resource-constrained microcontrollers (MCUs), e.g., Arm Cortex-M, are more prevalent, orders of magnitude cheaper, and less power-hungry than application processors (APUs) or graphical processing units (GPUs). Thus, enabling intelligence at the so-called deep/extreme edge is the zeitgeist, with researchers focusing on unveiling novel approaches to deploy artificial neural networks (ANN) on these constrained devices. Quantization is a well-established technique that has proved effective, i.e., negligible impact on accuracy, in enabling the deployment of neural networks on MCUs; however, it is still an open question to understand the robustness of quantized neural networks (QNNs) in the face of well-known adversarial examples. To fill this gap, we empirically evaluate the effectiveness of attacks and defenses from (full-precision) ANNs on (constrained) QNNs. Our evaluation suite includes three QNNs targeting TinyML applications, ten attacks, and six defenses. With this study, we draw a set of interesting findings. First, quantization increases the point distance to the decision boundary and leads the gradient estimated by some attacks to explode or vanish. Second, quantization can act as a noise attenuator or amplifier, depending on the noise magnitude, and causes gradient misalignment. Regarding adversarial defenses, we conclude that input pre-processing defenses show impressive results on small perturbations; however, they fall short as the perturbation increases. At the same time, train-based defenses increase adversarial robustness by increasing the average point distance to the decision boundary, which holds even after quantization. However, we argue that train-based defenses still need to smooth the quantization-shift and gradient misalignment phenomenons to counteract adversarial example transferability to QNNs. All artifacts are open-sourced to enable independent validation of results and encourage further exploration of the robustness of QNNs.
Sandro Pinto 0001
EuroS&P2
2024 WiP Paper: BUSted Second Stop! A First Step for Breaking Cryptographic Applications on MCU-based IoT Devices
André Barbosa, Cristiano Rodrigues, Tiago Gomes 0001, Sandro Pinto 0001
EWSN4
2024 BUSted!!! Microarchitectural Side-Channel Attacks on the MCU Bus Interconnect
abstract
Spectre and Meltdown have pushed the research community toward an otherwise-unavailable understanding of the security implications of processors’ microarchitecture. Notwithstanding, research efforts have concentrated on high-end processors (e.g., Intel, AMD, Arm Cortex-A), and very little has been done for microcontrollers (MCU) that power billions of small embedded and IoT devices. In this paper, we present BUSted. BUSted is a novel side-channel attack that explores the side effects of the MCU bus interconnect arbitration logic to bypass security guarantees enforced by memory protection primitives. Side-channel attacks on MCUs pose incremental and unforeseen challenges, which are strictly tied to the resource-constrained nature of these systems (e.g., single-core CPU, stateless bus). We devise a unique approach that relies on the concept of hardware gadgets. We present practical attacks on state-of-the-art Armv8-M MCUs with TrustZone-M, running the Trusted Firmware-M (TF-M). In contrast to the Nemesis attack, our attack is practical on Arm Cortex-M MCUs, and our findings suggest that it can scale across the full MCU spectrum.
Cristiano Rodrigues, Daniel Oliveira 0003, Sandro Pinto 0001
SP3
2024 GEML: a graph-enhanced pre-trained language model framework for text classification via mutual learning
Rui Song 0008, Sandro Pinto 0001, Tiago Gomes 0001, Adriano Tavares, Hao Xu 0012
Appl. Intell.3
2024 ChamelIoT: a tightly- and loosely-coupled hardware-assisted OS framework for low-end IoT devices
abstract
Abstract The evergrowing Internet of Things (IoT) ecosystem continues to impose new requirements and constraints on every device. At the edge, low-end devices are getting pressured by increasing workloads and stricter timing deadlines while simultaneously are desired to minimize their power consumption, form factor, and memory footprint. Field-Programmable Gate Arrays (FPGAs) emerge as a possible solution for the increasing demands of the IoT. Reconfigurable IoT platforms enable the offloading of software tasks to hardware, enhancing their performance and determinism. This paper presents ChamelIoT, an agnostic hardware operating systems (OSes) framework for reconfigurable IoT devices. The framework provides hardware acceleration for kernel services of different IoT OSes by leveraging the RISC-V open-source instruction set architecture (ISA). The ChamelIoT hardware accelerator can be deployed in a tightly- or loosely-coupled approach and implements the following kernel services: thread management, scheduling, synchronization mechanisms, and inter-process communication (IPC). ChamelIoT allows developers to run unmodified applications of three well-established OSes, RIOT, Zephyr, and FreeRTOS. The experiments conducted on both coupling approaches consisted of microbenchmarks to measure the API latency, the Thread Metric benchmark suite to evaluated the system performance, and tests to the FPGA resource consumption. The results show that the latency can be reduced up to 92.65% and 89.14% for the tightly- and loosely-coupled approaches, respectively, the jitter removed, and the execution performance increased by 199.49% and 184.85% for both approaches.
Tiago Gomes 0001, Mongkol Ekpanyapong, Adriano Tavares, Sandro Pinto 0001
Real Time Syst.5
2024 A Heterogeneous RISC-V Based SoC for Secure Nano-UAV Navigation
abstract
The rapid advancement of energy-efficient parallel ultra-low-power (ULP)$\mu$controllers units (MCUs) is enabling the development of autonomous nano-sized unmanned aerial vehicles (nano-UAVs). These sub-10cm drones represent the next generation of unobtrusive robotic helpers and ubiquitous smart sensors. However, nano-UAVs face significant power and payload constraints while requiring advanced computing capabilities akin to standard drones, including real-time Machine Learning (ML) performance and the safe co-existence of general-purpose and real-time OSs. Although some advanced parallel ULP MCUs offer the necessary ML computing capabilities within the prescribed power limits, they rely on small main memories ($<$1MB) and$\mu$controller-class CPUs with no virtualization or security features, and hence only support simple bare-metal runtimes. In this work, we present Shaheen, a 9mm$^{\textbf{2}}$200mW SoC implemented in 22nm FDX technology. Differently from state-of-the-art MCUs, Shaheen integrates a Linux-capable RV64 core, compliant with the v1.0 ratified Hypervisor extension and equipped with timing channel protection, along with a low-cost and low-power memory controller exposing up to 512MB of off-chip low-cost low-power HyperRAM directly to the CPU. At the same time, it integrates a fully programmable energy-and area-efficient multi-core cluster of RV32 cores optimized for general-purpose DSP as well as reduced-and mixed-precision ML. To the best of the authors’ knowledge, it is the first silicon prototype of a ULP SoC coupling the RV64 and RV32 cores in a heterogeneous host+accelerator architecture fully based on the RISC-V ISA. We demonstrate the capabilities of the proposed SoC on a wide range of benchmarks relevant to nano-UAV applications including general-purpose DSP as well as inference and online learning of quantized DNNs. The cluster can deliver up to 90GOp/s and up to 1.8TOp/s/W on 2-bit integer kernels and up to 7.9GFLOp/s and up to 150GFLOp/s/W on 16-bit FP kernels.
Luca Valente, Alessandro Nadalini, Asif Veeran, Mattia Sinigaglia, Bruno Sá, Nils Wistoff, Yvan Tortorella, Simone Benatti, Rafail Psiakis, Ari Kulmala, Baker Mohammad, Sandro Pinto 0001, Daniele Palossi, Luca Benini, Davide Rossi 0001
IEEE Trans. Circuits Syst. I Regul. Pap.12
2023 Holistic RISC-V Virtualization: CVA6-based SoC
abstract
This work describes our efforts to provide a holistic hardware RISC-V virtualization SoC based on the CVA6 core. At the core level, we implemented hardware support for virtualization through the ratified Hypervisor instruction set architecture (ISA) extension version 1.0. At the system level, we are working on providing reference open-source IPs for two non-ISA components needed to build a virtualization-aware platform: (i) the advanced interrupt architecture (AIA) to enable hardware support for interrupt virtualization; (ii) the input/output memory management unit (IOMMU) to protect memory accesses from direct memory access (DMA) devices. All these IPs will be open and freely available to the RISC-V community under permissive open-source licenses.
Bruno Sá, Francisco Marques, José Martins 0004, Sandro Pinto 0001
CF5
2023 Feature Visualization and Attribution Analysis of Confusion for Massive Open Online Course
Tongyu Zhao, Jiaying Gao, Jian Li 0080, Yatong Zu, Sandro Pinto 0001, Adriano Tavares, Hao Xu 0012
CogSci6
2023 H2T-FAST: Head-to-Tail Feature Augmentation by Style Transfer for Long-Tailed Recognition
abstract
Deep learning algorithms perform poorly on long-tailed datasets because there is insufficient data in the tail classes to recover its original distribution, resulting in an under-representation of the tail classes in the model. In this work, we propose H2T-FAST, a Head-to-Tail Feature Augmentation method by Style Transfer to improve the performance of the tail. H2T-FAST has the following advantages: (1) It is a fast and universal method that acts on the feature space and so, it can be applied to different backbone networks as well as easily integrated into various imbalanced algorithms with stable performance gains; and (2) it is used only in the training phase and therefore, imposes no additional burden on the deep neural network in the testing phase. In particular, we firstly and randomly select the same number of head samples as the tail ones in each training mini batch. Secondly, the style of the head is transferred to the tail to generate new tail data containing the head style, as a way to increase the number of the tail and get better feature representations. We test our methods on several benchmark vision tasks with state-of-the-art performances.
Ziyao Meng 0001, Xue Gu, Qiang Shen 0005, Adriano Tavares, Sandro Pinto 0001, Hao Xu 0012
ECAI5
2023 Shaheen: An Open, Secure, and Scalable RV64 SoC for Autonomous Nano-UAVs
abstract
Open Source Hardware, the way it should be!
Luca Valente, Asif Veeran, Mattia Sinigaglia, Yvan Tortorella, Alessandro Nadalini, Nils Wistoff, Bruno Sá, Angelo Garofalo, Rafail Psiakis, M. Tolba, Ari Kulmala, Nimisha Limaye, Ozgur Sinanoglu, Sandro Pinto 0001, Daniele Palossi, Luca Benini, Baker Mohammad, Davide Rossi 0001
HCS14
2023 Shedding Light on Static Partitioning Hypervisors for Arm-based Mixed-Criticality Systems
abstract
In this paper, we aim to understand the properties and guarantees of static partitioning hypervisors (SPH) for Arm-based mixed-criticality systems (MCS). To this end, we performed a comprehensive empirical evaluation of popular open-source SPH, i.e., Jailhouse, Xen (Dom0-less), Bao, and seL4 CAmkES VMM, focusing on two key requirements of modern MCS: real-time and safety. The goal of this study is twofold. Firstly, to empower industrial practitioners with hard data to reason about the different trade-offs of SPH. Secondly, we aim to raise awareness of the research and open-source communities to the still open problems in SPH by unveiling new insights regarding lingering weaknesses. All artifacts will be open-sourced to enable independent validation of results and encourage further exploration on SPH.
José Martins 0004, Sandro Pinto 0001
RTAS2
2023 IRQ Coloring and the Subtle Art of Mitigating Interrupt-Generated Interference
abstract
Integrating workloads with differing criticality levels presents a formidable challenge in achieving the stringent spatial and temporal isolation requirements imposed by safety-critical standards such as ISO26262. The shift towards high-performance multicore platforms has been posing increasing issues to the so-called mixed-criticality systems (MCS) due to the reciprocal interference created by consolidated subsystems vying for access to shared (microarchitectural) resources (e.g., caches, bus interconnect, memory controller). The research community has acknowledged all these challenges. Thus, several techniques, such as cache partitioning and memory throttling, have been proposed to mitigate such interference; however, these techniques have some drawbacks and limitations that impact performance, memory footprint, and availability. In this work, we look from a different perspective. Departing from the observation that safety-critical workloads are typically event- and thus interrupt-driven, we mask “colored” interrupts based on the Quality of Service (QoS) assessment, providing fine-grain control to mitigate interference on critical workloads without entirely suspending non-critical workloads. We propose the so-called IRQ coloring technique. We implement and evaluate the IRQ Coloring on a reference high-performance multicore platform, i.e., Xilinx ZCU102. Results demonstrate negligible performance overhead, i.e., < 1% for a 100 microseconds period, and reasonable throughput guarantees for medium-critical workloads. We argue that the IRQ coloring technique presents predictability and intermediate guarantees advantages compared to state-of-art mechanisms.
Luca Cuomo, Daniel Oliveira 0003, Ida Maria Savino, Bruno Morelli, José Martins 0004, Alessandro Biasci, Sandro Pinto 0001
RTCSA8
2023 CVA6 RISC-V Virtualization: Architecture, Microarchitecture, and Design Space Exploration
abstract
Virtualization is a key technology used in a wide range of applications, from cloud computing to embedded systems. Over the last few years, mainstream computer architectures were extended with hardware virtualization support, giving rise to a set of virtualization technologies (e.g., Intel VT and Arm VE) that are now proliferating in modern processors and systems on chip (SoCs). In this article, we describe our work on hardware virtualization support in the RISC-V CVA6 core. Our contribution is multifold and encompasses architecture, microarchitecture, and design space exploration (DSE). In particular, we highlight the design of a set of microarchitectural enhancements [i.e., G-stage translation lookaside buffer (GTLB) and second-level TLB (L2 TLB)] to alleviate the virtualization performance overhead. We also perform a DSE and accompanying postlayout simulations (based on 22-nm FDX technology) to assess performance, power, and area (PPA). Furthermore, we map design variants on a field-programmable gate array (FPGA) platform (Genesys 2) to assess the functional performance–area tradeoff. Based on the DSE, we select an optimal design point for the CVA6 with hardware virtualization support. For this optimal hardware configuration, we collected functional performance results by running the MiBench benchmark on Linux atop Bao hypervisor for a single-core configuration. We observed a performance speedup of up to 16% (approximately 12.5% on average) compared with virtualization-aware nonoptimized design at the minimal cost of 0.78% in area and 0.33% in power. Finally, all works described in this article are publicly available and open-sourced for the community to further evaluate additional design configurations and software stacks.
Bruno Sá, Luca Valente, José Martins 0004, Davide Rossi 0001, Luca Benini, Sandro Pinto 0001
IEEE Trans. Very Large Scale Integr. Syst.6
2022 Agnostic Hardware-Accelerated Operating System for Low-End IoT
abstract
There is increasing pressure to optimize Internet of things (IoT) low-end devices. The ever-growing number of requirements and constraints is pushing towards maximizing performance and real-time, but simultaneously minimizing power consumption, form factor, and memory footprint. This has motivated the adoption of Field-Programmable Gate Array (FPGA) technology to accelerate computing-intensive workloads in hardware. However, and despite the ongoing trend of migrating application-level tasks to hardware, recently, the offload of system software such as operating system (OS) services has received little attention. This paper presents CHAMELIOT, a framework for FPGA-based IoT platforms that provides agnostic hardware acceleration to OS services by leveraging RISC-V technology. CHAMELIOT allows for developers to run unmodified applications in a set of well-established IoT OSes. Currently, the framework has support for RIOT, Zephyr, and FreeRTOS. The evaluation showed that latency and determinism can be enhanced up to 10x while the system’s performance can be increased to nearly 200%. CHAMELIOT will be open-sourced.
Tiago Gomes 0001, Sandro Pinto 0001
RTCSA3
2022 ReZone: Disarming TrustZone with TEE Privilege Reduction
David Cerdeira, José Martins 0004, Nuno Santos 0001, Sandro Pinto 0001
USENIX Security Symposium4
2022 A First Look at RISC-V Virtualization From an Embedded Systems Perspective
abstract
This article describes the first public implementation and evaluation of the latest version of the RISC-V hypervisor extension (H-extension v0.6.1) specification in a Rocket chip core. To perform a meaningful evaluation for modern multi-core embedded and mixed-criticality systems, we have ported Bao, an open-source static partitioning hypervisor, to RISC-V. We have also extended the RISC-V platform-level interrupt controller (PLIC) to enable direct guest interrupt injection with low and deterministic latency and we have enhanced the timer infrastructure to avoid trap and emulation overheads. Experiments were carried out in FireSim, a cycle-accurate, FPGA-accelerated simulator, and the system was also successfully deployed and tested in a Zynq UltraScale+ MPSoC ZCU104. Our hardware implementation was open-sourced and is currently in use by the RISC-V community towards the ratification of the H-extension specification.
Bruno Sá, José Martins 0004, Sandro Pinto 0001
IEEE Trans. Computers3
2022 Shifting Capsule Networks from the Cloud to the Deep Edge
abstract
Capsule networks (CapsNets) are an emerging trend in image processing. In contrast to a convolutional neural network, CapsNets are not vulnerable to object deformation, as the relative spatial information of the objects is preserved across the network. However, their complexity is mainly related to the capsule structure and the dynamic routing mechanism, which makes it almost unreasonable to deploy a CapsNet, in its original form, in a resource-constrained device powered by a small microcontroller (MCU). In an era where intelligence is rapidly shifting from the cloud to the edge, this high complexity imposes serious challenges to the adoption of CapsNets at the very edge. To tackle this issue, we present an API for the execution of quantized CapsNets in Arm Cortex-M and RISC-V MCUs. Our software kernels extend the Arm CMSIS-NN and RISC-V PULP-NN to support capsule operations with 8-bit integers as operands. Along with it, we propose a framework to perform post-training quantization of a CapsNet. Results show a reduction in memory footprint of almost 75%, with accuracy loss ranging from 0.07% to 0.18%. In terms of throughput, our Arm Cortex-M API enables the execution of primary capsule and capsule layers with medium-sized kernels in just 119.94 and 90.60 ms, respectively (STM32H755ZIT6U, Cortex-M7 @ 480 MHz). For the GAP-8 SoC (RISC-V RV32IMCXpulp @ 170 MHz), the latency drops to 7.02 and 38.03 ms, respectively.
Tiago Gomes 0001, Sandro Pinto 0001
ACM Trans. Intell. Syst. Technol.4
2021 Self-secured devices: High performance and secure I/O access in TrustZone-based systems
abstract
Arm TrustZone is a hardware technology that adds significant value to the ongoing security picture. TrustZone-based systems typically consolidate multiple environments into the same platform, requiring resources to be shared among them. Currently, hardware devices on TrustZone-enabled system-on-chip (SoC) solutions can only be configured as secure or non-secure, which means the dual-world concept of TrustZone is not spread to the inner logic of the devices. The traditional passthrough model dictates that both worlds cannot use the same device concurrently. Furthermore, existing shared device access methods have been proven to cause a negative impact on the overall system in terms of security and performance. This work introduces the concept of self-secured devices, a novel approach for shared device access in TrustZone-based architectures. This concept extends the TrustZone dual-world model to the device itself, providing a secure and non-secure logical interface in a single device instance. The solution was deployed and evaluated on the LTZVisor, an open-source and lightweight TrustZone-assisted hypervisor . The obtained results are encouraging, demonstrating that our solution requires only a few additional hardware resources when compared with the native device implementation, while providing a secure solution for device sharing.
Sandro Pinto 0001, Daniel Oliveira 0003, David Cerdeira, Tiago Gomes 0001
J. Syst. Archit.1
2020 SoK: Understanding the Prevailing Security Vulnerabilities in TrustZone-assisted TEE Systems
abstract
Hundreds of millions of mobile devices worldwide rely on Trusted Execution Environments (TEEs) built with Arm TrustZone for the protection of security-critical applications (e.g., DRM) and operating system (OS) components (e.g., Android keystore). TEEs are often assumed to be highly secure; however, over the past years, TEEs have been successfully attacked multiple times, with highly damaging impact across various platforms. Unfortunately, these attacks have been possible by the presence of security flaws in TEE systems. In this paper, we aim to understand which types of vulnerabilities and limitations affect existing TrustZone-assisted TEE systems, what are the main challenges to build them correctly, and what contributions can be borrowed from the research community to overcome them. To this end, we present a security analysis of popular TrustZone-assisted TEE systems (targeting Cortex-A processors) developed by Qualcomm, Trustonic, Huawei, Nvidia, and Linaro. By studying publicly documented exploits and vulnerabilities as well as by reverse engineering the TEE firmware, we identified several critical vulnerabilities across existing systems which makes it legitimate to raise reasonable concerns about the security of commercial TEE implementations.
David Cerdeira, Nuno Santos 0001, Pedro Fonseca 0001, Sandro Pinto 0001
SP4
2019 Towards a Heterogeneous Fault-Tolerance Architecture based on Arm and RISC-V Processors
abstract
Computer systems are permanently present in our daily basis in a wide range of applications. In systems with mixed-criticality requirements, e.g., autonomous driving or aerospace applications, devices are expected to continue operating properly even in the event of a failure. An approach to improve the robustness of the device's operation lies in enabling fault-tolerant mechanisms during the system's design. This article proposes Lock-V, a heterogeneous architecture that explores a Dual-Core Lockstep (DCLS) fault-tolerance technique in two different processing units: a hard-core Arm Cortex-A9 and a soft-core RISC-V-based processor. It resorts a System-on-Chip (SoC) solution with software programmability (available trough the hard-core Arm Cortex-A9) and field-programmable gate array (FPGA) technology, taking advantages from the latter to support the deployment of the RISC-V soft-core along with dedicated hardware accelerators towards the realization of the DCLS.
Cristiano Rodrigues, Ivo Marques, Sandro Pinto 0001, Tiago Gomes 0001, Adriano Tavares
IECON3
2019 Virtualization on TrustZone-Enabled Microcontrollers? Voilà!
abstract
With predictions pointing to more than 20 billion Internet-enabled 'things' by 2020 and much more to come, smart sensor nodes are expected to be predominant in the Internet of Things (IoT) era. As these systems are connected to the Internet and tend to implement an ever-growing number of mixed-criticality features, there is huge pressure for strong isolation to guarantee a reliable, secure, and predictable infrastructure. While virtualization has been a game-changer for consolidation and isolation in mid-to high-end embedded applications, for low-end and low-cost systems it is still in its infancy, and only a limited number of solutions have been proposed so far. This work aims at developing a lightweight hypervisor which provides strong isolation on resource-constrained devices. Our approach leverages TrustZone technology available on modern Arm microcontrollers (TrustZone-M) to implement a predictable virtualization infrastructure for low-end and low-cost systems. Experiments conducted on an Arm Musca-A multi-core platform demonstrate our solution achieves low memory footprint, high efficiency, and strict timing predictability.
Sandro Pinto 0001, Hugo Araújo, Daniel Oliveira 0003, José Martins 0004, Adriano Tavares
RTAS1
2019 Operating Systems for Internet of Things Low-End Devices: Analysis and Benchmarking
abstract
In the era of the Internet of Things (IoT), billions of wirelessly connected embedded devices rapidly became part of our daily lives. As a key tool for each Internet-enabled object, embedded operating systems (OSes) provide a set of services and abstractions which eases the development and speedups the deployment of IoT solutions at scale. This article starts by discussing the requirements of an IoT-enabled OS, taking into consideration the major concerns when developing solutions at the network edge, followed by a deep comparative analysis and benchmarking on Contiki-NG, RIOT, and Zephyr. Such OSes were considered as the best representative of their class considering the main key-points that best define an OS for resource-constrained IoT devices: low-power consumption, real-time capabilities, security awareness, interoperability, and connectivity. While evaluating each OS under different network conditions, the gathered results revealed distinct behaviors for each OS feature, mainly due to differences in kernel and network stack implementations.
David Cerdeira, Sandro Pinto 0001, Tiago Gomes 0001
IEEE Internet Things J.3
2019 ChamelIoT: An Agnostic Operating System Framework for Reconfigurable IoT Devices
abstract
This letter proposes ChamelIoT, an agnostic operating system (OS) framework for reconfigurable Internet of Things (IoT) devices. ChamelIoT is bringing a reconfigurable hardware OS stack supported by a semantically enriched infrastructure that aims at offering an easy-to-use tool for building (mainly) low-power and low-cost IoT sensors with performance, real-time, and power consumption advantages.
Adriano Tavares, Tiago Gomes 0001, Sandro Pinto 0001
IEEE Internet Things J.4
2018 A 6LoWPAN Accelerator for Internet of Things Endpoint Devices
abstract
The Internet of Things (IoT) is revolutionizing the Internet of the future and the way smart devices, people, and processes interact with each other. Challenges in endpoint devices (EDs) communication are due not only to the security and privacy-related requirements but also due to the ever-growing amount of data transferred over the network that needs to be handled, which can considerably reduce the system availability to perform other tasks. This paper presents an IPv6 over low power wireless personal area networks (6LoWPAN) accelerator for the IP-based IoT networks, which is able to process and filter IPv6 packets received by an IEEE 802.15.4 radio transceiver without interrupting the microcontroller. It offers nearly 13.24% of performance overhead reduction for the packet processing and the address filtering tasks, while guaranteeing full system availability when unwanted packets are received and discarded. The contribution is meant to be deployed on heterogeneous EDs at the network edge, which include on the same architecture, field-programmable gate array (FPGA) technology beside a microcontroller and an IEEE 802.15.4-compliant radio transceiver. Performed evaluations include the low FPGA hardware resources utilization and the energy consumption analysis when the processing and accepting/rejection of a received packet is performed by the 6LoWPAN accelerator.
Tiago Gomes 0001, Filipe Salgado, Sandro Pinto 0001, Jorge Cabral 0001, Adriano Tavares
IEEE Internet Things J.3
2017 LTZVisor: TrustZone is the Key
abstract
Virtualization technology starts becoming more and more widespread in the embedded systems arena, driven by the upward trend for integrating multiple environments into the same hardware platform. The penalties incurred by standard software-based virtualization, altogether with the strict timing requirements imposed by real-time virtualization are pushing research towards hardware-assisted solutions. Among existing commercial off-the-shelf (COTS) technologies, ARM TrustZone promises to be a game-changer for virtualization, despite of this technology still being seen with a lot of obscurity and scepticism. In this paper we present a Lightweight TrustZone-assisted Hypervisor (LTZVisor) as a tool to understand, evaluate and discuss the benefits and limitations of using TrustZone hardware to assist virtualization. We demonstrate how TrustZone can be adequately exploited for meeting the real-time needs, while presenting a low performance cost on running unmodified rich operating systems. While ARM continues to spread TrustZone technology from the applications processors to the smallest of microcontrollers, it is undeniable that this technology is gaining an increasing relevance. Our intent is to encourage research and drive the next generation of TrustZone-assisted virtualization solutions.
Sandro Pinto 0001, Tiago Gomes 0001, Adriano Tavares, Jorge Cabral 0001
ECRTS1
2017 Lightweight multicore virtualization architecture exploiting ARM TrustZone
abstract
Virtualization technology is well established in the server and desktop spaces, and has been spreading across embedded system market. This technology allows for the coexistence and execution of multiples operating systems on top of the same hardware platform, with proven technological and economic benefits. Hardware extensions for easing virtualization have been added into several commercial off-the-shelf processors. Among existing technologies, ARM TrustZone is gaining particular attention due to its broadly availability into ARM processors. However, existent TrustZone-assisted virtualization solutions are limited to a dual-guest and single-core configuration, which can lead to the starvation of the non-secure side when the secure world does not yield the processor. This work presents the extension of a TrustZone-assisted hypervisor to an asymmetric multi-processing configuration. We describe and demonstrate how to run a general-purpose operating system side-by-side with an real-time operating system in a Xilinx Zynq-based platform, enhanced with a dual ARM Cortex-A9. The achieved results demonstrate that the implemented multicore approach not only completely eliminates starvation, but also increases the general-purpose operating system's performance, especially when the real-time workload is demanding.
Sandro Pinto 0001, Jorge Cabral 0001, João Monteiro 0001, Adriano Tavares
IECON1
2016 Towards an FPGA-based network layer filter for the Internet of Things edge devices
abstract
In the near future, billions of new smart devices will connect the big network of the Internet of Things, playing an important key role in our daily life. Allowing IPv6 on the low-power resource constrained devices will lead research to focus on novel approaches that aim to improve the efficiency, security and performance of the 6LoWPAN adaptation layer. This work in progress paper proposes a hardware-based Network Packet Filtering (NPF) and an IPv6 Link-local address calculator which is able to filter the received IPv6 packets, offering nearly 18% overhead reduction. The goal is to obtain a System-on-Chip implementation that can be deployed in future IEEE 802.15.4 radio modules.
Tiago Gomes 0001, Filipe Salgado, Sandro Pinto 0001, Jorge Cabral 0001, Adriano Tavares
ETFA3
2015 RT-SHADOWS: Real-time system hardware for agnostic and deterministic OSes within softcore
abstract
Multithreading is a key feature for dealing with the complexity of current generation embedded devices. Real-Time Operating Systems (RTOSes) provide a higher level of abstraction that alleviates design complexity and time-to-market pressure, inducing, however, undesired latencies and unpredictable execution times. Research works have been focusing on improving performance, determinism or agnosticism, but none simultaneously tackled the three metrics. This work in progress paper presents RT-SHADOWS, a co-designed architecture that promotes configurability, determinism and agnosticism from the outset. A multi-thread ARM softcore processor was exploited and extended by offloading the scheduling-related features into the hardware layer. The hardware multi-thread support is transparent for the standard RTOS applications and independent of the Operating System (OS). Promising preliminary results demonstrate huge improvements on performance and determinism, with only a small cost on hardware.
Tiago Gomes 0001, Sandro Pinto 0001, Paulo Garcia, Adriano Tavares
ETFA2
2015 Towards an FPGA-based edge device for the Internet of Things
abstract
With the growing ubiquity of Internet of Things (IoT), myriads of smart devices connect and share important information over the Internet. In order to provide connectivity and interoperability of all the existing heterogeneous wireless devices, a full communication stack is proposed by the IoT Architecture Reference Model (IoT-ARM). From the sensor to the cloud, the proposed stack can be implemented on all IoT devices avoiding the battle for the wireless standard that will be adopted. This work in progress paper proposes an FPGA-based edge device for IoT, which uses SoC (System-on-Chip) FPGA technology to offload critical features of the communication stack to dedicated hardware, aiming to increase systems performance.
Tiago Gomes 0001, Sandro Pinto 0001, Adriano Tavares, Jorge Cabral 0001
ETFA2
2015 FreeTEE: When real-time and security meet
abstract
The pervasive use of embedded computing systems in modern societies altogether with the industry trend towards consolidating workloads, openness and interconnectedness, have raised security, safety, and real-time concerns. Virtualization has been used as an enabler for safety and security, but research works have proven that it must be extended and improved with hardware-based security foundations. ARM Trustzone has been used for the realization of Trusted Environments, however in this case real-time requirements are completely disregarded. This work in progress paper presents FreeTEE, an embedded architecture that emphasizes and preserves the real-time properties of the system but still guarantees security from the outset. TrustZone technology is exploited to implement the basic building blocks of a Trusted Execution Environment (TEE) as a lower-priority thread of a RTOS. Preliminary results demonstrated that the real-time properties of the RTOS remain practically intact.
Sandro Pinto 0001, Daniel Oliveira 0003, Jorge Cabral 0001, Adriano Tavares
ETFA1
2015 HcM-FreeRTOS: Hardware-centric FreeRTOS for ARM multicore
abstract
Migration to multicore is inevitable. To harness the potential of this technology, embedded system designers need to have available operating systems (OSes) with built-in capabilities for multicore hardware. When designed to meet real-time requirements, multicore SMP (Symmetric Multiprocessing) OSes not only face the inherent problem of concurrent access to shared kernel resources, but still suffer from a bifid priority space, dictated by the co-existence of threads and interrupts. This work in progress paper presents the offloading of the FreeRTOS kernel components to a commercial-off-the-shelf (COTS) multicore hardware. The ARM Generic Interrupt Controller (GIC) is exploited to implement a multicore hardware-centric version of the FreeRTOS that not only solves the priority inversion problem, but also removes the need of internal software synchronization points. Promising preliminary results on performance and determinism are presented, and the research roadmap is discussed.
Esam A. Al Qaralleh, Tiago Gomes 0001, Adriano Tavares, Sandro Pinto 0001
ETFA5
2014 Towards a lightweight embedded virtualization architecture exploiting ARM TrustZone
abstract
Virtualization has been used as the de facto technology to allow multiple operating systems (virtual machines) to run on top of the same hardware platform. In the embedded systems domain, virtualization research has focused on the coexistence of real-time requirements with non-real-time characteristics. However, existent standard software-based virtualization solutions have been shown to negatively impact the overall system, especially in performance, memory footprint and determinism. This work in progress paper presents the implementation of an embedded virtualization architecture through commodity hardware. ARM TrustZone technology is exploited to implement a lightweight virtualization solution with low overhead and high determinism, corroborated by promising preliminary results. Research roadmap is also pointed and discussed.
Sandro Pinto 0001, Daniel Oliveira 0003, Nuno Cardoso, Mongkol Ekpanyapong, Jorge Cabral 0001, Adriano Tavares
ETFA1
2012 Shifting SOA to MPSoC: An exploratory example of application
abstract
This paper presents an innovative Multi Processor System on Chip (MPSoC) architecture based on Service Oriented Architecture (SOA) concepts. The proposed architecture promotes the design flexibility and interoperability by setting an implementation agnostic services' integration interface and being reconfiguration aware by design. The concept is tried with an example of an embedded legacy support system, a Dynamic Binary Translation (DBT) engine based on three services. DBT can solve the x86 architecture virtualization limitations. A SOA based MPSoC endowed of autonomous dynamic reconfiguration capabilities is introduced and pointed as a research opportunity. The introduction in the architecture of a Service Module responsible for Partial Runtime Reconfiguration (PRTR) is expected to dynamically adjust the architecture to better match the workload.
Filipe Salgado, Paulo Garcia, Tiago Gomes 0001, João Vale, Sandro Pinto 0001, Jorge Cabral 0001, Mongkol Ekpanyapong
ETFA5