Alessandro Verosimile

dblp:363/8813 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0000-1814-9338ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Shakan: Training-Inference co-design for Oblique Random Forests on Embedded Devices
abstract
Embedded systems are increasingly leveraging Artificial Intelligence of Things (AIoT) to enable real-time decision-making in critical applications, such as autonomous navigation and medical diagnostics. In these contexts, Random Forests (RFs) have been widely adopted due to their inherent parallelism. However, RFs rely on axis-aligned splits, which limit their ability to model complex decision boundaries. Oblique Random Forests (ORFs), which employ hyperplane-based splits, offer a more expressive alternative by improving classification accuracy. Despite their advantages, inference of ORFs is resource-consuming, prohibiting the implementation of such models on resource-constrained hardware devices.In this work, we present Shakan, a novel framework for Oblique Decision Trees (ODTs) inference on embedded systems. We introduce a new training technique designed to mitigate both training complexity and overfitting while enabling low-latency inference in hardware, along with a new architecture that maximizes performance and optimizes resource usage. Shakan enables, on resource-constrained devices, the inference of several ORFs configurations that can provide either a significant increase in accuracy or a notable speedup in terms of inference latency compared to state-of-the-art accelerators for traditional RFs on embedded devices. The most accurate configurations provide average accuracy improvements above 5% with similar latency, while the fastest configurations achieve speedups of 1140×, 214×, and 29× for tree depths of 5, 7, and 9, respectively, with comparable accuracy.
Alessandro Annechini, Alessandro Verosimile, Marco D. Santambrogio
DATE2
2025 Moyogi: A Memory-Centric Accelerator for Low-Latency Random Forest Inference on Embedded Devices
abstract
The convergence of Artificial Intelligence (AI) and Internet of Things (IoT) is driving the need for real-time, low-latency architectures to trust the inference of complex Machine Learning (ML) models in critical applications like autonomous vehicles and smart healthcare. While traditional cloud-based solutions introduce latency due to the need to transmit data to and from centralized servers, edge computing offers lower response times by processing data locally. In this context, Random Forests (RFs) are highly suited for building hardware accelerators over resource-constrained edge devices due to their inherent parallelism. Nevertheless, maintaining a low latency as the size of the RF grows is still critical for state-of-the-art (SoA) approaches. To address this challenge, this paper proposes Moyogi, a hardware-software codesign framework for memory-centric RF inference that optimizes the architecture for the target ML model, employing RFs with Decision Trees (DTs) of multiple depths and exploring several architectural variations to find the best-performing configuration. We propose a resource estimation model based on the most relevant architectural features to enable effective Design Space Exploration. Moyogi achieves a geomean latency reduction of 3.88x on RFs trained on relevant IoT datasets, compared to the best-performing SoA memory-centric architecture.
Alessandro Verosimile, Francesco Peverelli, Marco D. Santambrogio
FCCM1
2024 YoseUe: "trimming" Random Forest's training towards resource-constrained inference
abstract
Endowing artificial objects with intelligence is a longstanding computer science and engineering vision that recently converged under the umbrella of Artificial Intelligence of Things (AIoT). Nevertheless, AIoT’s mission cannot be fulfilled if objects rely on the cloud for their “brain,” at least concerning inference. Thanks to heterogeneous hardware, it is possible to bring Machine Learning (ML) inference on resource-constrained embedded devices, but this requires careful co-optimization between model training and its hardware acceleration. This work proposes YoseUe, a memory-centric hardware co-processor for Random Forests (RFs) inference, which significantly reduces the waste of memory resources by exploiting a novel train-acceleration co-optimization. YoseUe proposes a novel ML model, the Multi-Depth Random Forest Classifier (MDRFC), in which a set of RFs are trained at decreasing depths and then weighted, exploiting a Neural Network (NN) tailored to counteract potential accuracy losses w.r.t. classical RFs. With this modeling technique, first proposed in this paper, it becomes possible to accelerate the inference of RFs that count up to 2 orders of magnitude more Decision Trees (DTs) than those the current state-of-the-art architectures can fit on embedded devices. Furthermore, this is achieved without losing accuracy with respect to classical, full-depth RF in their most relevant configurations.
Alessandro Verosimile, Alessandro Tierno, Andrea Damiani, Marco D. Santambrogio
ASPDAC1
2024 SATL: A Spatial Architecture Rapid Prototyping Framework for Irregular Applications Acceleration
abstract
Modern FPGA HLS tools are proficient at accelerating datapath applications, but they generate considerable overhead when dealing with irregular, control-driven workloads. Conversely, RTL-based approaches significantly increase development time and system integration effort. To address this tooling gap, we present SATL, a Chisel-based rapid prototyping framework for building FPGA-based spatial architectures targeting irregular workloads. We use it to re-implement YoseUe, a state-of-the-art accelerator for inferring Decision Tree Ensemble Machine Learning models. Compared to the original HLS-based work, SATL yields an average 3.4 x throughput improvement and reduces the architecture's resource consumption, allowing inference on Ensemble Models with up to 3.6 x more trees, enabling the deployment of larger models on resource-constrained devices.
Francesco Peverelli, Alessandro Verosimile, Davide Conficconi, Andrea Damiani, Marco D. Santambrogio
ICCD2