EDBT 2026 Demo / reviewers in the wild / expert
Francesco Peverelli
dblp:224/1513
· DBLP profile ↗
4ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-0569-3446ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Moyogi: A Memory-Centric Accelerator for Low-Latency Random Forest Inference on Embedded DevicesabstractThe convergence of Artificial Intelligence (AI) and Internet of Things (IoT) is driving the need for real-time, low-latency architectures to trust the inference of complex Machine Learning (ML) models in critical applications like autonomous vehicles and smart healthcare. While traditional cloud-based solutions introduce latency due to the need to transmit data to and from centralized servers, edge computing offers lower response times by processing data locally. In this context, Random Forests (RFs) are highly suited for building hardware accelerators over resource-constrained edge devices due to their inherent parallelism. Nevertheless, maintaining a low latency as the size of the RF grows is still critical for state-of-the-art (SoA) approaches. To address this challenge, this paper proposes Moyogi, a hardware-software codesign framework for memory-centric RF inference that optimizes the architecture for the target ML model, employing RFs with Decision Trees (DTs) of multiple depths and exploring several architectural variations to find the best-performing configuration. We propose a resource estimation model based on the most relevant architectural features to enable effective Design Space Exploration. Moyogi achieves a geomean latency reduction of 3.88x on RFs trained on relevant IoT datasets, compared to the best-performing SoA memory-centric architecture. Alessandro Verosimile, Francesco Peverelli, Marco D. Santambrogio |
FCCM | 2 |
| 2025 | DFlows: A Flow-Based Programming Approach for a Polyglot Design-Space Exploration FrameworkabstractCurrent architectural Design-Space Exploration (DSE) tools specify the exploration problem through annotations or pragmas. However, this approach is inherently language-dependent and limits the applicability to one specific target language and synthesis toolchain. Additionally, the rapid development of new hardware Domain-Specific Languages, programming models, and different exploration heuristics calls for a language-agnostic and modular approach. To address this need, we present a DSE formalization to facilitate the integration of new components and customized flows and leverage it to implement DFlows , a flow-based-programming DSE tool that decouples problem definition, code generation, exploration, and evaluation strategies. DFlows ’s compiler-based frontend provides language-agnostic generation of design points through Abstract Syntax Tree manipulation. We show how DFlows can integrate custom performance models from complex state-of-the-art accelerators for Verilog, VHDL, Chisel, and HLS designs. We compare the runtimes of our DSE process against a state-of-the-art Chisel-based DSE tool, achieving up to 3.74× speedup while identifying the same set of optimal solutions. Additionally, we integrate in DFlows a custom exploration heuristic leveraging genetic algorithms and a novel online learning fitness function approximation methodology. This approximation yields a negligible hypervolume difference with the exhaustive search Pareto-front while improving DSE runtime by up to 2.67×. Francesco Peverelli, Daniele Paletti, Davide Conficconi |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2024 | SATL: A Spatial Architecture Rapid Prototyping Framework for Irregular Applications AccelerationabstractModern FPGA HLS tools are proficient at accelerating datapath applications, but they generate considerable overhead when dealing with irregular, control-driven workloads. Conversely, RTL-based approaches significantly increase development time and system integration effort. To address this tooling gap, we present SATL, a Chisel-based rapid prototyping framework for building FPGA-based spatial architectures targeting irregular workloads. We use it to re-implement YoseUe, a state-of-the-art accelerator for inferring Decision Tree Ensemble Machine Learning models. Compared to the original HLS-based work, SATL yields an average 3.4 x throughput improvement and reduces the architecture's resource consumption, allowing inference on Ensemble Models with up to 3.6 x more trees, enabling the deployment of larger models on resource-constrained devices. Francesco Peverelli, Alessandro Verosimile, Davide Conficconi, Andrea Damiani, Marco D. Santambrogio |
ICCD | 1 |
| 2019 | Automated Acceleration of Dataflow-Oriented C Applications on FPGA-Based SystemsabstractThe acceleration of compute-intensive applications on FPGA-based systems has become an increasingly common trend thanks to their availability as cloud commodities. This trend has also been accompanied by wider support of High-Level Synthesis tools. Despite these solutions reduce the learning curve for hardware development, the programmer still requires specific expertise in order to achieve efficient implementations. In this paper, we propose an automated approach for the acceleration of C applications into dataflow kernels on FPGAs. Francesco Peverelli, Marco Rabozzi, Salvatore Cardamone, Emanuele Del Sozzo, Alex J. W. Thom, Marco D. Santambrogio, Lorenzo Di Tucci |
FCCM | 1 |