EDBT 2026 Demo / reviewers in the wild / expert
Andrea Mondelli
dblp:162/5147
· DBLP profile ↗
7ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-4740-7081ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Architecture-Level CPU Modeling Framework for Power and Other Design QualitiesabstractPower efficiency is a critical design objective in modern microprocessor design. To evaluate the impact of architectural-level design decisions, an accurate yet efficient architecture-level power model is desired. However, widely adopted analytical power models like McPAT and Wattch have been criticized for their unreliable accuracy, while machine learning (ML) methods like McPAT-Calib rely on sufficient known designs for training and perform poorly when available designs are limited, which is the case in realistic scenarios. In this work, we propose PANDA, an innovative architecture-level solution that combines the advantages of analytical and ML power models. It achieves unprecedented high accuracy on unknown new designs even when there are very limited designs for training. Besides being an excellent average power model, we also extend PANDA to support the time-based power trace prediction, which can enable the analysis of peak power, power fluctuations, and voltage fluctuation. This is highly challenging at the architecture level. Other qualities, such as area, performance, and energy accurately, can also be supported. In addition to single design quality, PANDA can model the tradeoffs among different design qualities, such as the tradeoff between power and timing, by predicting the Pareto-optimal curve. Finally, PANDA can further support power prediction for unknown new technology nodes. Our experiment shows that, for average power prediction, our method can achieve high accuracy with a correlation coefficient R of 0.99 and mean absolute percentage error (MAPE) of 7.91% even when only one configuration is known, outperforming McPAT-Calib which has R of -0.24 and MAPE of 35.96%. For time-based power trace prediction, our method can achieve a low MAPE of 4.34%, outperforming the state-of-the-art method Powertrain which has a MAPE of 53.8%. Qijun Zhang, Mengming Li, Andrea Mondelli, Zhiyao Xie |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | DeVAS: Decoupled Virtual Address SpacesabstractThe constant growth of workload size in modern applications is making address translation a performance bottleneck. In principle, increasing the virtual page size could be advantageous, as it would allow each cached address translation to cover a larger memory space. Nevertheless, the utilization of larger pages introduces challenges, such as issues related to memory fragmentation and physical page management. In this paper, we present Decoupled Virtual Address Spaces (DeVAS), a virtual memory proposal that enables the decoupling of address translation and memory allocation, by allowing different sizes for virtual and physical pages, aiming to exploit the benefits of both. DeVAS introduces an intermediate virtual address space allocated by the Operating System employing larger pages (e.g., 2MiB), corresponding to the virtual page size seen by the processor, and a memory controller extension devoted to their allocation in physical memory at a smaller granularity (e.g., 4KiB). We show that DeVAS achieves an average 1.13 × performance improvement over a traditional configuration (i.e., with 4KiB pages) with no specific architectural modifications and for memory-intensive benchmarks. Moreover, DeVAS strategy enables architectural modifications for increasing performance, such as simplified/optimized TLB structure and L1-cache design flexibility. When considering these adjustments, DeVAS achieves a speedup of up to 1.20 × compared to the same reference. Furthermore, it matches and even edges the performance of an ideal (i.e., not implementable in practice) virtual memory configuration based on 2MiB pages only. Mirco Mannino, Biagio Peccerillo, Andrea Mondelli, Sandro Bartolini |
SBAC-PAD | 3 |
| 2023 | Energy and Performance Improvements for Convolutional Accelerators Using Lightweight Address Translation SupportabstractThe growing demand for deep learning applications has led to the design and development of several hardware accelerators to increase performance and energy efficiency. In particular, convolutional accelerators are among those receiving the most attention due to their applicability in many fields. Another aspect that is gaining increasing attention is the use of a shared virtual address space between processor and accelerators. It can provide several advantages such as programmability and security. The use of a shared address space relies on a time-consuming IOMMU to satisfy address translation requests. In this work, we analyze convolutional workloads in convolutional accelerators, identifying the sensitivity of performance to IOMMU activity. Additionally, based on the analysis done on convolutional workloads, we propose the use of dedicated accelerator registers (Translation Registers) to reduce costly IOMMU accesses. Translation Registers allow reducing execution time by about 20% and the energy consumption related to address translation up to about 55%. Mirco Mannino, Biagio Peccerillo, Andrea Mondelli, Sandro Bartolini |
CF | 3 |
| 2022 | A survey on hardware accelerators: Taxonomy, trends, challenges, and perspectivesabstractIn recent years, the limits of the multicore approach emerged in the so-called “dark silicon” issue and diminishing returns of an ever-increasing core count. Hardware manufacturers, out of necessity, switched their focus to accelerators, a new paradigm that pursues specialization and heterogeneity over generality and homogeneity. They are special-purpose hardware structures separated from the CPU with aspects that exhibit a high degree of variability. We define a taxonomy based on fourteen of these aspects, grouped in four macro-categories: general aspects, host coupling, architecture, and software aspects. According to it, we categorize around 100 accelerators of the last decade from both industry and academia, and critically analyze emerging trends. We complete our discussion with throughput and efficiency figures. Then, we discuss some prominent open challenges that accelerators are facing, analyzing state-of-the-art solutions, and suggesting prospective research directions for the future. Biagio Peccerillo, Mirco Mannino, Andrea Mondelli, Sandro Bartolini |
J. Syst. Archit. | 3 |
| 2021 | SeMPE: Secure Multi Path Execution Architecture for Removing Conditional Branch Side ChannelsabstractOne prevalent source of side channel vulnerabilities is the secret-dependent behavior of conditional branches (SDBCB). The state-of-the-art solution relies on Constant-Time Expressions, which require high programming effort and incur high performance overheads. In this paper, we propose SeMPE, an architecture support to eliminate SDBCB without requiring much programming effort while incurring low performance overheads. When a secret-dependent branch is encountered, SeMPE fetches, executes, and commits both paths of the branch, preventing the adversary from inferring secret values from the branching behavior of the program. SeMPE outperforms code generated by FaCT, a constant-time expression language, by up to 18×. Andrea Mondelli, Paul Gazzillo, Yan Solihin |
DAC | 1 |
| 2015 | Dataflow Support in x86_64 Multicore Architectures through Small Hardware ExtensionsabstractThe path towards future high performance computers requires architectures able to efficiently run multi-threaded applications. In this context, dataflow-based execution models can improve the performance by limiting the synchronization overhead, thanks to a simple producer-consumer approach. This paper advocates the ISE of standard cores with a small hardware extension for efficiently scheduling the execution of threads on the basis of dataflow principles. A set of dedicated instructions allow the code to interact with the scheduler. Experimental results demonstrate that, the combination of dedicated scheduling units and a dataflow execution model improve the performance when compared with other techniques for code parallelization (e.g., OpenMP, Cilk). Andrea Mondelli, Nam Ho, Alberto Scionti, Marco Solinas, Antoni Portero, Roberto Giorgi |
DSD | 1 |
| 2015 | Revisiting Clustered Microarchitecture for Future Superscalar Cores: A Case for Wide Issue ClustersabstractDuring the past 10 years, the clock frequency of high-end superscalar processors has not increased. Performance keeps growing mainly by integrating more cores on the same chip and by introducing new instruction set extensions. However, this benefits only some applications and requires rewriting and/or recompiling these applications. A more general way to accelerate applications is to increase the IPC, the number of instructions executed per cycle. Although the focus of academic microarchitecture research moved away from IPC techniques, the IPC of commercial processors was continuously improved during these years. We argue that some of the benefits of technology scaling should be used to raise the IPC of future superscalar cores further. Starting from microarchitecture parameters similar to recent commercial high-end cores, we show that an effective way to increase the IPC is to allow the out-of-order engine to issue more micro-ops per cycle. But this must be done without impacting the clock cycle. We propose combining two techniques: clustering and register write specialization. Past research on clustered microarchitectures focused on narrow issue clusters, as the emphasis at that time was on allowing high clock frequencies. Instead, in this study, we consider wide issue clusters, with the goal of increasing the IPC under a constant clock frequency. We show that on a wide issue dual cluster, a very simple steering policy that sends 64 consecutive instructions to the same cluster, the next 64 instructions to the other cluster, and so forth, permits tolerating an intercluster delay of three cycles. We also propose a method for decreasing the energy cost of sending results from one cluster to the other cluster. Pierre Michaud, Andrea Mondelli, André Seznec |
ACM Trans. Archit. Code Optim. | 2 |