VLDB 2026 Research / reviewers in the wild / expert
Pietro Mercati
dblp:130/1383
· DBLP profile ↗
21ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0003-2842-7201ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 6 first-author · 8 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LogHD: Robust Compression of Hyperdimensional Classifiers via Logarithmic Class-Axis ReductionabstractHyperdimensional computing (HDC) suits memory, energy, and reliability-constrained systems, yet the standard "one prototype per class" design requires $O(CD)$ memory (with $C$ classes and dimensionality $D$). Prior compaction reduces $D$ (feature axis), improving storage/compute but weakening robustness. We introduce LogHD, a logarithmic class-axis reduction that replaces the $C$ per-class prototypes with $n\!\approx\!\lceil\log_k C\rceil$ bundle hypervectors (alphabet size $k$) and decodes in an $n$-dimensional activation space, cutting memory to $O(D\log_k C)$ while preserving $D$. LogHD uses a capacity-aware codebook and profile-based decoding, and composes with feature-axis sparsification. Across datasets and injected bit flips, LogHD attains competitive accuracy with smaller models and higher resilience at matched memory. Under equal memory, it sustains target accuracy at roughly $2.5$-$3.0\times$ higher bit-flip rates than feature-axis compression; an ASIC instantiation delivers $498\times$ energy efficiency and $62.6\times$ speedup over an AMD Ryzen 9 9950X and $24.3\times$/$6.58\times$ over an NVIDIA RTX 4090, and is $4.06\times$ more energy-efficient and $2.19\times$ faster than a feature-axis HDC ASIC baseline. Sanggeon Yun, Hyunwoo Oh, Ryozo Masukawa, Pietro Mercati, Nathaniel D. Bastian, Mohsen Imani |
DATE | 4 |
| 2025 | LLM-IMC: Automating Analog In-Memory Computing Architecture Generation with Large Language ModelsabstractResistive crossbars enabling analog In-Memory Computing (IMC) have garnered significant attention from academia and industry as a promising architecture for Deep Neural Network (DNN) acceleration, thanks to their high memory access bandwidth and in-situ computing capabilities. However, the knowledge-intensive hardware design process and the lack of high-quality circuit netlists have constrained design space exploration and optimization of analog IMC to behavioral system-level tools. In this one-page abstract, we introduce LLM-IMC, a novel fine-tune-free Large Language Model (LLM) framework, supported by a Python-based tool, designed for analog IMC SPICE code generation. LLM-IMC systematically addresses these limitations by automating the creation of diverse IMC simulation scripts, enabling efficient design space exploration through LLM-driven performance, and outlining an integration roadmap for hardware-oriented neuromorphic crossbar design flows. Deepak Vungarala, Md Hasibul Amin, Pietro Mercati, Arman Roohi, Ramtin Zand, Shaahin Angizi |
FCCM | 3 |
| 2025 | Event-Driven Spatiotemporal Processing-In-Sensor with Phase Change Memory-based Optical Acceleration
Mehrdad Morsali, Deniz Najafi, Amin Shafiee, Sepehr Tabrizchi, Pietro Mercati, Mohsen Imani, Arman Roohi, Navid Khoshavi, Mahdi Nikdast, Shaahin Angizi |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | Opto-ViT: Architecting a Near-Sensor Region of Interest-Aware Vision Transformer Accelerator with Silicon PhotonicsabstractVision Transformers (ViTs) have emerged as a powerful architecture for computer vision tasks due to their ability to model long-range dependencies and global contextual relationships. However, their substantial compute and memory demands hinder efficient deployment in scenarios with strict energy and bandwidth limitations. In this work, we propose Opto-ViT, the first near-sensor, region-aware ViT accelerator leveraging silicon photonics (SiPh) for real-time and energy-efficient vision processing. Opto-ViT features a hybrid electronic-photonic architecture, where the optical core handles compute-intensive matrix multiplications using Vertical-Cavity Surface-Emitting Lasers (VCSELs) and Microring Resonators (MRs), while nonlinear functions and normalization are executed electronically. To reduce redundant computation and patch processing, we introduce a lightweight Mask Generation Network (MGNet) that identifies regions of interest in the current frame and prunes irrelevant patches before ViT encoding. We further co-optimize the ViT backbone using quantization-aware training and matrix decomposition tailored for photonic constraints. Experiments across device fabrication, circuit and architecture co-design, to classification, detection, and video tasks demonstrate that Opto-ViT achieves 100.4 KFPS/W with up to 84% energy savings with less than 1.6% accuracy loss, while enabling scalable and efficient ViT deployment at the edge. Mehrdad Morsali, Chengwei Zhou, Deniz Najafi, Sreetama Sarkar, Pietro Mercati, Navid Khoshavi, Peter A. Beerel, Mahdi Nikdast, Gourav Datta, Shaahin Angizi |
ICCAD | 5 |
| 2024 | Multi-Objective Software-Hardware Co-Optimization for HD-PIM via Noise-Aware Bayesian OptimizationabstractIn hardware accelerator design, software-hardware co-optimization requires intricate trade-offs and tight integration between software algorithms and hardware design to optimize performance, power efficiency, and area (PPA) while ensuring high accuracy. Furthermore, the inherent non-ideality in some emerging hardware technologies poses extra challenges to the co-optimization problem. This paper proposes a novel software-hardware co-optimization framework for hyperdimensional (HD) computing accelerators with emerging ReRAM-based processing in-memory (PIM) technologies, which have shown superior performance and energy efficiency over conventional machine learning accelerators. We first comprehensively characterize the non-trivial trade-offs between design parameters in HD-PIM and PPA and accuracy metrics in HD-PIM. Then, we develop a multi-objective noise-aware Bayesian optimization algorithm to find the Pareto set (optimal trade-offs between metrics) of the HD-PIM design. Our methodology uniquely addresses the stochastic nature of ReRAM by integrating error characteristics into the optimization process, thereby enhancing the quality of the generated designs. Experimental results show that our configurations achieve up to 4.28% accuracy improvement, 35.38% power reduction, 49x timing improvement, and 10% area reduction over a non-optimized design. Chien-Yi Yang, Minxuan Zhou, Flavio Ponzina, Suraj Sathya Prakash, Raid Ayoub, Pietro Mercati, Mahesh Subedar, Tajana Rosing |
ICCAD | 6 |
| 2023 | Efficient Off-Policy Reinforcement Learning via Brain-Inspired ComputingabstractReinforcement Learning (RL) has opened up new opportunities to enhance existing smart systems that generally include a complex decision-making process. However, modern RL algorithms, e.g., Deep Q-Networks (DQN), are based on deep neural networks, resulting in high computational costs. In this paper, we propose QHD, an off-policy value-based Hyperdimensional Reinforcement Learning, that mimics brain properties toward robust and real-time learning. QHD relies on a lightweight brain-inspired model to learn an optimal policy in an unknown environment. On both desktop and power-limited embedded platforms, QHD achieves significantly better overall efficiency than DQN while providing higher or comparable rewards. QHD is also suitable for highly-efficient reinforcement learning with great potential for online and real-time learning. Our solution supports a small experience replay batch size that provides 12.3 times speedup compared to DQN while ensuring minimal quality loss. Our evaluation shows QHD capability for real-time learning, providing 34.6 times speedup and significantly better quality of learning than DQN. Yang Ni 0001, Danny Abraham, Mariam Issa, Yeseong Kim, Pietro Mercati, Mohsen Imani |
ACM Great Lakes Symposium on VLSI | 5 |
| 2023 | Brain-Inspired Trustworthy Hyperdimensional Computing with Efficient Uncertainty QuantificationabstractRecent advancement in emerging brain-inspired computing has pointed out a promising path to Machine Learning (ML) algorithms with high efficiency. Particularly, research in the field of HyperDimensional Computing (HDC) brings orders of magnitude speedup to both ML model training and inference compared to their deep learning counterparts. However, current HDC-based ML algorithms generally lack uncertainty estimation, despite having shown good results in various practical applications and outstanding energy efficiency. On the other hand, existing solutions such as the Bayesian Neural Networks (BNN) are generally much slower than regular neural networks and lead to high energy consumption. In this paper, we propose a hyperdimensional Bayesian framework called DiceHD, which enables uncertainty estimation for the HDC-based regression algorithm. The core of our framework is a specially designed HDC encoder that maps input features to the high dimensional space with an extra layer of randomness, i.e., a small number of dimensions are randomly dropped for each input. Our key insight is that by using this encoder, DiceHD implements Bayesian inference while maintaining the efficiency advantage of HDC. We verify our framework with both toy regression tasks and real-world datasets. We compare our DiceHD to several widely-used BNN baselines in terms of performance and efficiency. The results on CPU show that DiceHD provides comparable uncertainty estimations while achieving significant speedup compared to the BNN baseline. We also deploy DiceHD on two FPGA platforms with different acceleration capabilities, showing that DiceHD provides up to 84× (3740×) better energy efficiency for training (inference). Yang Ni 0001, Hanning Chen, Prathyush Poduval, Zhuowen Zou, Pietro Mercati, Mohsen Imani |
ICCAD | 5 |
| 2023 | Dynamic Reliability Management of Multigateway IoT Edge Computing SystemsabstractThe emerging paradigm of edge computing envisions to overcome the shortcomings of cloud-centric Internet of Things (IoT) by providing data processing and storage capabilities closer to the source of data. Accordingly, IoT edge devices, with the increasing demand of computation workloads on them, are prone to failures more than ever. Hard failures in hardware due to aging and reliability degradation are particularly important since they are irrecoverable, requiring maintenance for the replacement of defective parts, at high costs. In this article, we propose a novel dynamic reliability management (DRM) technique for multigateway IoT edge computing systems to mitigate degradation and defer early hard failures. Taking advantage of the edge computing architecture, we utilize gateways for computation offloading with the primary goal of maximizing the battery lifetime of edge devices, while satisfying the Quality of Service (QoS) and reliability requirements. We present a two-level management scheme, which work together to 1) choose the offloading rates of edge devices; 2) assign edge devices to gateways; and 3) decide multihop data flow routes and rates in the network. The offloading rates are selected by a hierarchical multitimescale distributed controller. We assign edge devices by solving a bottleneck generalized assignment problem (BGAP) and compute optimal flows in a fully distributed fashion, leveraging the subgradient method. Our results, based on real measurements and trace-driven simulation, demonstrate that the proposed scheme can achieve a similar battery lifetime and better QoS compared to the state-of-the-art approaches while satisfying reliability requirements, where other approaches fail by a large margin. Kazim Ergun, Raid Ayoub, Pietro Mercati, Tajana Rosing |
IEEE Internet Things J. | 3 |
| 2022 | DPM-NFV: Dynamic Power Management Framework for 5G User Plane Function using Bayesian OptimizationabstractNetwork Function Virtualization (NFV), the replacement of purpose-built network appliances with software functions running on general purpose compute servers, is ubiquitous in today's telecommunication networks. The 5G User Plane Function (UPF) is an important example of an NFV workload, which enables 5G and internet communications. The UPF has strict packet drop requirements and because user traffic load can vary dramatically throughout the day, the selection of a single static configuration leads to over-provisioning of server resources. To reduce the cost of ownership, network operators can reduce power consumption during periods of low traffic load, but to do so they must ensure that packet drop requirements are met. In this paper we present DPM-NFV, a machine learning based framework that enables dynamic tuning of a real NFV system. Our methodology is composed of two phases: (1) Offline, targeted automated studies use Bayesian Optimization to infer the best configurations for various load levels; (2) Online, a run-time classifier dynamically selects the best configuration for the current load. Our results obtained on a real system demonstrate that the UPF can meet strict packet drop requirements while reducing power consumption by up to 52% with smooth traffic and up to 46% with bursty traffic. Jaroslaw J. Sydir, Bin Li 0018, Pietro Mercati, Tsung-Yuan Charlie Tai, Ravi R. Iyer 0001, Michael Kishinevsky, Boris Serafimov |
GLOBECOM | 3 |
| 2022 | Reinforcement learning based reliability-aware routing in IoT networks
Kazim Ergun, Raid Ayoub, Pietro Mercati, Tajana Rosing |
Ad Hoc Networks | 3 |
| 2021 | Energy and QoS-Aware Dynamic Reliability Management of IoT Edge Computing SystemsabstractThe Internet of Things (IoT) systems, as any electronic or mechanical system, are prone to failures. Hard failures in hardware due to aging and degradation are particularly important since they are irrecoverable, requiring maintenance for the replacement of defective parts, at high costs. In this paper, we propose a novel dynamic reliability management (DRM) technique for IoT edge computing systems to satisfy the Quality of Service (QoS) and reliability requirements while maximizing the remaining energy of the edge device batteries. We formulate a state-space optimal control problem with a battery energy objective, QoS, and terminal reliability constraints. We decompose the problem into low-overhead subproblems and solve it employing a hierarchical and multi-timescale control approach, distributed over the edge devices and the gateway. Our results, based on real measurements and trace-driven simulation demonstrate that the proposed scheme can achieve a similar battery lifetime compared to the state-of-the-art approaches while satisfying reliability requirements, where other approaches fail to do so. Kazim Ergun, Raid Ayoub, Pietro Mercati, Dancheng Liu, Tajana Rosing |
ASP-DAC | 3 |
| 2021 | MOBO-NFV: Automated Tuning of a Network Function Virtualization System using Multi-Objective Bayesian Optimization
Pietro Mercati, Bin Li 0018, Mesut Ali Ergin, Tsung-Yuan Charlie Tai, Michael Kishinevsky, Boris Serafimov, Subhiksha Ravisundar, Eoin Walsh, Thomas Long |
IM | 1 |
| 2019 | Dynamic Optimization of Battery Health in IoT NetworksabstractThe reliability and maintainability of the Internet of Things (IoT) devices become highly important as the number of "things" grows rapidly. The majority of the IoT devices have batteries which age, degrade, and eventually require maintenance. Existing work focuses on ensuring that batteries have sufficient amount of stored charge to operate until they can recharge, but does not consider battery degradation. This leads to high replacement and maintenance costs in large IoT networks. In this paper, we formulate the problem of minimizing battery degradation to improve the lifetime of IoT networks and solve it with Model Predictive Control (MPC) leveraging models for battery dynamics and State of Health (SoH). The battery SoH is modeled using a realistic non-linear model while taking ambient temperature into account. We demonstrate that our solution can improve network lifetime up to 68.5% compared to conventional energy consumption focused algorithms, which use simple linear battery models. The proposed approach achieves near-optimal performance in terms of preserving battery health, staying within 8.7% SoH with respect to an ideal oracle solution on average. Kazim Ergun, Raid Ayoub, Pietro Mercati, Tajana Rosing |
ICCD | 3 |
| 2017 | Multi-variable Dynamic Power Management for the GPU SubsystemabstractIn this work, we present a control-theoretic algorithm to improve the energy efficiency of the GPU targeting deadline-driven graphics applications. Our algorithm dynamically controls multiple power knobs within the GPU (DVFS and number of active slices) that have different control time granularities. We developed a multi-rate predictive control to overcome the time granularity constraints in the control variables and reduce runtime overhead. To enable predictive control, we developed runtime analytical predictive models for performance and power of the GPU, that take input from hardware counters and temperature sensor readings. We evaluated our approach on the latest generation of Intel Core i5 platform. Our experimental results demonstrate significant average GPU energy savings of 25% compared to the state-of-the-art algorithm at negligible performance overhead. Pietro Mercati, Raid Ayoub, Michael Kishinevsky, Eric Samson, Marc Beuchat, Francesco Paterna, Tajana Rosing |
DAC | 1 |
| 2017 | P4: Phase-based power/performance prediction of heterogeneous systems via neural networksabstractThe emergence of Internet of Things increases the complexity and the heterogeneity of computing platforms. Migrating workload between various platforms is one way to improve both energy efficiency and performance. Effective migration decisions require accurate estimates of its costs and benefits. To date, these estimates were done by either instrumenting the source code/binaries, thus causing high overhead, or by using power estimates from hardware performance counters, which work well for individual machines, but until now have not been accurate for predicting across different architectures. In this paper, we propose P4, a new Phase-based Power and Performance Prediction framework which identifies cross-platform application power and performance at runtime for heterogeneous computing systems. P4analyzes and detects machine-independent application phases by characterizing computing platforms offline with a set of benchmarks, and then builds neural network-based models to automatically identify and generalize the complex cross-platform relationships for each benchmark phase. It then leverages these models along with performance counter measurements collected at runtime to estimate performance and power consumption if it were running on a completely different computing platform, including a different CPU architecture, without ever having to run it on there. We evaluate the proposed framework on four commercial heterogeneous platforms, ranging from X86 servers to mobile ARM-based architecture, with 129 industry-standard benchmarks. Our experimental results show that P4can predict the power and performance changes with only 6.8% and 5.6% error, respectively, even for completely different architectures from the ones applications ran on. Yeseong Kim, Pietro Mercati, Ankit More, Emily Shriver, Tajana Rosing |
ICCAD | 2 |
| 2017 | WARM: Workload-Aware Reliability Management in Linux/AndroidabstractWith CMOS scaling beyond 14 nm, reliability is a major concern for IC manufacturers. Reliability-aware design has a non-negligible overhead and cannot account for user experience in mobile devices. An alternative is dynamic reliability management (DRM), which counteracts degradation by adapting the operating conditions at runtime. In this paper, for the first time we formulate DRM as an optimization problem that accounts for reliability, temperature and performance. We develop an optimal policy for multicores using convex optimization, and show that it is not feasible to implement on real systems. For this reason, we propose workload-aware reliability management (WARM), a fast DRM technique adapting to diverse workload requirements to trade reliability and user experience. WARM is implemented and tested on a real Android device. WARM approximates the solution of the convex solver within 5% on average, while executing more than $400 {\times }$ faster. WARM integrates a thermal controller that allocates tasks to meet thermal constraints. This is required since degradation strongly depends on temperature. We show that WARM meets temperature constraints within 5% in 87.5% more cases than the state-of-the-art. We show that WARM task allocation achieves up to one year lifetime improvement for a multicore platform. It can achieve up to 100% of performance improvement on cluster architectures, such as big.LITTLE, while still guaranteeing the reliability target. Finally, we show that it achieves performance in the 4% of the maximum for a broad range of a applications, while meeting the reliability constraints. Pietro Mercati, Francesco Paterna, Andrea Bartolini, Luca Benini, Tajana Rosing |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2016 | VarDroid: Online Variability Emulation in Android/Linux PlatformsabstractVariability is the real big challenge for integrated circuits. Today, simulators help to estimate the effect of variability, but fail to capture real workload dynamics and user interactions, which are fundamental to mobile devices. This paper presents VarDroid, a low-overhead tool to emulate power and performance variability on real platforms, running on top of the Android operating system. VarDroid enables analyzing the effect of variability in power and performance while capturing the complex interactions characteristic of mobile workloads, thus relating to user's quality of experience. The paper presents use cases to show the utility of VarDroid to test applications, device and OS robustness under the effects of variability. Our results show that a variability-agnostic OS can incur in a performance penalty of up to 60% and a power penalty of up to 20%. Pietro Mercati, Francesco Paterna, Andrea Bartolini, Mohsen Imani, Luca Benini, Tajana Rosing |
ACM Great Lakes Symposium on VLSI | 1 |
| 2014 | A Linux-governor based Dynamic Reliability Manager for android mobile devicesabstractReliability is a major concern in multiprocessors. Dynamic Reliability Management (DRM) aims at trading off processor performance with lifetime. The state-of-the-art publications study only the theory supported by simulation. This paper presents the first complete software implementation, working on a real hardware, of a low-overhead, Android-compatible workload-aware DRM Governor for mobile multiprocessors. We discuss the design challenges and the run-time overhead involved. We show the effectiveness of our governor in guaranteeing the predefined target lifetime and show that it achieves up to 100% of lifetime improvement with respect to traditional governors, while providing comparable performance for critical applications. Pietro Mercati, Andrea Bartolini, Francesco Paterna, Tajana Rosing, Luca Benini |
DATE | 1 |
| 2014 | An On-line Reliability Emulation FrameworkabstractTechnology scaling made reliability a primary concern for integrated circuits. Increased power and temperature exasperate the impact of degradation phenomena and shorten processors lifetime. This issue is particularly dramatic for mobile processors, characterized by variable workload and environmental conditions. Due to the different time scales at which reliability phenomena and computation happens, state-of-theartDRM solutions are evaluated using high-level workload and system models. To enable the design of workload-aware DRM with accurate reliability models, in this work we propose a software framework for virtualizing the processors reliability.Our framework captures the effect of variable workload and environmental conditions and allows to emulate longer degradation in a short time scale. We implement the framework on a realAndroid device and exploit it to enable workload-aware DynamicReliability Management (DRM). Pietro Mercati, Andrea Bartolini, Francesco Paterna, Luca Benini, Tajana Rosing |
EUC | 1 |
| 2014 | Dynamic variability management in mobile multicore processors under lifetime constraintsabstractVariability is a key issue in modern multiprocessors, resulting in performance and lifetime uncertainty, and high design margins. The margins can be reduced by exposing variability to software and then adapting at runtime. In this work we use sensors to monitor the variable operating conditions and the degradation rate. Based on the sensor data, our variability-aware OS scheduling algorithm assigns the workload to the cores and sets the power/performance tradeoffs to meet the mobile processor's lifetime constraints while adjusting to variability and improving the overall performance. We implement our algorithm in Android OS on a mobile phone and show that it achieves up to 160% performance improvement over the state-of-the-art while meeting the lifetime constraints. Pietro Mercati, Francesco Paterna, Andrea Bartolini, Luca Benini, Tajana Rosing |
ICCD | 1 |
| 2013 | Workload and user experience-aware dynamic reliability management in multicore processorsabstractReliability is a major concern for nanoscale CMOS circuits. Degradation phenomena such as Electromigration, Negative Bias Temperature Instability, Time Dependent Dielectric Breakdown worsen with transistor scaling. Dynamic Reliability Management (DRM) techniques reduce reliability loss at runtime by constraining operating points, but they face the challenge of reducing user experience degradation while meeting a lifetime target. In this work we propose a sensor based hierarchical controller for multicore processor DRM, exploiting the major gap between the time scales of workload variations and reliability loss. We improve performance and user experience by locally relaxing reliability-induced operating point constraints, while meeting them over the large time windows relevant for reliability. With respect to the state-of-the-art, our solution guarantees timely execution of 100% of latency-critical applications, and have a 4% performance improvement over the whole lifetime. Pietro Mercati, Andrea Bartolini, Francesco Paterna, Tajana Rosing, Luca Benini |
DAC | 1 |