Diana Trojaniello

dblp:171/7677 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0001-8935-5593ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DQT: Dynamic Quantization Training via Dequantization-Free Nested Integer Arithmetic
abstract
The deployment of deep neural networks on resource-constrained devices relies on quantization. While static, uniform quantization applies a fixed bit-width to all inputs, it fails to adapt to their varying complexity. Dynamic, instance-based mixed-precision quantization promises a superior accuracy-efficiency trade-off by allocating higher precision only when needed. However, a critical bottleneck remains: existing methods require a costly dequantize-to-float and requantize-to-integer cycle to change precision, breaking the integer-only hardware paradigm and compromising performance gains. This paper introduces Dynamic Quantization Training (DQT), a novel framework that removes this bottleneck. At the core of DQT is a nested integer representation where lower-precision values are bit-wise embedded within higher-precision ones. This design, coupled with custom integer-only arithmetic, allows for on-the-fly bit-width switching through a near-zero-cost bit-shift operation. This makes DQT the first quantization framework to enable both dequantization-free static mixed-precision of the backbone network, and truly efficient dynamic, instance-based quantization through a lightweight controller that decides at runtime how to quantize each layer. We demonstrate DQT state-of-the-art performance on ResNet18 on CIFAR-10 and ResNet50 on ImageNet. On ImageNet, our 4-bit dynamic ResNet50 achieves 77.00% top-1 accuracy, an improvement over leading static (LSQ, 76.70%) and dynamic (DQNET, 76.94%) methods at a comparable BitOPs budget. Crucially, DQT achieves this with a bit-width transition cost of only 28.3M simple bit-shift operations, a drastic improvement over the 56.6M costly Multiply-Accumulate (MAC) floating-point operations required by previous dynamic approaches - unlocking a new frontier in efficient, adaptive AI.
Hazem Hesham Yousef Shalby, Fabrizio Pittorino, Francesca Palermo, Diana Trojaniello, Manuel Roveri
AAAI4
2026 Energy-efficient Dynamic Partitioning and Tensors Compression of AI Applications in Smart Eyewears
abstract
Resource-constrained smart eyewear (SEW) devices face significant challenges when deploying deep neural networks due to limited computational capacity and battery life. Computational offloading to companion devices like smartphones and cloud servers addresses processing limitations, but data transmission becomes a critical bottleneck, consuming over 50% of total energy in some scenarios. Although lossless compression methods provide limited data reduction for intermediate tensors, lossy techniques such as Vector Quantization (VQ) offer higher compression ratios (requiring only 3.3 bits per float) at the expense of inference accuracy degradation. This paper presents an adaptive multi-stage compression framework that dynamically balances these trade-offs across the SEW-phone-cloud continuum. We employ VQ at the SEW-phone interface where aggressive compression is essential (achieving 89.6% tensor size reduction with 90% retained accuracy), followed by adaptive selection between quantization and run-length encoding for phone-to-cloud transmission based on network conditions. A Deep Q-Network (DQN) agent jointly optimizes network partitioning points and compression strategies to minimize energy consumption while preserving accuracy and meeting latency constraints. A large simulation campaign considering object detection and human pose estimation tasks demonstrate that our method achieves 55--70% energy savings and 86--91% violation reduction compared to Neurosurgeon (a dynamic partitioning baseline without compression), 45.8% energy savings versus local execution, and 61.1% savings over uncompressed offloading, with latency violation rates below 9% and acceptable accuracy loss (8.0--8.1%). These results enable practical deployment of AI applications on battery-limited SEW devices.
Abednego Wamuhindo Kambale, Samin Shokrivahed, Giacomo Verticale, Francesca Palermo, Diana Trojaniello, Danilo Ardagna
ICPE5
2025 Benchmarking Energy and Latency in TinyML: A Novel Method for Resource-Constrained AI
abstract
The rise of IoT has increased the need for on-edge machine learning, with tinyML emerging as a promising solution for resource-constrained devices such as Microcontroller Units (MCUs). However, benchmarking their energy efficiency, latency, and computational capabilities remains challenging due to diverse architectures and application scenarios. Current solutions, such as MLCommons’ “TinyML Perf: Inference” method, have limitations, including the need for separate setups for latency, accuracy, and energy measurements, as well as reliance on energy monitors that power the device, reducing flexibility. Moreover, the absence of a clear distinction between inference and ancillary operations can compromise the accuracy of performance estimations. This work introduces an alternative benchmarking methodology that integrates energy and latency measurements while distinguishing three execution phases—pre-inference, inference, and post-inference—to enable precise profiling. A dual-trigger approach is used to separate these phases, ensuring accurate and repeatable measurements. Additionally, the setup ensures that the device operates without being powered by an external measurement unit, while automated testing can be leveraged to enhance statistical significance. To evaluate our setup, we tested the STM32N6 MCU, which includes a Neural Processing Unit (NPU) for executing convolutional neural networks. Two configurations were considered: one running at maximum performance with the highest clock frequencies and core voltage, and another with reduced settings. The variation of the Energy Delay Product (EDP) was analyzed separately for each phase, providing insights into the impact of hardware configurations on energy efficiency. Each MLPerf model was tested 1000 times to ensure statistically relevant results. Our findings demonstrate that reducing the core voltage and clock frequency improves the efficiency of pre- and post-processing without significantly affecting network execution performance. This approach can also be used for cross-platform comparisons to determine the most efficient inference platform and to quantify how pre- and post-processing overhead varies across different hardware implementations.
Pietro Bartoli, Christian Veronesi, Andrea Giudici, David Siorpaes, Diana Trojaniello, Franco Zappa
IJCNN5
2025 Maximizing Eye Segmentation Accuracy with YOLO
abstract
Semantic eye segmentation has proven useful in different fields, including biometrics, eye-tracking, and physiological signal extraction. For this reason, identifying the areas of an image regarding the three main parts of the eye surface (sclera, iris, pupil) is of utmost importance, as it allows one to extract various meaningful information. Most studies found in the literature focus on the development of custom Deep Learning models, with an architecture tailored specifically for eye segmentation. Moreover, for tasks such as eye-tracking, the iris and the pupil are often modeled as ellipses, thus including sections of the skin in the segmented area. In this paper, we propose a method for eye segmentation based on You Only Look Once (YOLO) models, generally employed for object detection, solving the previously described issues. First, a large version of YOLOv11 (YOLOv11L-seg) was used to perform a semiautomatic labeling of the dataset, starting from a reduced number of manually generated masks. The so-obtained dataset was then used to train two different versions of YOLO (YOLOv8 nano and YOLOv11 nano), much more compact than the model used for semi-automated labeling. The two models achieved a mean Intersection Over Union (mIOU) over the three classes of 93% and 92%, respectively, while also presenting high generalization capabilities over data acquired with different hardware compared to the current state-of-the-art models.
Arianna De Vecchi, Sabrina Azzi, Pietro Bartoli, Marco Paracchini, Diana Trojaniello, Federica Villa
IJCNN5
2025 Federated Reinforcement Learning for Runtime Optimization of AI Applications in Smart Eyewears
Hamta Sedghani, Abednego Wamuhindo Kambale, Federica Filippini, Francesca Palermo, Diana Trojaniello, Danilo Ardagna
MASCOTS5
2021 Secure store for FHIR resources with Parquet encryption
abstract
The ability for medical professionals to efficiently process vast amounts of data is critical. We present a solution which offers cloud-secure analytics on healthcare data, utilizing FHIR, the latest standard from the HL7 organization, for the exchange of health care data, together with Apache Parquet Modular Encryption.
Eliot E. Salant, Maya Anderson, Diana Trojaniello
SYSTOR3
2021 A design methodology for matching smart health requirements
abstract
Summary As now well established, the world population is aging rapidly and, according to World Health Organization (WHO), the amount of people aged 60 years and older is expected to total 2 billion in 2050. For this reason, an emerging important issue is the definition of a new generation of healthcare platforms capable of monitoring people's quality of life. In this article, we propose a new methodology that supports the entire requirement elicitation process starting from the initial phase of gathering the requirements, both clinical, technological, and end‐user, up to the choice of the most suitable solution. Our proposal provides a new new iterative model in the smart healthcare field research area. Furthermore we apply our proposal in a real scenario and we report the end‐to‐end implementation of the proposed methodology.
Valerio Bellandi, Paolo Ceravolo, Alessia Cristiano, Ernesto Damiani, Alberto Sanna, Diana Trojaniello
Concurr. Comput. Pract. Exp.6