EDBT 2026 Demo / reviewers in the wild / expert
Eleonora D'Arnese
dblp:202/6627
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-6967-5079ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive AIE-PL Systems for Efficient End-to-End Pyramidal 3D Image RegistrationabstractModern accelerators maximize throughput through aggressive specialization. However, in many real-world applications, workloads often vary at runtime, requiring multiple bitstreams to handle such changes. As a result, frequent reconfigurations introduce substantial overhead that can dominate end-to-end execution time. This issue is particularly evident in AIE–PL systems, where statically scheduled AI Engines (AIEs) achieve high performance through compile-time optimization and are therefore typically tailored to fixed workloads. Although AIEs support Runtime Parameters (RTPs) under Processing System (PS) orchestration, RTPs are impractical for discrete hosts. For this reason, we present a structured approach to designing single-bitstream, runtime-adaptable AIE–PL accelerators that does not rely on RTPs, suitable for discrete hosts. We exploit the Programmable Logic (PL) to generate and stream a compact metadata packet that distributes workload configuration across a directed AIE graph before computation. By doing so, we deliberately trade a fraction of fixed-instance efficiency for flexibility. We validate our approach by devising PeterPan, a software-programmable AIE–PL accelerator for 3D image registration. PeterPan supports runtime-varying problem sizes and integrates seamlessly into multi-stage pipelines, such as pyramidal (coarse-to-fine) registration. To maximize PeterPan utilization, we couple it with an ad-hoc software module that employs a novel heuristic to rapidly select informative sub-volumes, keeping the accelerator continuously fed and preventing input-side stalls. On a VCK5000, PeterPan matches state-of-the-art accelerator performance while retaining software programmability. In the end-to-end task, instead, PeterPan delivers a 3.06× speedup and a 2.74× higher energy efficiency than the state-of-the-art AIE-PL accelerator. Giuseppe Sorrentino, Paolo Salvatore Galfano, Claudio Di Salvo, Eleonora D'Arnese, Davide Conficconi |
FCCM | 4 |
| 2025 | Soaring with TRILLI: An HW/SW Heterogeneous Accelerator for Multi-Modal Image Registrationabstract3D rigid image registration is a pivotal procedure in computer vision that aligns a floating volume with a reference one to correct positional and rotational distortions. It serves either as a stand-alone process or as a pre-processing step for non-rigid registration, where the rigid part dominates the computational cost. Various hardware accelerators have been proposed to optimize its compute-intensive components: geometric transformation with interpolation and similarity metric computation. However, existing solutions fail to address both components effectively, as GPUs excel at image transformation, while FPGAs in similarity metric computation. To close this gap, we propose TRILLI, a novel Versal-based accelerator for image transformation and interpolation. TRILLI optimally maps each computational step on the proper heterogeneous hardware component. TRILLI achieves speedup of 5.32× against the top performing GPU-based solution, and an energy efficiency improvement of 36.75 × against the most efficient one. Moreover, we integrate it with an FPGA-based similarity metric from literature to complete a rigid image registration step (i.e., transformation, interpolation, and similarity metric) attaining a speedup of 18.60 × against the top performing GPU-based solution, while being 36.11 ×more efficient than the most energy efficient one. Giuseppe Sorrentino, Paolo Salvatore Galfano, Eleonora D'Arnese, Davide Conficconi |
FCCM | 3 |
| 2025 | A Multiscale Attention-Based Deep Learning Method for DCE-MRI Breast Tumor SegmentationabstractBreast cancer, the most diagnosed cancer among women, demands accurate diagnosis for effective treatment. Dynamic Contrast-Enhanced Magnetic Resonance Imaging (DCE-MRI) provides detailed spatial insights into tissue characteristics, making it essential for tumor analysis. However, segmenting circumscribed mass and diffuse non-mass lesions remains a challenge, as existing solutions often rely on single-context images or multiscale approaches that overlook the imaged structural organization of breast anatomy. To address these gaps, we propose VENUS, a multiscale image and feature attention-based network for breast tumor segmentation in DCE-MRI. Inspired by physicians’ image inspection routine, it utilizes a multiscale encoder with Convolutional Feature Fusion Blocks utilizing early fusion to combine full-breast views with detailed single-breast zoom-ins and improve cross-context semantic modeling ability. A novel attention-based decoder with Attention Gating enhances skip connections by prioritizing critical features for accurate reconstruction. Experiments reveal significant performance gains of 11.05% and 15.07% Dice Similarity Coefficient over single-context state-of-the-art methods on clinical and public DCE-MRI datasets, respectively. Pablo Giaccaglia, Isabella Poles, Valentina Lidoni, Veronica Rizzo, Michele Gentili, Federica Pediconi, Marco D. Santambrogio, Eleonora D'Arnese |
ICIP | 8 |
| 2025 | VOTED - Versal Optimization Toolkit for Education and Heterogeneous Systems DevelopmentabstractDespite classic educational approaches proving their effectiveness for classic hardware acceleration, they suffer novel heterogeneous systems such as Versal, requiring a deeper system-level awareness. Versal devices merge FPGA on-field programma-bility with the performance of hardened VLIW processors, namely AI Engine, at the cost of facing novel challenges when integrating these two different layers. Given the system complexity of such devices and the low-level knowledge required to leverage them, Versal potentialities have yet to be fully exploited. Therefore, we present VOTED, a Versal Optimization Toolkit for Education and Heterogeneous Systems Development, that guides users of any expertise in discovering, using, and optimizing Versal-based applications. VOTED proved to be effective, leading five different groups of students, at their first experience with heterogeneous system design, at devising quite complex Versal-based applications capable of reaching the final stages of a design competition and even winning it with four months of work. Giuseppe Sorrentino, Paolo Salvatore Galfano, Eleonora D'Arnese, Davide Conficconi |
ISCAS | 3 |
| 2024 | Co-Designing a 3D Transformation Accelerator for Versal-Based Image RegistrationabstractRigid image registration is pivotal in modern imaging for correcting distortions of images acquired with different modalities or at different time instants. Literature accelerates its main compute-intensive steps, image transformation and similarity metric computation, through GPUs or FPGAs to meet performance-efficiency constraints. However, GPUs lack energy efficiency while FPGAs lack performance for image transformation. Therefore, we adopt a single heterogeneous system through PEGASO, a methodology and its implementation to co-design the image transformation algorithm on Versal system. We maximize performance by combining custom data layout and hardware optimizations, attaining a 19x speedup over the best GPU-based transformation accelerator. When integrated with FPGA-based similarity metric, PEGASO achieves 82.73x and 20.73x speedup against FPGA- and GPU-based solutions while improving the corresponding energy efficiency of 318.18x and 52.62x. Paolo Salvatore Galfano, Giuseppe Sorrentino, Eleonora D'Arnese, Davide Conficconi |
ICCD | 3 |
| 2024 | Letting Osteocytes Teach SR-MicroCT Bone Lacunae Segmentation: A Feature Variation Distillation Method via Diffusion Denoising
Isabella Poles, Marco D. Santambrogio, Eleonora D'Arnese |
MICCAI (9) | 3 |
| 2024 | Starlight: A kernel optimizer for GPU processingabstractOver the past few years, GPUs have found widespread adoption in many scientific domains, offering notable performance and energy efficiency advantages compared to CPUs. However, optimizing GPU high-performance kernels poses challenges given the complexities of GPU architectures and programming models. Moreover, current GPU development tools provide few high-level suggestions and overlook the underlying hardware. Here we present Starlight, an open-source, highly flexible tool for enhancing GPU kernel analysis and optimization. Starlight autonomously describes Roofline Models, examines performance metrics, and correlates these insights with GPU architectural bottlenecks. Additionally, Starlight predicts potential performance enhancements before altering the source code. We demonstrate its efficacy by applying it to literature genomics and physics applications, attaining speedups from 1.1× to 2.5× over state-of-the-art baselines. Furthermore, Starlight supports the development of new GPU kernels, which we exemplify through an image processing application, showing speedups of 12.7× and 140× when compared against state-of-the-art FPGA- and GPU-based solutions. Alberto Zeni, Emanuele Del Sozzo, Eleonora D'Arnese, Davide Conficconi, Marco D. Santambrogio |
J. Parallel Distributed Comput. | 3 |
| 2024 | NERONE: The Fast Way to Efficiently Execute Your Deep Learning Algorithm at the EdgeabstractSemantic segmentation and classification are pivotal in many clinical applications, such as radiation dose quantification and surgery planning. While manually labeling images is highly time-consuming, the advent of Deep Learning (DL) has introduced a valuable alternative. Nowadays, DL models inference is run on Graphics Processing Units (GPUs), which are power-hungry devices, and, therefore, are not the most suited solution in constrained environments where Field Programmable Gate Arrays (FPGAs) become an appealing alternative given their remarkable performance per watt ratio. Unfortunately, FPGAs are hard to use for non-experts, and the creation of tools to open their employment to the computer vision community is still limited. For these reasons, we propose NERONE, which allows end users to seamlessly benefit from FPGA acceleration and energy efficiency without modifying their DL development flows. To prove the capability of NERONE to cover different network architectures, we have developed four models, one for each of the chosen datasets (three for segmentation and one for classification), and we deployed them, thanks to NERONE, on three different embedded FPGA-powered boards achieving top average energy efficiency improvements of 3.4× and 1.9× against a mobile and a datacenter GPU devices, respectively. Raffaele Berzoini, Eleonora D'Arnese, Davide Conficconi, Marco D. Santambrogio |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Hephaestus: Codesigning and Automating 3D Image Registration on Reconfigurable ArchitecturesabstractHealthcare is a pivotal research field, and medical imaging is crucial in many applications. Therefore finding new architectural and algorithmic solutions would benefit highly repetitive image processing procedures. One of the most complex tasks in this sense is image registration, which finds the optimal geometric alignment among 3D image stacks and is widely employed in healthcare and robotics. Given the high computational demand of such a procedure, hardware accelerators are promising real-time and energy-efficient solutions, but they are complex to design and integrate within software pipelines. Therefore, this work presents an automation framework called Hephaestus that generates efficient 3D image registration pipelines combined with reconfigurable accelerators. Moreover, to alleviate the burden from the software, we codesign software-programmable accelerators that can adapt at run-time to the image volume dimensions. Hephaestus features a cross-platform abstraction layer that enables transparently high-performance and embedded systems deployment. However, given the computational complexity of 3D image registration, the embedded devices become a relevant and complex setting being constrained in memory; thus, they require further attention and tailoring of the accelerators and registration application to reach satisfactory results. Therefore, with Hephaestus , we also propose an approximation mechanism that enables such devices to perform the 3D image registration and even achieve, in some cases, the accuracy of the high-performance ones. Overall, Hephaestus demonstrates 1.85× of maximum speedup, 2.35× of efficiency improvement with respect to the State of the Art, a maximum speedup of 2.51× and 2.76× efficiency improvements against our software, while attaining state-of-the-art accuracy on 3D registrations. Giuseppe Sorrentino, Marco Venere, Davide Conficconi, Eleonora D'Arnese, Marco D. Santambrogio |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2023 | Faber: A Hardware/SoftWare Toolchain for Image RegistrationabstractImage registration is a well-defined computation paradigm widely applied to align one or more images to a target image. This paradigm, which builds upon three main components, is particularly compute-intensive and represents many image processing pipelines’ bottlenecks. State-of-the-art solutions leverage hardware acceleration to speed up image registration, but they are usually limited to implementing a single component. We present Faber, an open-source HW/SW CAD toolchain tailored to image registration. The Faber toolchain comprises HW/SW highly-tunable registration components, supports users with different expertise in building custom pipelines, and automates the design process. In this direction, Faber provides both default settings for entry-level users and latency and resource models to guide HW experts in customizing the different components. Finally, Faber achieves from 1.5× to 54× in speedup and from 2× to 177× in energy efficiency against state-of-the-art tools on a Xeon Gold. Eleonora D'Arnese, Davide Conficconi, Emanuele Del Sozzo, Luigi Fusco, Donatella Sciuto, Marco D. Santambrogio |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | On the Automation of Radiomics-Based Identification and Characterization of NSCLCabstractProper detection and accurate characterization of Non-Small Cell Lung Cancer (NSCLC) are an open challenge in the imaging field. Biomedical imaging is fundamental in lung cancer assessment and offers the possibility of calculating predictive biomarkers impacting patients' management. Within this context, radiomics, which consists of extracting quantitative features from digital images, shows encouraging results for clinical applications, but the sub-optimal standardization of the procedure and the lack of definitive results are still a concern in the field. For these reasons, this work proposes the design and development of LuCIFEx, a fully-automated pipeline for non-invasive in-vivo characterization of NSCLC, aiming to speed up the analysis process and enable an early diagnosis of the tumor.LuCIFEx pipeline relies on routinely acquired [18F]FDG-PET/CT images for the automatic segmentation of the cancer lesion, allowing the computation of accurate radiomic features, then employed for cancer characterization through Machine Learning algorithms. The proposed multi-stage segmentation process can identify the lesion with a mean accuracy of 94.2±5.0%. Finally, the proposed data analysis pipeline demonstrates the potential of PET/CT features for the automatic recognition of lung metastases and NSCLC histological subtypes, while highlighting the main current limitations of the radiomic approach. Eleonora D'Arnese, Guido Walter Di Donato, Emanuele Del Sozzo, Martina Sollini, Donatella Sciuto, Marco D. Santambrogio |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | A Framework for Customizable FPGA-based Image Registration AcceleratorsabstractImage Registration is a highly compute-intensive optimization procedure that determines the geometric transformation to align a floating image to a reference one. Generally, the registration targets are images taken from different time instances, acquisition angles, and/or sensor types. Several methodologies are employed in the literature to address the limiting factors of this class of algorithms, among which hardware accelerators seem the most promising solution to boost performance. However, most hardware implementations are either closed-source or tailored to a specific context, limiting their application to different fields. For these reasons, we propose an open-source hardware-software framework to generate a configurable architecture for the most compute-intensive part of registration algorithms, namely the similarity metric computation. This metric is the Mutual Information, a well-known calculus from the Information Theory, used in several optimization procedures. Through different design parameters configurations, we explore several design choices of our highly-customizable architecture and validate it on multiple FPGAs. We evaluated various architectures against an optimized Matlab implementation on an Intel Xeon Gold, reaching a speedup up to 2.86x, and remarkable performance and power efficiency against other state-of-the-art approaches. Davide Conficconi, Eleonora D'Arnese, Emanuele Del Sozzo, Donatella Sciuto, Marco D. Santambrogio |
FPGA | 2 |