Sara Moccia

dblp:188/3835 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
11since 2021 · last 2025
0000-0002-4494-8907ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Online Knowledge Distillation and Deep Supervision in HRNet: Green AI for Preterm Infants' Pose Estimation
abstract
The current approach to deep-learning research is exemplified by the pursuit of red AI models—designs that show increasingly high performance but with increasingly high costs, in terms of economic requirements and environmental footprint. This approach is particularly detrimental in sectors like healthcare, which typically have limited resources. Meanwhile, green AI prioritizes efficiency and sustainability, by reducing the environmental footprint and making advanced technologies accessible. Following the green AI principles, this study focuses on the combined use of two techniques, namely Knowledge Distillation (KD) and Deep Supervision (DS), to reduce the costs of HRNet, a convolutional neural network designed for human pose estimation, here applied to support the diagnosis of neurological impairments in preterm infants. All the experiments are carried out on the BabyPose dataset, a collection of videos from a depth camera showing hospitalized preterm infants. By combining KD and DS, we can use a sub-network of HRNet that needs only 27.5% of its parameters and 61.7% of its FLOPs, without affecting performance (−0.59 percentage points in average precision). This achievement can have deep implications in the actual clinical practice, as it fosters democratization of high-quality technologies. Our codes are available at https://github.com/geronimaw/OnlineKD-HRNet-Human-Pose-Estimation.git .
Alessandro Cacciatore, Daniele Berardini, Vito Scaraggi, Adriano Mancini, Sara Moccia, Lucia Migliorelli
ACM Trans. Comput. Heal.5
2024 Placental vessel segmentation and registration in fetoscopy: Literature review and MICCAI FetReg2021 challenge findings
abstract
Fetoscopy laser photocoagulation is a widely adopted procedure for treating Twin-to-Twin Transfusion Syndrome (TTTS). The procedure involves photocoagulation pathological anastomoses to restore a physiological blood exchange among twins. The procedure is particularly challenging, from the surgeon's side, due to the limited field of view, poor manoeuvrability of the fetoscope, poor visibility due to amniotic fluid turbidity, and variability in illumination. These challenges may lead to increased surgery time and incomplete ablation of pathological anastomoses, resulting in persistent TTTS. Computer-assisted intervention (CAI) can provide TTTS surgeons with decision support and context awareness by identifying key structures in the scene and expanding the fetoscopic field of view through video mosaicking. Research in this domain has been hampered by the lack of high-quality data to design, develop and test CAI algorithms. Through the Fetoscopic Placental Vessel Segmentation and Registration (FetReg2021) challenge, which was organized as part of the MICCAI2021 Endoscopic Vision (EndoVis) challenge, we released the first large-scale multi-center TTTS dataset for the development of generalized and robust semantic segmentation and video mosaicking algorithms with a focus on creating drift-free mosaics from long duration fetoscopy videos. For this challenge, we released a dataset of 2060 images, pixel-annotated for vessels, tool, fetus and background classes, from 18 in-vivo TTTS fetoscopy procedures and 18 short video clips of an average length of 411 frames for developing placental scene segmentation and frame registration for mosaicking techniques. Seven teams participated in this challenge and their model performance was assessed on an unseen test dataset of 658 pixel-annotated images from 6 fetoscopic procedures and 6 short clips. For the segmentation task, overall baseline performed was the top performing (aggregated mIoU of 0.6763) and was the best on the vessel class (mIoU of 0.5817) while team RREB was the best on the tool (mIoU of 0.6335) and fetus (mIoU of 0.5178) classes. For the registration task, overall the baseline performed better than team SANO with an overall mean 5-frame SSIM of 0.9348. Qualitatively, it was observed that team SANO performed better in planar scenarios, while baseline was better in non-planner scenarios. The detailed analysis showed that no single team outperformed on all 6 test fetoscopic videos. The challenge provided an opportunity to create generalized solutions for fetoscopic scene understanding and mosaicking. In this paper, we present the findings of the FetReg2021 challenge, alongside reporting a detailed literature review for CAI in TTTS fetoscopy. Through this challenge, its analysis and the release of multi-center fetoscopic data, we provide a benchmark for future research in this field.
Sophia Bano, Alessandro Casella, Francisco Vasconcelos 0001, Abdul Qayyum 0002, Abdessalam Benzinou, Moona Mazher, Fabrice Mériaudeau, Chiara Lena, Ilaria A. Cintorrino, Gaia Romana De Paolis, Jessica Biagioli, Daria Grechishnikova, Jing Jiao, Bizhe Bai, Yanyan Qiao, Binod Bhattarai, Rebati Raman Gaire, Ronast Subedi, Eduard Vazquez, Szymon Plotka, Aneta Lisowska, Arkadiusz Sitek, George Attilakos, Ruwan Wimalasundera, Anna L. David, Dario Paladini, Jan Deprest, Elena De Momi, Leonardo S. Mattos, Sara Moccia, Danail Stoyanov
Medical Image Anal.30
2024 A deep-learning framework running on edge devices for handgun and knife detection from indoor video-surveillance cameras
abstract
Abstract The early detection of handguns and knives from surveillance videos is crucial to enhance people’s safety. Despite the increasing development of Deep Learning (DL) methods for general object detection, weapon detection from surveillance videos still presents open challenges. Among these, the most significant are: (i) the very small size of the weapons with respect to the camera field of view and (ii) the need of a real-time feedback, even when using low-cost edge devices for computation. Complex and recently-developed DL architectures could mitigate the former challenge but do not satisfy the latter one. To tackle such limitation, the proposed work addresses the weapon-detection task from an edge perspective. A double-step DL approach was developed and evaluated against other state-of-the-art methods on a custom indoor surveillance dataset. The approach is based on a first Convolutional Neural Network (CNN) for people detection which guides a second CNN to identify handguns and knives. To evaluate the performance in a real-world indoor environment, the approach was deployed on a NVIDIA Jetson Nano edge device which was connected to an IP camera. The system achieved near real-time performance without relying on expensive hardware. The results in terms of both COCO Average Precision (AP = 79.30) and Frames per Second (FPS = 5.10) on the low-power NVIDIA Jetson Nano pointed out the goodness of the proposed approach compared with the others, encouraging the spread of automated video surveillance systems affordable to everyone.
Daniele Berardini, Lucia Migliorelli, Alessandro Galdelli, Emanuele Frontoni, Adriano Mancini, Sara Moccia
Multim. Tools Appl.6
2023 Deep learning-based approaches for human motion decoding in smart walkers for rehabilitation
Carolina Gonçalves, João M. Lopes, Sara Moccia, Daniele Berardini, Lucia Migliorelli, Cristina P. Santos 0001
Expert Syst. Appl.3
2023 A review on deep-learning algorithms for fetal ultrasound-image analysis
Maria Chiara Fiorentino, Francesca Pia Villani, Mariachiara Di Cosmo, Emanuele Frontoni, Sara Moccia
Medical Image Anal.5
2023 Beyond rankings: Learning (more) from algorithm validation
abstract
Challenges have become the state-of-the-art approach to benchmark image analysis algorithms in a comparative manner. While the validation on identical data sets was a great step forward, results analysis is often restricted to pure ranking tables, leaving relevant questions unanswered. Specifically, little effort has been put into the systematic investigation on what characterizes images in which state-of-the-art algorithms fail. To address this gap in the literature, we (1) present a statistical framework for learning from challenges and (2) instantiate it for the specific task of instrument instance segmentation in laparoscopic videos. Our framework relies on the semantic meta data annotation of images, which serves as foundation for a General Linear Mixed Models (GLMM) analysis. Based on 51,542 meta data annotations performed on 2,728 images, we applied our approach to the results of the Robust Medical Instrument Segmentation Challenge (ROBUST-MIS) challenge 2019 and revealed underexposure, motion and occlusion of instruments as well as the presence of smoke or other objects in the background as major sources of algorithm failure. Our subsequent method development, tailored to the specific remaining issues, yielded a deep learning model with state-of-the-art overall performance and specific strengths in the processing of images in which previous methods tended to fail. Due to the objectivity and generic applicability of our approach, it could become a valuable tool for validation in the field of medical image analysis and beyond.
Tobias Roß, Pierangela Bruno, Annika Reinke, Manuel Wiesenfarth, Lisa Koeppel, Peter M. Full, Bünyamin Pekdemir, Patrick Godau, Darya Trofimova, Fabian Isensee, Tim Adler, Thuy Nuong Tran, Sara Moccia, Francesco Calimeri, Beat P. Müller-Stich, Annette Kopp-Schneider, Lena Maier-Hein
Medical Image Anal.13
2022 Autonomous Intraluminal Navigation of a Soft Robot using Deep-Learning-based Visual Servoing
abstract
Navigation inside luminal organs is an arduous task that requires non-intuitive coordination between the movement of the operator's hand and the information obtained from the endoscopic video. The development of tools to automate certain tasks could alleviate the physical and mental load of doctors during interventions allowing them to focus on diagnosis and decision-making tasks. In this paper we present a synergic solution for intraluminal navigation consisting of a 3D printed endoscopic soft robot that can move safely inside luminal structures. Visual servoing based on Convolutional Neural Networks (CNNs) is used to achieve the autonomous navigation task. The CNN is trained with phantoms and in-vivo data to segment the lumen and a model-less approach is presented to control the movement in constrained environments. The proposed robot is validated in anatomical phantoms in different path configurations. We analyze the movement of the robot using different metrics such as task completion time smoothness error in the steady-state mean and maximum error. We show that our method is suitable to navigate safely in hollow environments and conditions which are different than the ones the network was originally trained on.
Jorge F. Lazo, Chun-Feng Lai, Sara Moccia, Benoit Rosa, Michele Catellani, Michel de Mathelin, Giancarlo Ferrigno, Paul Breedveld, Jenny Dankelman, Elena De Momi
IROS3
2022 An accurate estimation of preterm infants' limb pose from depth images using deep neural networks with densely connected atrous spatial convolutions
Lucia Migliorelli, Emanuele Frontoni, Sara Moccia
Expert Syst. Appl.3
2021 End-to-end semantic joint detection and limb-pose estimation from depth images of preterm infants in NICUs
abstract
Continuous evaluation of preterm infants' spontaneous motility proved to be a decisive tool for timely diagnosing the presence of neurodevelopmental disorders. Automatic infants' limbs pose estimation is a powerful ally to support clinicians in infant's monitoring. This work proposes an end-to-end pipeline for limb-pose estimation based on a region-based convolutional neural network, named Mask R-CNN. The framework was validated on a custom dataset of 6000 depth images from 30 videos of 19 preterm infants acquired in a neonatal intensive care unit during the actual clinical practice. Leave-one-infant-out cross-validation was performed to evaluate the framework performance. Results for joints' detection showed a mean average precision equal to 0.9 with a standard deviation of 0.2. For limb-pose estimation, median root mean square error [pixel] was equal to 6.8 (right arm), 6.7 (left arm), 6.5 (right leg), 6.5 (left leg). The interquartile ranges [pixels] were 1.1, 1.2, 0.6, 1.2 for each limb, respectively. This end - to-end framework represents a step toward embedded monitoring solutions for on-the-edge computation.
Matteo Carbonari, Greta Vallasciani, Lucia Migliorelli, Emanuele Frontoni, Sara Moccia
ISCC5
2021 Real-time human pose estimation on a smart walker using convolutional neural networks
Manuel Palermo, Sara Moccia, Lucia Migliorelli, Emanuele Frontoni, Cristina P. Santos 0001
Expert Syst. Appl.2
2021 A shape-constraint adversarial framework with instance-normalized spatio-temporal features for inter-fetal membrane segmentation
abstract
BACKGROUND AND OBJECTIVES: During Twin-to-Twin Transfusion Syndrome (TTTS), abnormal vascular anastomoses in the monochorionic placenta can produce uneven blood flow between the fetuses. In the current practice, this syndrome is surgically treated by closing the abnormal connections using laser ablation. Surgeons commonly use the inter-fetal membrane as a reference. Limited field of view, low fetoscopic image quality and high inter-subject variability make the membrane identification a challenging task. However, currently available tools are not optimal for automatic membrane segmentation in fetoscopic videos, due to membrane texture homogeneity and high illumination variability. METHODS: To tackle these challenges, we present a new deep-learning framework for inter-fetal membrane segmentation on in-vivo fetoscopic videos. The framework enhances existing architectures by (i) encoding a novel (instance-normalized) dense block, invariant to illumination changes, that extracts spatio-temporal features to enforce pixel connectivity in time, and (ii) relying on an adversarial training, which constrains macro appearance. RESULTS: We performed a comprehensive validation using 20 different videos (2000 frames) from 20 different surgeries, achieving a mean Dice Similarity Coefficient of 0.8780±0.1383. CONCLUSIONS: The proposed framework has great potential to positively impact the actual surgical practice for TTTS treatment, allowing the implementation of surgical guidance systems that can enhance context awareness and potentially lower the duration of the surgeries.
Alessandro Casella, Sara Moccia, Dario Paladini, Emanuele Frontoni, Elena De Momi, Leonardo S. Mattos
Medical Image Anal.2
2020 NephCNN: A deep-learning framework for vessel segmentation in nephrectomy laparoscopic videos
abstract
Objective: In the last years, Robot-assisted partial nephrectomy (RAPN) is establishing as elected treatment for renal cell carcinoma (RCC). Reduced field of view, field occlusions by surgical tools, and reduced maneuverability may potentially cause accidents, such as unwanted vessel resection with consequent bleeding. Surgical Data Science (SDS) can provide effective context-aware tools for supporting surgeons. However, currently no tools have been exploited for automatic vessels segmentation from nephrectomy laparoscopic videos. Herein, we propose a new approach based on adversarial Fully Convolutional Neural Networks (FCNNs) to kidney vessel segmentation from nephrectomy laparoscopic vision. Methods: The proposed approach enhances existing segmentation framework by (i) encoding 3D kernels for spatio-temporal features extraction to enforce pixel connectivity in time, and (ii) perform training in adversarial fashion, which constrains vessels shape. Results: We performed a preliminary study using 8 different RAPN videos (1871 frames), the first in the field, achieving a median Dice Similarity Coefficient of 71.76%. Conclusions: Results showed that the proposed approach could be a valuable solution with a view to assist surgeon during RAPN.
Alessandro Casella, Sara Moccia, Chiara Carlini, Emanuele Frontoni, Elena De Momi, Leonardo S. Mattos
ICPR2
2020 A Lumen Segmentation Method in Ureteroscopy Images based on a Deep Residual U-Net architecture
abstract
U reteroscopy is becoming the first surgical treatment option for the majority of urinary affections. This procedure is performed using an endoscope which provides the surgeon with the visual information necessary to navigate inside the urinary tract. Having in mind the development of surgical assistance systems, that could enhance the performance of surgeon, the task of lumen segmentation is a fundamental part since this is the visual reference which marks the path that the endoscope should follow. This is something that has not been analyzed in ureteroscopy data before. However, this task presents several challenges given the image quality and the conditions itself of ureteroscopy procedures. In this paper, we study the implementation of a Deep Neural Network which exploits the advantage of residual units in an architecture based on U-Net. For the training of these networks, we analyze the use of two different color spaces: gray-scale and RGB data images. We found that training on gray-scale images gives the best results obtaining mean values of Dice Score, Precision, and Recall of 0.73, 0.58, and 0.92 respectively. The results obtained shows that the use of residual U-Net could be a suitable model for further development for a computer-aided system for navigation and guidance through the urinary system.
Jorge F. Lazo, Aldo Marzullo, Sara Moccia, Michele Catellani, Benoit Rosa, Francesco Calimeri, Michel de Mathelin, Elena De Momi
ICPR3
2020 Evaluating the autonomy of children with autism spectrum disorder in washing hands: a deep-learning approach
abstract
Monitoring children with Autism Spectrum Dis-order (ASD) during the execution of the Applied Behaviour Analysis (ABA) program is crucial to assess the progresses while performing actions. Despite its importance, this monitoring procedure still relies on ABA operators’ visual observation and manual annotation of the significant events. In this work a deep learning (DL) based approach has been proposed to evaluate the autonomy of children with ASD while performing the hand-washing task. The goal of the algorithm is the automatic detection of RGB frames in which the ASD child washes his/her hands autonomously (no-aid frames) or is supported by the operator (aid frames). The proposed approach relies on a pre-trained VGG16 convolutional network (CNN) modified to fulfill the binary classification task. The performance of the fine-tuned VGG16 was compared against that of other CNN architectures. The fine-tuned VGG16 achieved the best performance with a recall of 0.92 and 0.89 for the no-aid and aid class, respectively. These results prompt the possibility of translating the presented methodology into the actual monitoring practice. The integration of the presented tool with other computer-aided monitoring systems into a single framework, will provide fully support to ABA operators during the therapy session.
Daniele Berardini, Lucia Migliorelli, Sara Moccia, Marcello Naldini, Gioia De Angelis, Emanuele Frontoni
ISCC3
2019 Preterm infants' limb-pose estimation from depth images using convolutional neural networks
abstract
Preterm infants' limb-pose estimation is a crucial but challenging task, which may improve patients' care and facilitate clinicians in infant's movements monitoring. Work in the literature either provides approaches to whole-body segmentation and tracking, which, however, has poor clinical value, or retrieve a posteriori limb pose from limb segmentation, increasing computational costs and introducing inaccuracy sources. In this paper, we address the problem of limb-pose estimation under a different point of view. We proposed a 2D fully-convolutional neural network for roughly detecting limb joints and joint connections, followed by a regression convolutional neural network for accurate joint and joint-connection position estimation. Joints from the same limb are then connected with a maximum bipartite matching approach. Our analysis does not require any prior modeling of infants' body structure, neither any manual interventions. For developing and testing the proposed approach, we built a dataset of four videos (video length = 90 s) recorded with a depth sensor in a neonatal intensive care unit (NICU) during the actual clinical practice, achieving median root mean square distance [pixels] of 10.790 (right arm), 10.542 (left arm), 8.294 (right leg), 11.270 (left leg) with respect to the ground-truth limb pose. The idea of estimating limb pose directly from depth images may represent a future paradigm for addressing the problem of preterm-infants' movement monitoring and offer all possible support to clinicians in NICUs.
Sara Moccia, Lucia Migliorelli, Rocco Pietrini, Emanuele Frontoni
CIBCB1
2017 Physiological Parameter Estimation from Multispectral Images Unleashed
Sebastian J. Wirkert, Anant Suraj Vemuri, Hannes Kenngott, Sara Moccia, Michael Götz, Benjamin F. B. Mayer, Klaus H. Maier-Hein, Daniel S. Elson, Lena Maier-Hein
MICCAI (3)4