Cosimo Patruno

dblp:172/3161 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0001-8624-5444ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Analysis of Input Data Configurations in CNN-based Human Action Recognition for Assembly Task
abstract
Human Action Recognition (HAR) plays a vital role in manufacturing assembly tasks, addressing key areas such as worker safety, operational support, production optimization, employee training, and facilitating human-robot collaboration. This paper introduces a skeleton-based action recognition approach based on a CNN deep neural network architecture. Joint-to-joint distances are used to represent human movements during assembly tasks, enabling the model to capture intricate motion patterns. The primary focus of this work is on structuring the input data in various ways to analyze how these variations influence the network performance. Studying the spatial configurations of input data for human action recognition in an assembly task is an insightful and challenging research topic. In assembly tasks, the high similarity between actions and the operator-specific execution variations make distinguishing actions more complex. This work investigates how the arrangement of input data impacts model accuracy. In particular, two input data configurations are analyzed: onechannel and multi-channel types. The assembly actions are classified using a CNN-based architecture. So, the different data configurations directly influence the type of CNN applied, which can be 2D or 3D. The proposed approach is evaluated on the publicly available HA4M dataset. The obtained results showed that the proposed data structure greatly influences the model performance measurement.
Cosimo Patruno, Grazia Cicirelli, Laura Romeo, Tiziana D'Orazio
CoDIT1
2025 Deep Learning Methods with Iterative-Boosting for performing Human Action Recognition in Manufacturing Scenarios
L. Romeo, Cosimo Patruno, Grazia Cicirelli, Tiziana D'Orazio
CoDIT2
2025 Multi-View Skeleton Analysis for Human Action Segmentation Tasks
Laura Romeo, Cosimo Patruno, Grazia Cicirelli, Tiziana D'Orazio
ICPRAM2
2025 Transformer-based Human Action Recognition for Fine-Grained Industrial Assembly Tasks
abstract
Human Action Recognition (HAR) in industrial assembly scenarios presents significant challenges, primarily due to the slight differences in motion patterns across fine-grained actions. In this work, we address the problem of action recognition in assembly tasks by employing skeleton data to represent detailed human movements. To effectively capture the spatial and temporal dependencies among joints, we apply a Transformer-based architecture. We conducted an extensive evaluation by varying the model dimension to analyze its effect on recognition performance. Furthermore, given the high semantic and structural similarity between certain action classes, we propose a class merging strategy that combines highly similar actions into unified categories. This not only simplifies the classification task, but also improves overall recognition performance by reducing ambiguity. Experimental results demonstrate the effectiveness of Transformers for fine-grained action recognition in industrial settings, highlighting the importance of both architectural tuning and label refinement when dealing with closely related human actions.
Mayra Vanessa Alvear Gallón, Cosimo Patruno, Gadea Mata, César Domínguez 0001, Grazia Cicirelli
IECON2
2019 People re-identification using skeleton standard posture and color descriptors from RGB-D data
Cosimo Patruno, Roberto Marani, Grazia Cicirelli, Ettore Stella, Tiziana D'Orazio
Pattern Recognit.1
2015 An Embedded Vision System for Real-Time Autonomous Localization Using Laser Profilometry
abstract
In this paper, we propose an embedded vision system based on laser profilometry able to get the pose of a vehicle and its relative displacements with reference to the constitutive media of a structured environment. Fundamental equations for laser triangulation are developed and encoded for their actual implementation on an embedded system. It is made of a laser source that projects a line-shaped beam onto the environment and an on-chip camera able to frame the laser light. Images are then sent to the inexpensive Raspberry Pi onboard computer, which is responsible for processing tasks. For the first time, laser profilometry is coupled with the correlation of laser signatures on a low-cost and low-resource processing board for vehicle localization purposes. Several validation tests of the proposed sensor have proven the effectiveness of the system with respect to commercially available sensors such as inductive sensors and standard odometers, which fail when the vehicle crosses path interceptions or its wheels undergo unavoidable slippages. Moreover, further comparisons with other vision-based techniques have also proven the good performances of this embedded system for real-time localization of vehicles.
Cosimo Patruno, Roberto Marani, Massimiliano Nitti, Tiziana D'Orazio, Ettore Stella
IEEE Trans. Intell. Transp. Syst.1