David Ortiz-Perez

dblp:331/4117 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
0009-0008-4890-8217ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Multimodal pain assessment with transformers
abstract
Pain assessment is a critical challenge in healthcare, requiring accurate and objective measurement to enhance patient care. Traditional methods rely on subjective self-reporting, which lacks reliability, particularly for patients with communication difficulties. This work presents PainFusion+, a multimodal transformer architecture for pain assessment that integrates physiological signals and facial expressions to improve pain assessment accuracy. For physiological signals, our approach first employs convolutional neural networks to extract local patterns from short signal fragments, capturing essential pain-related features. These localized representations are then processed by a transformer encoder, which models long-range dependencies to form a comprehensive global representation. For facial video data, we leverage a frozen video transformer to extract expressive features without requiring fine-tuning, significantly reducing computational costs. Finally, both feature spaces are fused using a transformer encoder, allowing effective cross-modal learning. Experiments on publicly available datasets demonstrate that PainFusion+ outperforms existing models. In biomedical signal processing, our method achieves over a 16 % improvement in accuracy. For multimodal pain estimation, it achieves 35.40 % accuracy on the BioVid dataset, setting a new state-of-the-art benchmark.
Manuel Benavent-Lledó, Maria Dolores Lopez-Valle, David Ortiz-Perez, David Mulero-Pérez, José García Rodríguez 0001, Alexandra Psarrou
Neurocomputing3
2025 HierADL: Automatic Hierarchical Structures to Improve Daily Activity Recognition
abstract
Human action recognition is a widely explored problem in computer vision, with applications in domains such as robotics and surveillance. While most existing methods focus on predicting a single action class, human actions are not isolated events but range from simple movements to complex behaviors. Recent work has approached action recognition from a hierarchical perspective, relying on manually annotated labels available in some domains. In contrast, the Activities of Daily Living (ADL) domain is constrained by a lack of structured datasets for this purpose. In this work, we introduce HierADL, a framework for automatically generating action hierarchies in the ADL domain. HierADL leverages semantic features of action labels and visual features from video clips to group fine-grained actions into coarse-grained categories. These hierarchies are integrated with the HierADL classifier, which allows simultaneous prediction of both fine-grained and coarse-grained actions to improve accuracy, besides being compatible with most video classification architectures. In addition, we conduct an ablation study to identify the most effective method for generating action hierarchies based on visual and semantic cues. We evaluate HierADL hierarchies qualitatively using t-SNE plots, and quantitatively on two ADL datasets: ETRI-Activity and Hierarchical TSU. Our results demonstrate that HierADL outperforms fine-grained-only approaches on several state-of-the-art video backbones, achieving an accuracy improvement of over 3.5%.
Manuel Benavent-Lledó, David Ortiz-Perez, David Mulero-Pérez, Pablo Ruiz-Ponce, José García Rodríguez 0001
IJCNN2
2025 Obj3Dify: Occlusion-Invariant 3D Reconstruction of Hand-Held Objects
abstract
Reconstructing high-quality 3D textured models from single images of hand-held objects is a challenging task due to occlusions, complex interactions between hands and objects, and the need for precise segmentation and texture reconstruction. In this paper, we introduce Obj3Dify, a novel framework that addresses these challenges by integrating state-of-the-art techniques in computer vision, generative modeling, and 3D reconstruction. The pipeline includes object detection and classification with LLaVA-NeXT, segmentation with Lang-SAM, occlusion inpainting with Stable Diffusion XL, and 3D model generation using TRELLIS.We evaluate our framework on the HO3D dataset, leveraging its comprehensive 3D annotations and diverse object shapes for robust benchmarking. Comparative analyses with In-Hand3D and TRELLIS demonstrate that Obj3Dify significantly improves geometric fidelity and reduces noise in 3D reconstructions, achieving results closer to ground-truth models. An ablation study further validates the contributions of each pipeline stage, highlighting the critical role of segmentation and occlusion inpainting in enhancing 3D model quality. Furthermore, a multiview approach is implemented and tested, demonstrating its usefulness for generating asymmetric and organic figures.Our results establish Obj3Dify as an effective solution for occlusion-free 3D object reconstruction, advancing 3D modeling for real-world hand-object scenarios with applications in AR, robotics, and virtual content creation.
David Mulero-Pérez, Pablo Ruiz-Ponce, Manuel Benavent-Lledó, David Ortiz-Perez, José García Rodríguez 0001
IJCNN4
2025 Effectiveness of bird species identification using Birdnet: Case study at the La Mata coastal lagoon
abstract
Monitoring avian species is fundamental to detect any negative impact from human activity. In our case study, we focus in particular on the La Mata lagoon in Torrevieja (Alicante, Spain), aiming at the monitoring of two gull species that often share the same environment: Larus Michahellis (Yellow-Legged Gull), and Ichthyaetus Audouinii (Audouin’s Gull). As of 2020, Ichtyaetus Audouinii has been included within the International Union for Conservation of Nature (IUCN) Red List, with the status of vulnerable as the global population has experienced a rapid decrease, which is expected to be approaching 40% between 2006–2030, and is projected to continue declining at a similar rate over the next three generations. Since conservation of these species requires informed, prompt and efficient decision-making, we propose constantly monitoring the ecosystem by deploying an AI-enabled IoT infrastructure to listen to bird sounds, and automatically identify bird species. To this end, we relied on the BirdNET artificial neural network for the acoustic analysis. However, using BirdNET models must be carefully planned to produce insightful data that can drive informed decisions. For this reason, we discuss our methodology and the logic behind tuning the most important parameters according to the proposed use case.
Ousman Seye, Esther Sebastián-González, David Ortiz-Perez, Carlos T. Calafate, José M. Cecilia
IE3
2025 Enhancing action recognition by leveraging the hierarchical structure of actions and textual context
abstract
We propose a novel approach to improve action recognition by exploiting the hierarchical organization of actions and by incorporating contextualized textual information, including location and previous actions, to reflect the action’s temporal context . To achieve this, we introduce a transformer architecture tailored for action recognition that employs both visual and textual features. Visual features are obtained from RGB and optical flow data, while text embeddings represent contextual information. Furthermore, we define a joint loss function to simultaneously train the model for both coarse- and fine-grained action recognition, effectively exploiting the hierarchical nature of actions. To demonstrate the effectiveness of our method, we extend the Toyota Smarthome Untrimmed (TSU) dataset by incorporating action hierarchies, resulting in the Hierarchical TSU dataset , a hierarchical dataset designed for monitoring activities of the elderly in home environments. An ablation study assesses the performance impact of different strategies for integrating contextual and hierarchical data. Experimental results demonstrate that the proposed method consistently outperforms SOTA methods on the Hierarchical TSU dataset, Assembly101 and IkeaASM, achieving over a 17% improvement in top-1 accuracy. • We propose a vision-language transformer model that improves action recognition using contextual information and action hierarchies, outperforming state-of-the-art visual-only models on the Hierarchical TSU, Assembly101 and IkeaASM action recognition benchmarks. • We introduce the Hierarchical TSU dataset for contextual action recognition with structured action hierarchies for activities of daily living. • We conduct an extensive ablation study on integrating contextual and hierarchical data to enhance action recognition performance.
Manuel Benavent-Lledó, David Mulero-Pérez, David Ortiz-Perez, José García Rodríguez 0001, Antonis A. Argyros
Comput. Vis. Image Underst.3
2025 Integrating advanced vision-language models for context recognition in risks assessment
abstract
This study proposes an open-environment, multi-label human risk classification framework, capable of identifying possible risks to which individuals appearing on input video data are exposed. The framework consists of an ensemble of models covering object detection, action recognition, context understanding and text classification tasks. Each model is evaluated separately in the context of home environments, with the overall framework performing well in each evaluation after fine tuning. The models were evaluated using a combination of several datasets, including Charades, ETRI-Activity3D, and custom video question answering and risk datasets. This study exploits the ability of large language models to interpret semantic visual features combined with textual input in order to understand the context in which the person is placed. The framework’s ability to output multiple risks and its cross-domain capabilities make it a powerful tool that can enhance current risk management systems in a variety of scenarios, such as homes, construction sites and industry.
Javier Rodríguez-Juan, David Ortiz-Perez, José García Rodríguez 0001, David Tomás 0001, Grzegorz J. Nalepa
Neurocomputing2
2025 CogniAlign: Word-level multimodal speech alignment with gated cross-attention for Alzheimer's detection
abstract
Early detection of cognitive disorders such as Alzheimer’s disease is critical for enabling timely clinical intervention and improving patient outcomes. In this work, we introduce CogniAlign, a multimodal architecture for Alzheimer’s detection that integrates audio and textual modalities, two non-intrusive sources of information that offer complementary insights into cognitive health. Unlike prior approaches that fuse modalities at a coarse level, CogniAlign leverages a word-level temporal alignment strategy that synchronizes audio embeddings with corresponding textual tokens based on transcription timestamps. This alignment supports the development of token-level fusion techniques, enabling more precise cross-modal interactions. To fully exploit this alignment, we propose a Gated Cross-Attention Fusion mechanism, where audio features attend over textual representations, guided by the superior unimodal performance of the text modality. In addition, we incorporate prosodic cues, specifically interword pauses, by inserting pause tokens into the text and generating audio embeddings for silent intervals, further enriching both streams. We evaluate CogniAlign on the ADReSSo dataset, where it achieves an accuracy of 87.35% over a Leave-One-Subject-Out setup and of 90.36% over a 5 fold Cross-Validation, outperforming existing state-of-the-art methods. A detailed ablation study confirms the advantages of our alignment strategy, attention-based fusion, and prosodic modeling. Finally, we perform a corpus analysis to assess the impact of the proposed prosodic features and apply Integrated Gradients to identify the most influential input segments used by the model in predicting cognitive health outcomes.
David Ortiz-Perez, Manuel Benavent-Lledó, Javier Rodríguez-Juan, José García Rodríguez 0001, David Tomás 0001
Knowl. Based Syst.1
2024 UnrealFall: Overcoming Data Scarcity through Generative Models
abstract
Humans perform a variety of actions, some of which are infrequent but crucial for data collection. Synthetic generation techniques are highly effective in these situations, enhancing the data for such rare actions. In response to this need, we present UnrealFall, a robust framework developed within Unreal Engine 5, designed for the generation of human action video data in hyper-realistic virtual scenes. It addresses the scarcity and limited diversity in existing datasets for actions like falls by leveraging synthetic motion generation through text-guided generative models, Gaussian Splatting technology, and MetaHumans. The usefulness of the framework is demonstrated by its capability to produce a synthetic video dataset featuring elderly individuals falling in various settings. The value of the dataset is demonstrated by its successful use in training a VideoMAE model, in conjunction with the UCF101 and various fall-specific datasets. This versatility in generating data across a spectrum of actions and environments positions our framework as a valuable tool for broader applications such as digital twin creation and dataset augmentation. The code and data are available for research at project website, darkviid.github.io/UnrealFall/.
David Mulero-Pérez, Manuel Benavent-Lledó, David Ortiz-Perez, José García Rodríguez 0001
IJCNN3
2024 Multimodal Fusion Strategies for Emotion Recognition
abstract
Emotions play a crucial role in our daily lives, influencing how we face challenges throughout the day and shaping our behavior, even when we are not consciously aware of them. Detecting emotional states in others relies on comprehending the collective impact of a variety of actions that emotions can produce, such as facial expressions, posture, tone of voice, or speech. To address this challenge, we propose a multimodal transformer-based model designed to recognize the emotional moods of individuals using video data, including audio and text transcriptions. Consequently, our model extracts the most relevant information from each modality to make a final prediction. Throughout this work, different fusion architectures have been integrated with transformer-based models to determine the optimal combination. This study examines the performance of individual modalities and their combinations using the CMU-MOSEI dataset. This dataset encompasses preprocessed video, audio, and text data. Our best model achieves a weighted accuracy of 85.59% on this dataset, surpassing previous works for this task.
David Ortiz-Perez, Manuel Benavent-Lledó, David Mulero-Pérez, David Tomás 0001, José García Rodríguez 0001
IJCNN1
2024 Cognitive Insights Across Languages: Enhancing Multimodal Interview Analysis
David Ortiz-Perez, José García Rodríguez 0001, David Tomás 0001
INTERSPEECH1
2023 A Deep Learning-Based Multimodal Architecture to predict Signs of Dementia
abstract
This paper proposes a multimodal deep learning architecture combining text and audio information to predict dementia, a disease which affects around 55 million people all over the world and makes them in some cases dependent people. The system was evaluated on the DementiaBank Pitt Corpus dataset, which includes audio recordings as well as their transcriptions for healthy people and people with dementia. Different models have been used and tested, including Convolutional Neural Networks (CNN) for audio classification, Transformers for text classification, and a combination of both in a multimodal ensemble. These models have been evaluated on a test set, obtaining the best results by using the text modality, achieving 90.36% accuracy on the task of detecting dementia. Additionally, an analysis of the corpus has been conducted for the sake of explainability, aiming to obtain more information about how the models generate their predictions and identify patterns in the data.
David Ortiz-Perez, Pablo Ruiz-Ponce, David Tomás 0001, José García Rodríguez 0001, Maria Flores Vizcaya-Moreno, Marco Leo
Neurocomputing1