EDBT 2026 Demo / reviewers in the wild / expert
José García Rodríguez 0001
dblp:44/3507
· DBLP profile ↗
88ranked-venue papers
11as first author
24since 2021 · last 2026
0000-0002-7798-3055ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 76 · 11 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?abstractAnticipating actions before they occur is a core challenge in action understanding research. While conventional methods rely on extracting and aggregating temporal information from videos, as humans we can often predict upcoming actions by observing a single moment from a scene, when given sufficient context. Can a model achieve this competence? The short answer is yes, although its effectiveness depends on the complexity of the task. In this work, we investigate to what extent video aggregation can be replaced with alternative modalities. To this end, based on recent advances in visual feature extraction and language-based reasoning, we introduce AAG, a method for Action Anticipation at a Glimpse. AAG combines RGB features with depth cues from a single frame for enhanced spatial reasoning, and incorporates prior action information to provide long-term context. This context is obtained either through textual summaries from Vision-Language Models, or from predictions generated by a single-frame action recognizer. Our results demonstrate that multimodal single-frame action anticipation using AAG can perform competitively compared to both temporally aggregated video baselines and state-of-the-art methods across three instructional activity datasets: IKEA-ASM, Meccano, and Assembly101. Manuel Benavent-Lledó, Konstantinos Bacharidis, Victoria Manousaki, Konstantinos E. Papoutsakis, Antonis A. Argyros, José García Rodríguez 0001 |
WACV | 6 |
| 2026 | Multimodal pain assessment with transformersabstractPain assessment is a critical challenge in healthcare, requiring accurate and objective measurement to enhance patient care. Traditional methods rely on subjective self-reporting, which lacks reliability, particularly for patients with communication difficulties. This work presents PainFusion+, a multimodal transformer architecture for pain assessment that integrates physiological signals and facial expressions to improve pain assessment accuracy. For physiological signals, our approach first employs convolutional neural networks to extract local patterns from short signal fragments, capturing essential pain-related features. These localized representations are then processed by a transformer encoder, which models long-range dependencies to form a comprehensive global representation. For facial video data, we leverage a frozen video transformer to extract expressive features without requiring fine-tuning, significantly reducing computational costs. Finally, both feature spaces are fused using a transformer encoder, allowing effective cross-modal learning. Experiments on publicly available datasets demonstrate that PainFusion+ outperforms existing models. In biomedical signal processing, our method achieves over a 16 % improvement in accuracy. For multimodal pain estimation, it achieves 35.40 % accuracy on the BioVid dataset, setting a new state-of-the-art benchmark. Manuel Benavent-Lledó, Maria Dolores Lopez-Valle, David Ortiz-Perez, David Mulero-Pérez, José García Rodríguez 0001, Alexandra Psarrou |
Neurocomputing | 5 |
| 2025 | MixerMDM: Learnable Composition of Human Motion Diffusion ModelsabstractGenerating human motion guided by conditions such as textual descriptions is challenging due to the need for datasets with pairs of high-quality motion and their corresponding conditions. The difficulty increases when aiming for finer control in the generation. To that end, prior works have proposed to combine several motion diffusion models pre-trained on datasets with different types of conditions, thus allowing control with multiple conditions. However, the proposed merging strategies overlook that the optimal way to combine the generation processes might depend on the particularities of each pre-trained generative model and also the specific textual descriptions. In this context, we introduce MixerMDM, the first learnable model composition technique for combining pre-trained text-conditioned human motion diffusion models. Unlike previous approaches, MixerMDM provides a dynamic mixing strategy that is trained in an adversarial fashion to learn to combine the denoising process of each model depending on the set of conditions driving the generation. By using MixerMDM to combine single- and multi-person motion diffusion models, we achieve fine-grained control on the dynamics of every person individually, and also on the overall interaction. Furthermore, we propose a new evaluation technique that, for the first time in this task, measures the interaction and individual quality by computing the alignment between the mixed generated motions and their conditions as well as the capabilities of MixerMDM to adapt the mixing throughout the denoising process depending on the motions to mix. Pablo Ruiz-Ponce, Germán Barquero, Cristina Palmero, Sergio Escalera, José García Rodríguez 0001 |
CVPR | 5 |
| 2025 | HierADL: Automatic Hierarchical Structures to Improve Daily Activity RecognitionabstractHuman action recognition is a widely explored problem in computer vision, with applications in domains such as robotics and surveillance. While most existing methods focus on predicting a single action class, human actions are not isolated events but range from simple movements to complex behaviors. Recent work has approached action recognition from a hierarchical perspective, relying on manually annotated labels available in some domains. In contrast, the Activities of Daily Living (ADL) domain is constrained by a lack of structured datasets for this purpose. In this work, we introduce HierADL, a framework for automatically generating action hierarchies in the ADL domain. HierADL leverages semantic features of action labels and visual features from video clips to group fine-grained actions into coarse-grained categories. These hierarchies are integrated with the HierADL classifier, which allows simultaneous prediction of both fine-grained and coarse-grained actions to improve accuracy, besides being compatible with most video classification architectures. In addition, we conduct an ablation study to identify the most effective method for generating action hierarchies based on visual and semantic cues. We evaluate HierADL hierarchies qualitatively using t-SNE plots, and quantitatively on two ADL datasets: ETRI-Activity and Hierarchical TSU. Our results demonstrate that HierADL outperforms fine-grained-only approaches on several state-of-the-art video backbones, achieving an accuracy improvement of over 3.5%. Manuel Benavent-Lledó, David Ortiz-Perez, David Mulero-Pérez, Pablo Ruiz-Ponce, José García Rodríguez 0001 |
IJCNN | 5 |
| 2025 | Obj3Dify: Occlusion-Invariant 3D Reconstruction of Hand-Held ObjectsabstractReconstructing high-quality 3D textured models from single images of hand-held objects is a challenging task due to occlusions, complex interactions between hands and objects, and the need for precise segmentation and texture reconstruction. In this paper, we introduce Obj3Dify, a novel framework that addresses these challenges by integrating state-of-the-art techniques in computer vision, generative modeling, and 3D reconstruction. The pipeline includes object detection and classification with LLaVA-NeXT, segmentation with Lang-SAM, occlusion inpainting with Stable Diffusion XL, and 3D model generation using TRELLIS.We evaluate our framework on the HO3D dataset, leveraging its comprehensive 3D annotations and diverse object shapes for robust benchmarking. Comparative analyses with In-Hand3D and TRELLIS demonstrate that Obj3Dify significantly improves geometric fidelity and reduces noise in 3D reconstructions, achieving results closer to ground-truth models. An ablation study further validates the contributions of each pipeline stage, highlighting the critical role of segmentation and occlusion inpainting in enhancing 3D model quality. Furthermore, a multiview approach is implemented and tested, demonstrating its usefulness for generating asymmetric and organic figures.Our results establish Obj3Dify as an effective solution for occlusion-free 3D object reconstruction, advancing 3D modeling for real-world hand-object scenarios with applications in AR, robotics, and virtual content creation. David Mulero-Pérez, Pablo Ruiz-Ponce, Manuel Benavent-Lledó, David Ortiz-Perez, José García Rodríguez 0001 |
IJCNN | 5 |
| 2025 | ViT-VS: On the Applicability of Pretrained Vision Transformer Features for Generalizable Visual ServoingabstractVisual servoing enables robots to precisely position their end-effector relative to a target object. While classical methods rely on hand-crafted features and thus are universally applicable without task-specific training, they often struggle with occlusions and environmental variations, whereas learning-based approaches improve robustness but typically require extensive training. We present a visual servoing approach that leverages pretrained vision transformers for semantic feature extraction, combining the advantages of both paradigms while also being able to generalize beyond the provided sample. Our approach achieves full convergence in unperturbed scenarios and surpasses classical image-based visual servoing by up to 31.2% relative improvement in perturbed scenarios. Even the convergence rates of learning-based methods are matched despite requiring no task-or object-specific training. Real-world evaluations confirm robust performance in end-effector positioning, industrial box manipulation, and grasping of unseen objects using only a reference from the same category. Our code and simulation environment are available at: https://alessandroscherl.github.io/ViT-VS/ Alessandro Scherl, Stefan Thalhammer, Bernhard Neuberger, Wilfried Wöber, José García Rodríguez 0001 |
IROS | 5 |
| 2025 | Enhancing action recognition by leveraging the hierarchical structure of actions and textual contextabstractWe propose a novel approach to improve action recognition by exploiting the hierarchical organization of actions and by incorporating contextualized textual information, including location and previous actions, to reflect the action’s temporal context . To achieve this, we introduce a transformer architecture tailored for action recognition that employs both visual and textual features. Visual features are obtained from RGB and optical flow data, while text embeddings represent contextual information. Furthermore, we define a joint loss function to simultaneously train the model for both coarse- and fine-grained action recognition, effectively exploiting the hierarchical nature of actions. To demonstrate the effectiveness of our method, we extend the Toyota Smarthome Untrimmed (TSU) dataset by incorporating action hierarchies, resulting in the Hierarchical TSU dataset , a hierarchical dataset designed for monitoring activities of the elderly in home environments. An ablation study assesses the performance impact of different strategies for integrating contextual and hierarchical data. Experimental results demonstrate that the proposed method consistently outperforms SOTA methods on the Hierarchical TSU dataset, Assembly101 and IkeaASM, achieving over a 17% improvement in top-1 accuracy. • We propose a vision-language transformer model that improves action recognition using contextual information and action hierarchies, outperforming state-of-the-art visual-only models on the Hierarchical TSU, Assembly101 and IkeaASM action recognition benchmarks. • We introduce the Hierarchical TSU dataset for contextual action recognition with structured action hierarchies for activities of daily living. • We conduct an extensive ablation study on integrating contextual and hierarchical data to enhance action recognition performance. Manuel Benavent-Lledó, David Mulero-Pérez, David Ortiz-Perez, José García Rodríguez 0001, Antonis A. Argyros |
Comput. Vis. Image Underst. | 4 |
| 2025 | Integrating advanced vision-language models for context recognition in risks assessmentabstractThis study proposes an open-environment, multi-label human risk classification framework, capable of identifying possible risks to which individuals appearing on input video data are exposed. The framework consists of an ensemble of models covering object detection, action recognition, context understanding and text classification tasks. Each model is evaluated separately in the context of home environments, with the overall framework performing well in each evaluation after fine tuning. The models were evaluated using a combination of several datasets, including Charades, ETRI-Activity3D, and custom video question answering and risk datasets. This study exploits the ability of large language models to interpret semantic visual features combined with textual input in order to understand the context in which the person is placed. The framework’s ability to output multiple risks and its cross-domain capabilities make it a powerful tool that can enhance current risk management systems in a variety of scenarios, such as homes, construction sites and industry. Javier Rodríguez-Juan, David Ortiz-Perez, José García Rodríguez 0001, David Tomás 0001, Grzegorz J. Nalepa |
Neurocomputing | 3 |
| 2025 | CogniAlign: Word-level multimodal speech alignment with gated cross-attention for Alzheimer's detectionabstractEarly detection of cognitive disorders such as Alzheimer’s disease is critical for enabling timely clinical intervention and improving patient outcomes. In this work, we introduce CogniAlign, a multimodal architecture for Alzheimer’s detection that integrates audio and textual modalities, two non-intrusive sources of information that offer complementary insights into cognitive health. Unlike prior approaches that fuse modalities at a coarse level, CogniAlign leverages a word-level temporal alignment strategy that synchronizes audio embeddings with corresponding textual tokens based on transcription timestamps. This alignment supports the development of token-level fusion techniques, enabling more precise cross-modal interactions. To fully exploit this alignment, we propose a Gated Cross-Attention Fusion mechanism, where audio features attend over textual representations, guided by the superior unimodal performance of the text modality. In addition, we incorporate prosodic cues, specifically interword pauses, by inserting pause tokens into the text and generating audio embeddings for silent intervals, further enriching both streams. We evaluate CogniAlign on the ADReSSo dataset, where it achieves an accuracy of 87.35% over a Leave-One-Subject-Out setup and of 90.36% over a 5 fold Cross-Validation, outperforming existing state-of-the-art methods. A detailed ablation study confirms the advantages of our alignment strategy, attention-based fusion, and prosodic modeling. Finally, we perform a corpus analysis to assess the impact of the proposed prosodic features and apply Integrated Gradients to identify the most influential input segments used by the model in predicting cognitive health outcomes. David Ortiz-Perez, Manuel Benavent-Lledó, Javier Rodríguez-Juan, José García Rodríguez 0001, David Tomás 0001 |
Knowl. Based Syst. | 4 |
| 2025 | Holo4Care: a MR framework for assisting in activities of daily living by context-aware action recognitionabstractAbstract The evolution of virtual and augmented reality devices in recent years has encouraged researchers to develop new systems for different fields. This paper introduces Holo4Care, a context-aware mixed reality framework designed for assisting in activities of daily living (ADL) using the HoloLens 2. By leveraging egocentric cameras embedded in these devices, which offer a close-to-wearer perspective, our framework establishes a congruent relationship, facilitating a deeper understanding of user actions and enabling effective assistance. In our approach, we extend a previously established action estimation architecture after conducting a thorough review of state-of-the-art methods. The proposed architecture utilizes YOLO for hand and object detection, enabling action estimation based on these identified elements. We have trained new models on well-known datasets for object detection, incorporating action recognition annotations. The achieved mean Average Precision (mAP) is 33.2% in the EpicKitchens dataset and 26.4% on the ADL dataset. Leveraging the capabilities of the HoloLens 2, including spatial mapping and 3D hologram display, our system seamlessly presents the output of the action recognition architecture to the user. Unlike previous systems that focus primarily on user evaluation, Holo4Care emphasizes assistance by providing a set of global actions based on the user’s field of view and hand positions that reflect their intentions. Experimental results demonstrate Holo4Care’s ability to assist users in activities of daily living and other domains. Manuel Benavent-Lledó, David Mulero-Pérez, José García Rodríguez 0001, Ester Martínez-Martín, Maria Flores Vizcaya-Moreno |
Multim. Tools Appl. | 3 |
| 2024 | UnrealFall: Overcoming Data Scarcity through Generative ModelsabstractHumans perform a variety of actions, some of which are infrequent but crucial for data collection. Synthetic generation techniques are highly effective in these situations, enhancing the data for such rare actions. In response to this need, we present UnrealFall, a robust framework developed within Unreal Engine 5, designed for the generation of human action video data in hyper-realistic virtual scenes. It addresses the scarcity and limited diversity in existing datasets for actions like falls by leveraging synthetic motion generation through text-guided generative models, Gaussian Splatting technology, and MetaHumans. The usefulness of the framework is demonstrated by its capability to produce a synthetic video dataset featuring elderly individuals falling in various settings. The value of the dataset is demonstrated by its successful use in training a VideoMAE model, in conjunction with the UCF101 and various fall-specific datasets. This versatility in generating data across a spectrum of actions and environments positions our framework as a valuable tool for broader applications such as digital twin creation and dataset augmentation. The code and data are available for research at project website, darkviid.github.io/UnrealFall/. David Mulero-Pérez, Manuel Benavent-Lledó, David Ortiz-Perez, José García Rodríguez 0001 |
IJCNN | 4 |
| 2024 | Multimodal Fusion Strategies for Emotion RecognitionabstractEmotions play a crucial role in our daily lives, influencing how we face challenges throughout the day and shaping our behavior, even when we are not consciously aware of them. Detecting emotional states in others relies on comprehending the collective impact of a variety of actions that emotions can produce, such as facial expressions, posture, tone of voice, or speech. To address this challenge, we propose a multimodal transformer-based model designed to recognize the emotional moods of individuals using video data, including audio and text transcriptions. Consequently, our model extracts the most relevant information from each modality to make a final prediction. Throughout this work, different fusion architectures have been integrated with transformer-based models to determine the optimal combination. This study examines the performance of individual modalities and their combinations using the CMU-MOSEI dataset. This dataset encompasses preprocessed video, audio, and text data. Our best model achieves a weighted accuracy of 85.59% on this dataset, surpassing previous works for this task. David Ortiz-Perez, Manuel Benavent-Lledó, David Mulero-Pérez, David Tomás 0001, José García Rodríguez 0001 |
IJCNN | 5 |
| 2024 | Cognitive Insights Across Languages: Enhancing Multimodal Interview Analysis
David Ortiz-Perez, José García Rodríguez 0001, David Tomás 0001 |
INTERSPEECH | 2 |
| 2024 | Improving Landslides Prediction: Meteorological Data Preprocessing Based on Supervised and Unsupervised LearningabstractThe hazard of landslides has been demonstrated over time with numerous events causing damage to human lives and high material costs. Several previous studies have shown that one of the predominant factors in landslides is intensive rainfall. The present work proposes the use of data generated by weather stations to predict landslides. We give special treatment to precipitation information as the most influential factor and whose data are accumulated in time windows (3, 5, 7, 10, 15, 20, and 30 days) looking for the persistence of meteorological conditions. To optimize the dataset composed of geological, geomorphological, and climatological data, a feature selection process is applied to the meteorological variables. We use filter-based feature ranking and Self-Organizing Map (SOM) with Clustering as supervised and unsupervised machine learning techniques, respectively. This contribution was successfully verified by experimenting with different classification models, improving the test accuracy of the prediction, and obtaining 99.29% for Multilayer Perceptron, 96.80% for Random Forest, and 88.79% for Support Vector Machine. To validate the proposal, a geographical area sensitive to this phenomenon was selected, which is monitored by several meteorological stations. Practical use is a valuable tool for risk management decision making, can help save lives and reduce economic losses. Byron Guerrero Rodríguez, Jaime Salvador-Meneses, José García Rodríguez 0001, Christian Mejía Escobar |
Cybern. Syst. | 3 |
| 2024 | Challenges for Monocular 6-D Object Pose Estimation in RoboticsabstractObject pose estimation is a core perception task that enables, for example, object manipulation and scene understanding. The widely available, inexpensive, and high-resolution RGB sensors and CNNs that allow for fast inference make monocular approaches especially well-suited for robotics applications. We observe that previous surveys establish the state of the art for varying modalities, single- and multiview settings, and datasets and metrics that consider a multitude of applications. We argue, however, that those works' broad scope hinders the identification of open challenges that are specific to monocular approaches and the derivation of promising future challenges for their application in robotics. By providing a unified view on recent publications from both robotics and computer vision, we find that occlusion handling, pose representations, and formalizing and improving category-level pose estimation are still fundamental challenges that are highly relevant for robotics. Moreover, to further improve robotic performance, large object sets, novel objects, refractive materials, and uncertainty estimates are central and largely unsolved open challenges. In order to address them, ontological reasoning, deformability handling, scene-level reasoning, realistic datasets, and the ecological footprint of algorithms need to be improved. Stefan Thalhammer, Dominik Bauer, Peter Hönig, Jean-Baptiste Weibel, José García Rodríguez 0001, Markus Vincze |
IEEE Trans. Robotics | 5 |
| 2023 | A Deep Learning-Based Multimodal Architecture to predict Signs of DementiaabstractThis paper proposes a multimodal deep learning architecture combining text and audio information to predict dementia, a disease which affects around 55 million people all over the world and makes them in some cases dependent people. The system was evaluated on the DementiaBank Pitt Corpus dataset, which includes audio recordings as well as their transcriptions for healthy people and people with dementia. Different models have been used and tested, including Convolutional Neural Networks (CNN) for audio classification, Transformers for text classification, and a combination of both in a multimodal ensemble. These models have been evaluated on a test set, obtaining the best results by using the text modality, achieving 90.36% accuracy on the task of detecting dementia. Additionally, an analysis of the corpus has been conducted for the sake of explainability, aiming to obtain more information about how the models generate their predictions and identify patterns in the data. David Ortiz-Perez, Pablo Ruiz-Ponce, David Tomás 0001, José García Rodríguez 0001, Maria Flores Vizcaya-Moreno, Marco Leo |
Neurocomputing | 4 |
| 2023 | Self-supervised Vision Transformers for 3D pose estimation of novel objectsabstractObject pose estimation is important for object manipulation and scene understanding. In order to improve the general applicability of pose estimators, recent research focuses on providing estimates for novel objects, that is, objects unseen during training. Such works use deep template matching strategies to retrieve the closest template connected to a query image, which implicitly provides object class and pose. Despite the recent success and improvements of Vision Transformers over CNNs for many vision tasks, the state of the art uses CNN-based approaches for novel object pose estimation. This work evaluates and demonstrates the differences between self-supervised CNNs and Vision Transformers for deep template matching. In detail, both types of approaches are trained using contrastive learning to match training images against rendered templates of isolated objects. At test time such templates are matched against query images of known and novel objects under challenging settings, such as clutter, occlusion and object symmetries, using masked cosine similarity. The presented results not only demonstrate that Vision Transformers improve matching accuracy over CNNs but also that for some cases pre-trained Vision Transformers do not need fine-tuning to achieve the improvement. Furthermore, we highlight the differences in optimization and network architecture when comparing these two types of networks for deep template matching. Stefan Thalhammer, Jean-Baptiste Weibel, Markus Vincze, José García Rodríguez 0001 |
Image Vis. Comput. | 4 |
| 2023 | Detecting and locating trending places using multimodal social network dataabstractAbstract This paper presents a machine learning-based classifier for detecting points of interest through the combined use of images and text from social networks. This model exploits the transfer learning capabilities of the neural network architecture CLIP (Contrastive Language-Image Pre-Training) in multimodal environments using image and text. Different methodologies based on multimodal information are explored for the geolocation of the places detected. To this end, pre-trained neural network models are used for the classification of images and their associated texts. The result is a system that allows creating new synergies between images and texts in order to detect and geolocate trending places that has not been previously tagged by any other means, providing potentially relevant information for tasks such as cataloging specific types of places in a city for the tourism industry. The experiments carried out reveal that, in general, textual information is more accurate and relevant than visual cues in this multimodal setting. Luis Lucas, David Tomás 0001, José García Rodríguez 0001 |
Multim. Tools Appl. | 3 |
| 2022 | Predicting Human-Object Interactions in Egocentric VideosabstractEgocentric videos provide a rich source of hand-object interactions that support action recognition. However, prior to action recognition, one may need to detect the presence of hands and objects in the scene. In this work, we propose an action estimation architecture based on the simultaneous detection of the hands and objects in the scene. For the hand and object detection, we have adapted well known YOLO architecture, leveraging its inference speed and accuracy. We experimentally determined the best performing architecture for our task. After obtaining the hand and object bounding boxes, we select the most likely objects to interact with, i.e., the closest objects to a hand. The rough estimation of the closest objects to a hand is a direct approach to determine hand-object interaction. After identifying the scene and alongside a set of per-object and global actions, we could determine the most suitable action we are performing in each context. Manuel Benavent-Lledó, Sergiu Ovidiu-Oprea, John Alejandro Castro-Vargas, David Mulero-Pérez, José García Rodríguez 0001 |
IJCNN | 5 |
| 2022 | Synthetic contact maps to predict grasp regions on objectsabstractFrom early ages, we learn to naturally grasp objects from our trial-and-error interaction with the environment. We end up excelling at manipulating objects, such that grabbing a pan by its handle is all but unconscious. However, performing the same task from a machine's perspective is particularly challenging, as it requires an in-depth understanding of object manipulation. We could approximate such an understanding by capturing the hand-object contact resulting from human grasps, i.e., contact maps. This has already been done using cumbersome capture systems, and not on a large scale. In this work, we simplify and accelerate such a process by generating contact maps from our interaction with household objects in photorealistic virtual reality environments. To the best of our knowledge, we are the first to generate contact maps at scale from our interaction in a virtual scenario. We train an image-to-image translation method to predict grasp regions on objects, demonstrating the usefulness of our generated contact maps. We provide all the necessary tools, code, and dataset to foster applications where hand-object understanding is necessary. Pablo Martinez-Gonzalez, David Mulero-Pérez, Sergiu Ovidiu-Oprea, Manuel Benavent-Lledó, Sergio Orts, José García Rodríguez 0001 |
IJCNN | 6 |
| 2022 | A Review on Deep Learning Techniques for Video PredictionabstractThe ability to predict, anticipate and reason about future outcomes is a key component of intelligent decision-making systems. In light of the success of deep learning in computer vision, deep-learning-based video prediction emerged as a promising research direction. Defined as a self-supervised learning task, video prediction represents a suitable framework for representation learning, as it demonstrated potential capabilities for extracting meaningful representations of the underlying patterns in natural videos. Motivated by the increasing interest in this task, we provide a review on the deep learning methods for prediction in video sequences. We first define the video prediction fundamentals, as well as mandatory background concepts and the most used datasets. Next, we carefully analyze existing video prediction models organized according to a proposed taxonomy, highlighting their contributions and their significance in the field. The summary of the datasets and methods is accompanied with experimental results that facilitate the assessment of the state of the art on a quantitative basis. The paper is summarized by drawing some general conclusions, identifying open research challenges and by pointing out future research directions. Sergiu Ovidiu-Oprea, Pablo Martinez-Gonzalez, Alberto Garcia-Garcia, John Alejandro Castro-Vargas, Sergio Orts, José García Rodríguez 0001, Antonis A. Argyros |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Graph Convolutional Neural Networks-based 3D Hand Pose Estimation over Point CloudsabstractIn recent years we can find a multitude of approaches that aim to return the 3D pose of the hands. Most of them try to estimate the pose from RGB images or even include some geometrical information via depth maps. Furthermore, some proposals have shown promising results using point clouds as input data. However, the sparse nature of this type of data is often one of its drawbacks. To tackle this sparsity, different strategies have been brought to the table such as voxelizing or sorting the input data to impose a structure to the input domain. In this paper, we address this problem by means of a graph structure. This process implies that we should accommodate the point cloud onto a graph representation that connects its points. We connect each point to its neighborhood, a method that has been successfully used in similar proposals and whose clustering effect enables us to emulate an effect similar to kernels in image convolutions. The proposed architecture uses both graph and 2D convolutions. The first one aims to extract local features and build a feature map, from which the 2D convolutions will extract a second level of features used to estimate the pose. This proposal shows initial results to return a 3D pose of the hand from depth maps, which are projected on point clouds and redefined as graphs. Although the results diverge from other more established methods in the state of the art, it presents a proof of concept by which to address this problem without losing spatial information. John Alejandro Castro-Vargas, Pablo Martinez-Gonzalez, Sergiu Ovidiu-Oprea, Alberto Garcia-Garcia, Sergio Orts, José García Rodríguez 0001 |
IJCNN | 6 |
| 2021 | UnrealROX+: An Improved Tool for Acquiring Synthetic Data from Virtual 3D EnvironmentsabstractSynthetic data generation has become essential in last years for feeding data-driven algorithms, which surpassed traditional techniques performance in almost every computer vision problem. Gathering and labelling the amount of data needed for these data-hungry models in the real world may become unfeasible and error-prone, while synthetic data give us the possibility of generating huge amounts of data with pixel-perfect annotations. However, most synthetic datasets lack from enough realism in their rendered images. In that context UnrealROX generation tool was presented in 2019, allowing to generate highly realistic data, at high resolutions and framerates, with an efficient pipeline based on Unreal Engine, a cutting-edge videogame engine. UnrealROX enabled robotic vision researchers to generate realistic and visually plausible data with full ground truth for a wide variety of problems such as class and instance semantic segmentation, object detection, depth estimation, visual grasping, and navigation. Nevertheless, its workflow was very tied to generate image sequences from a robotic on-board camera, making hard to generate data for other purposes. In this work, we present UnrealROX+, an improved version of UnrealROX where its decoupled and easy-to-use data acquisition system allows to quickly design and generate data in a much more flexible and customizable way. Moreover, it is packaged as an Unreal plug-in, which makes it more comfortable to use with already existing Unreal projects, and it also includes new features such as generating albedo or a Python API for interacting with the virtual environment from Deep Learning frameworks. Pablo Martinez-Gonzalez, Sergiu Ovidiu-Oprea, John Alejandro Castro-Vargas, Alberto Garcia-Garcia, Sergio Orts, José García Rodríguez 0001, Markus Vincze |
IJCNN | 6 |
| 2021 | H-GAN: the power of GANs in your HandsabstractWe present HandGAN (H-GAN), a cycle-consistent adversarial learning approach implementing multi-scale perceptual discriminators. It is designed to translate synthetic images of hands to the real domain. Synthetic hands provide complete ground-truth annotations, yet they are not representative of the target distribution of real-world data. We strive to provide the perfect blend of a realistic hand appearance with synthetic annotations. Relying on image-to-image translation, we improve the appearance of synthetic hands to approximate the statistical distribution underlying a collection of real images of hands. H-GAN tackles not only the cross-domain tone mapping but also structural differences in localized areas such as shading discontinuities. Results are evaluated on a qualitative and quantitative basis improving previous works. Furthermore, we relied on the hand classification task to claim our generated hands are statistically similar to the real domain of hands. Sergiu Ovidiu-Oprea, Giorgos Karvounas, Pablo Martinez-Gonzalez, Nikolaos Kyriazis, Sergio Orts, Iasonas Oikonomidis, Alberto Garcia-Garcia, Aggeliki Tsoli, José García Rodríguez 0001, Antonis A. Argyros |
IJCNN | 9 |
| 2020 | COMBAHO: A deep learning system for integrating brain injury patients in societyabstractIn the last years, the care of dependent people, either by disease, accident, disability, or age, is one of the current priority research topics in developed countries. Moreover, such care is intended to be at patients home, in order to minimize the cost of therapies. Patients rehabilitation will be fulfilled when their integration in society is achieved, either in the family or in a work environment. To address this challenge, we propose the development and evaluation of an assistant for people with acquired brain injury or dependents. This assistant is twofold: in the patient’s home is based on the design and use of an intelligent environment with abilities to monitor and active learning, combined with an autonomous social robot for interactive assistance and stimulation. On the other hand, it is complemented with an outdoor assistant, to help patients under disorientation or complex situations. This involves the integration of several existing technologies and provides solutions to a variety of technological challenges. Deep leaning-based techniques are proposed as core technology to solve these problems. José García Rodríguez 0001, Francisco Gomez-Donoso, Sergiu Ovidiu-Oprea, Alberto Garcia-Garcia, Miguel Cazorla, Sergio Orts, Zuria Bauer, John Alejandro Castro-Vargas, Félix Escalona, David Ivorra-Piqueres, Pablo Martinez-Gonzalez, Eugenio Aguirre, Miguel García-Silvente, Marcelo García-Pérez, José María Cañas, Francisco Martín 0001, Jonatan Gines Clavero, Francisco Rivas-Montero |
Pattern Recognit. Lett. | 1 |
| 2019 | 3DCNN Performance in Hand Gesture Recognition Applied to Robot Arm InteractionabstractIn the past, methods for hand sign recognition have been successfully tested in Human Robot Interaction (HRI) using traditional methodologies based on static image features and machine learning. However, the recognition of gestures in video sequences is a problem still open, because current detection methods achieve low scores when the background is undefined or in unstructured scenarios. Deep learning techniques are being applied to approach a solution for this problem in recent years. In this paper, we present a study in which we analyse the performance of a 3DCNN architecture for hand gesture recognition in an unstructured scenario. The system yields a score of 73% in both accuracy and F1. The aim of the work is the implementation of a system for commanding robots with gestures recorded by video in real scenarios. John Alejandro Castro-Vargas, Brayan S. Zapata-Impata, Pablo Gil, José García Rodríguez 0001, Fernando Torres 0001 |
ICPRAM | 4 |
| 2019 | TactileGCN: A Graph Convolutional Network for Predicting Grasp Stability with Tactile SensorsabstractTactile sensors provide useful contact data during the interaction with an object which can be used to accurately learn to determine the stability of a grasp. Most of the works in the literature represented tactile readings as plain feature vectors or matrix-like tactile images, using them to train machine learning models. In this work, we explore an alternative way of exploiting tactile information to predict grasp stability by leveraging graph-like representations of tactile data, which preserve the actual spatial arrangement of the sensor's taxels and their locality. In experimentation, we trained a Graph Neural Network to binary classify grasps as stable or slippery ones. To train such network and prove its predictive capabilities for the problem at hand, we captured a novel dataset of ~ 5000 three-fingered grasps across 41 objects for training and 1000 grasps with 10 unknown objects for testing. Our experiments prove that this novel approach can be effectively used to predict grasp stability. Alberto Garcia-Garcia, Brayan S. Zapata-Impata, Sergio Orts, Pablo Gil, José García Rodríguez 0001 |
IJCNN | 5 |
| 2019 | A visually realistic grasping system for object manipulation and interaction in virtual reality environments
Sergiu Ovidiu-Oprea, Pablo Martinez-Gonzalez, Alberto Garcia-Garcia, John Alejandro Castro-Vargas, Sergio Orts, José García Rodríguez 0001 |
Comput. Graph. | 6 |
| 2019 | Evaluation of different chrominance models in the detection and reconstruction of faces and hands using the growing neural gas networkabstractPhysical traits such as the shape of the hand and face can be used for human recognition and identification in video surveillance systems and in biometric authentication smart card systems, as well as in personal health care. However, the accuracy of such systems suffers from illumination changes, unpredictability, and variability in appearance (e.g. occluded faces or hands, cluttered backgrounds, etc.). This work evaluates different statistical and chrominance models in different environments with increasingly cluttered backgrounds where changes in lighting are common and with no occlusions applied, in order to get a reliable neural network reconstruction of faces and hands, without taking into account the structural and temporal kinematics of the hands. First a statistical model is used for skin colour segmentation to roughly locate hands and faces. Then a neural network is used to reconstruct in 3D the hands and faces. For the filtering and the reconstruction we have used the growing neural gas algorithm which can preserve the topology of an object without restarting the learning process. Experiments conducted on our own database but also on four benchmark databases (Stirling's, Alicante, Essex, and Stegmann's) and on deaf individuals from normal 2D videos are freely available on the BSL signbank dataset. Results demonstrate the validity of our system to solve problems of face and hand segmentation and reconstruction under different environmental conditions. Anastassia Angelopoulou, José García Rodríguez 0001, Sergio Orts, Epaminondas Kapetanios, Xing Liang, Bencie Woll, Alexandra Psarrou |
Pattern Anal. Appl. | 2 |
| 2018 | Machine Learning Methods Based Preprocessing to Improve Categorical Data Classification
Zoila Ruiz, Jaime Salvador-Meneses, José García Rodríguez 0001 |
IDEAL (1) | 3 |
| 2018 | Data Pre-processing to Apply Multiple Imputation Techniques: A Case Study on Real-World Census Data
Zoila Ruiz, Jaime Salvador-Meneses, José García Rodríguez 0001, Antonio J. Tallón-Ballesteros |
IDEAL (2) | 3 |
| 2018 | Categorical Big Data Processing
Jaime Salvador-Meneses, Zoila Ruiz, José García Rodríguez 0001 |
IDEAL (1) | 3 |
| 2018 | Finding the Place: How to Train and Use Convolutional Neural Networks for a Dynamically Learning RobotabstractFor a robot, the ability to adapt his knowledge automatically and customize its behavior is a key feature. Furthermore, a robot should be able to carry out its tasks at a long-term basis, performing it seamlessly in presence of changes in their surroundings. To do that, it is essential that the robot dynamically learn from their environment, but to perform a fully retraining of a deep learning architecture when the model needs new knowledge is a highly time consuming task. This work focus on exploring several strategies to include new data to an already learned model, applied to the semantic localization problem focusing in the accuracy of the final model and their training time. Exhaustive experimentation is carried out and each result is discussed consequently. Edmanuel Cruz, José Carlos Rangel, Francisco Gomez-Donoso, Zuria Bauer, Miguel Cazorla, José García Rodríguez 0001 |
IJCNN | 6 |
| 2018 | A New Dataset and Performance Evaluation of a Region-based CNN for Urban Object DetectionabstractIn the last years, we have seen a large growth in the number of applications which use deep learning-based object detectors. Autonomous Driving Assistance Systems (ADAS) is one of the areas where it has more impact. In this work, we present a novel study that evaluates a state-of-the-art technique for urban object localization. In particular, we investigate the performance of the Faster R-CNN method to detect and localize urban objects in a variety of outdoor urban videos involving pedestrians, cars, bicycles and other objects moving in the scene. We propose a new dataset that is used for benchmarking the accuracy of a real-time object detector (Faster R-CNN). Part of the data was collected using an HD camera mounted in a vehicle. Besides, some of the data is weakly annotated so it can be used for testing weakly-supervised learning techniques. We have carried out extensive experiments demonstrating the effectiveness of the baseline approach, which achieved a 74.2% accuracy on the proposed dataset. Moreover, we have evaluated a baseline approach for traffic sign recognition achieving an accuracy of 98.1%. A ResNet-based architecture was trained and used for this purpose as a second stage of our object detector. The full dataset is available for download at http://www.rovit.ua.es/dataset/traffic/. Alex Dominguez-Sanchez, Sergio Orts, José García Rodríguez 0001, Miguel Cazorla |
IJCNN | 3 |
| 2018 | The RobotriX: An Extremely Photorealistic and Very-Large-Scale Indoor Dataset of Sequences with Robot Trajectories and InteractionsabstractEnter the RobotriX, an extremely photorealistic indoor dataset designed to enable the application of deep learning techniques to a wide variety of robotic vision problems. The RobotriX consists of hyperrealistic indoor scenes which are explored by robot agents which also interact with objects in a visually realistic manner in that simulated world. Photorealistic scenes and robots are rendered by Unreal Engine into a virtual reality headset which captures gaze so that a human operator can move the robot and use controllers for the robotic hands; scene information is dumped on a per-frame basis so that it can be reproduced offline using UnrealCV to generate raw data and ground truth labels. By taking this approach, we were able to generate a dataset of 38 semantic classes across 512 sequences totaling 8M stills recorded at +60 frames per second with full HD resolution. For each frame, RGB-D and 3D information is provided with full annotations in both spaces. Thanks to the high quality and quantity of both raw information and annotations, the RobotriX will serve as a new milestone for investigating 2D and 3D robotic vision tasks with large-scale data-driven techniques. Alberto Garcia-Garcia, Pablo Martinez-Gonzalez, Sergiu Ovidiu-Oprea, John Alejandro Castro-Vargas, Sergio Orts, José García Rodríguez 0001, Alvaro Jover-Alvarez |
IROS | 6 |
| 2018 | Expert systems: Special issue on "Machine Learning Methods Neural Networks applied to Vision and Robotics (MLMVR)"abstractThe International Joint Conference on Neural Networks (IJCNN) was held in Anchorage (Alaska) in May 2017. This top conference in the field of neural networks included many tracks and special sessions. In particular, a special session on Machine Learning Methods Neural Networks applied to Vision and Robotics (MLMVR) was organized by the authors receiving a large volume of excellent contributions. Only a small set of outstanding papers presented at this special session were invited to submit extended versions of their work. After a rigorous revision process, four of these papers were accepted. José García Rodríguez 0001, Sergio Escalera, Alexandra Psarrou, Isabelle Guyon, Andrew Lewis 0004, Jürgen Leitner |
Expert Syst. J. Knowl. Eng. | 1 |
| 2018 | Fast 2D/3D object representation with growing neural gasabstractThis work presents the design of a real-time system to model visual objects with the use of self-organising networks. The architecture of the system addresses multiple computer vision tasks such as image segmentation, optimal parameter estimation and object representation. We first develop a framework for building non-rigid shapes using the growth mechanism of the self-organising maps, and then we define an optimal number of nodes without overfitting or underfitting the network based on the knowledge obtained from information-theoretic considerations. We present experimental results for hands and faces, and we quantitatively evaluate the matching capabilities of the proposed method with the topographic product. The proposed method is easily extensible to 3D objects, as it offers similar features for efficient mesh reconstruction. Anastassia Angelopoulou, José García Rodríguez 0001, Sergio Orts, Gaurav Gupta 0001, Alexandra Psarrou |
Neural Comput. Appl. | 2 |
| 2018 | Bioinspired point cloud representation: 3D object tracking
Sergio Orts, José García Rodríguez 0001, Miguel Cazorla, Vicente Morell, Jorge Azorín López, Marcelo Saval-Calvo, Alberto Garcia-Garcia, Victor Villena-Martinez |
Neural Comput. Appl. | 2 |
| 2017 | LonchaNet: A sliced-based CNN architecture for real-time 3D object recognitionabstractIn the last few years, Convolutional Neural Networks (CNNs) had become the default paradigm to address classification problems, specially, but not only, in image recognition. This is mainly due to the high success rate that they provide. Despite there currently exist approaches that apply deep learning to the 3D recognition problem, they are either too slow for online uses or too error prone. To fill this gap, we propose LonchaNet, a deep learning architecture for point clouds classification. Our system successfully achieves a high accuracy yet providing a low computation cost. A dense set of experiments were carried out in order to validate our system in the frame of the ModelNet - a large-scale 3D CAD models dataset - challenge. Our proposal achieves a success rate of 94.37% in the ModelNet-10 classification task, the second place in the leaderboard as of today (November, 2016). Francisco Gomez-Donoso, Alberto Garcia-Garcia, José García Rodríguez 0001, Sergio Orts, Miguel Cazorla |
IJCNN | 3 |
| 2017 | A recurrent neural network based Schaeffer gesture recognition systemabstractSchaeffer language is considered an effective method to help autistic children overcome communicative disorders. Speech and language therapy results in an improvement in communication skills and understanding of language productions. In this work, a Schaeffer language recognition system is presented with the purpose of teaching children with autism disorder the correct way to communicate using gestures in combination with speech reproduction. The purpose is to accelerate the learning process and increase children interest using a technological approach. A Long Short-Term Memory (LSTM) model has been implemented for this purpose reporting a 93.13% classification success rate over a subset of 25 Schaeffer gestures. A comparison with vanilla RNNs and GRU-based models has been also carried out. Pose-based features such as angles and euclidean distances have been extracted from our gesture dataset by processing raw skeletal data from a Kinect v2 sensor. Sergiu Ovidiu-Oprea, Alberto Garcia-Garcia, José García Rodríguez 0001, Sergio Orts, Miguel Cazorla |
IJCNN | 3 |
| 2017 | A study of the effect of noise and occlusion on the accuracy of convolutional neural networks applied to 3D object recognition
Alberto Garcia-Garcia, José García Rodríguez 0001, Sergio Orts, Sergiu Ovidiu-Oprea, Francisco Gomez-Donoso, Miguel Cazorla |
Comput. Vis. Image Underst. | 2 |
| 2017 | Automatic selection of molecular descriptors using random forest: Application to drug discoveryabstractThe optimal selection of chemical features (molecular descriptors) is an essential pre-processing step for the efficient application of computational intelligence techniques in virtual screening for identification of bioactive molecules in drug discovery . The selection of molecular descriptors has key influence in the accuracy of affinity prediction. In order to improve this prediction, we examined a Random Forest (RF)-based approach to automatically select molecular descriptors of training data for ligands of kinases, nuclear hormone receptors , and other enzymes . The reduction of features to use during prediction dramatically reduces the computing time over existing approaches and consequently permits the exploration of much larger sets of experimental data. To test the validity of the method, we compared the results of our approach with the ones obtained using manual feature selection in our previous study (Perez-Sanchez, Cano, and Garcia-Rodriguez, 2014).The main novelty of this work in the field of drug discovery is the use of RF in two different ways: feature ranking and dimensionality reduction, and classification using the automatically selected feature subset. Our RF-based method outperforms classification results provided by Support Vector Machine (SVM) and Neural Networks (NN) approaches. Gaspar Cano, José García Rodríguez 0001, Alberto Garcia-Garcia, Horacio Emilio Pérez Sánchez, Jón Atli Benediktsson, Anil Thapa, Alastair Barr |
Expert Syst. Appl. | 2 |
| 2017 | Multi-sensor 3D object dataset for object recognition with full pose estimation
Alberto Garcia-Garcia, Sergio Orts, Sergiu Ovidiu-Oprea, José García Rodríguez 0001, Jorge Azorín López, Marcelo Saval-Calvo, Miguel Cazorla |
Neural Comput. Appl. | 4 |
| 2017 | Constrained self-organizing feature map to preserve feature extraction topology
Jorge Azorín López, Marcelo Saval-Calvo, Andrés Fuster Guilló, José García Rodríguez 0001, Higinio Mora Mora |
Neural Comput. Appl. | 4 |
| 2017 | Editorial: special issue on computational intelligence for vision and robotics
José García Rodríguez 0001, Isabelle Guyon, Sergio Escalera, Alexandra Psarrou, Andrew Lewis 0004, Miguel Cazorla |
Neural Comput. Appl. | 1 |
| 2017 | Evaluation of sampling method effects in 3D non-rigid registration
Marcelo Saval-Calvo, Jorge Azorín López, Andrés Fuster Guilló, José García Rodríguez 0001, Sergio Orts, Alberto Garcia-Garcia |
Neural Comput. Appl. | 4 |
| 2017 | Object recognition in noisy RGB-D data using GNG
José Carlos Rangel, Vicente Morell, Miguel Cazorla, Sergio Orts, José García Rodríguez 0001 |
Pattern Anal. Appl. | 5 |
| 2017 | A robotic platform for customized and interactive rehabilitation of persons with disabilities
Francisco Gomez-Donoso, Sergio Orts, Alberto Garcia-Garcia, José García Rodríguez 0001, John Alejandro Castro-Vargas, Sergiu Ovidiu-Oprea, Miguel Cazorla |
Pattern Recognit. Lett. | 4 |
| 2016 | PointNet: A 3D Convolutional Neural Network for real-time object class recognitionabstractDuring the last few years, Convolutional Neural Networks are slowly but surely becoming the default method solve many computer vision related problems. This is mainly due to the continuous success that they have achieved when applied to certain tasks such as image, speech, or object recognition. Despite all the efforts, object class recognition methods based on deep learning techniques still have room for improvement. Most of the current approaches do not fully exploit 3D information, which has been proven to effectively improve the performance of other traditional object recognition methods. In this work, we propose PointNet, a new approach inspired by VoxNet and 3D ShapeNets, as an improvement over the existing methods by using density occupancy grids representations for the input data, and integrating them into a supervised Convolutional Neural Network architecture. An extensive experimentation was carried out, using ModelNet - a large-scale 3D CAD models dataset - to train and test the system, to prove that our approach is on par with state-of-the-art methods in terms of accuracy while being able to perform recognition under real-time constraints. Alberto Garcia-Garcia, Francisco Gomez-Donoso, José García Rodríguez 0001, Sergio Orts, Miguel Cazorla, Jorge Azorín López |
IJCNN | 3 |
| 2016 | Group activity description and recognition based on trajectory analysis and neural networksabstractThe recognition of group activities using computer vision and pattern recognition methods has been, and still remains, a challenging problem. Most of the research on human behaviour has been focused on recognizing individual issues from actions to behaviours. However, the analysis and recognition of group activities, the relationships of different groups in the scene and the interaction of the individuals in the group is still considered an open problem. This paper proposes a novel representation method to analyse and recognise group activities, called Group Activity Descriptor Vector (GADV). It is calculated from the trajectory described by the group and by the individuals who form it. Specifically, the GADV describes three different components: the trajectory followed by the group, the coherence of the individual trajectories in the group and, finally, the movement relationships among different groups in the scene. The trajectory analysis allows a simple high level understanding of complex groups activities. The GADV representation has been evaluated with different self-organizing neural networks using Behave and Caviar dataset sequences obtaining great accuracy in the recognition of the group activities, outperforming the state of the art methods. Jorge Azorín López, Marcelo Saval-Calvo, Andrés Fuster Guilló, José García Rodríguez 0001, Miguel Cazorla, María Teresa Signes Pont |
IJCNN | 4 |
| 2016 | Automatic Schaeffer's gestures recognition systemabstractAbstract Schaeffer's sign language consists of a reduced set of gestures designed to help children with autism or cognitive learning disabilities to develop adequate communication skills. Our automatic recognition system for Schaeffer's gesture language uses the information provided by an RGB‐D camera to capture body motion and recognize gestures using dynamic time warping combined with k‐nearest neighbors methods. The learning process is reinforced by the interaction with the proposed system that accelerates learning itself thus helping both children and educators. To demonstrate the validity of the system, a set of qualitative experiments with children were carried out. As a result, a system which is able to recognize a subset of 11 gestures of Schaeffer's sign language online was achieved. Francisco Gomez-Donoso, Miguel Cazorla, Alberto Garcia-Garcia, José García Rodríguez 0001 |
Expert Syst. J. Knowl. Eng. | 4 |
| 2016 | A Novel Prediction Method for Early Recognition of Global Human Behaviour in Image Sequences
Jorge Azorín López, Marcelo Saval-Calvo, Andrés Fuster Guilló, José García Rodríguez 0001 |
Neural Process. Lett. | 4 |
| 2016 | 3D Surface Reconstruction of Noisy Point Clouds Using Growing Neural Gas: 3D Object/Scene Reconstruction
Sergio Orts, José García Rodríguez 0001, Vicente Morell, Miguel Cazorla, Jose Antonio Serra-Perez, Alberto Garcia-Garcia |
Neural Process. Lett. | 2 |
| 2016 | Editorial: Neural Processing Letters Special Issue on "Neural Networks for Vision and Robotics"
José García Rodríguez 0001, Alexandra Psarrou, Andrew Lewis 0004, Anastassia Angelopoulou, Miguel Cazorla |
Neural Process. Lett. | 1 |
| 2015 | Self-Organizing Activity Description Map to represent and classify human behaviourabstractThe automated understanding of people activities from video sequences is an open research topic in which the computer vision and pattern recognition areas have made big efforts in recent years. This paper proposes the Self Organizing Activity Description Map (SOADM). It is a novel neural network based on the self-organizing paradigm to classify high level of semantic understanding from video sequences. The neural network is able to deal with the big gap between human trajectories in a scene and the global behaviour associated to them. Specifically, using simple representations of people trajectories as input, the SOADM is able to both represent and classify human behaviours. Additionally, the map is able to preserve the topological information about the scene. Experiments have been carried out using the Shopping Centre dataset of the CAVIAR database taken into account the global behaviour of an individual. Results confirm the high accuracy of the proposal outperforming previous methods. Jorge Azorín López, Marcelo Saval-Calvo, Andrés Fuster Guilló, José García Rodríguez 0001, Sergio Orts |
IJCNN | 4 |
| 2015 | Processing point cloud sequences with Growing Neural GasabstractWe consider the problem of processing point cloud sequences. In particular, we represent and track objects in dynamic scenes acquired using low-cost sensors such as the Kinect. A neural network based approach is proposed to represent and estimate 3D objects motion. This system addresses multiple computer vision tasks such as object segmentation, representation, motion analysis and tracking. The use of a neural network allows the unsupervised estimation of motion and the representation of objects in the scene. This proposal avoids the problem of finding corresponding features while tracking moving objects. A set of experiments are presented that demonstrate the validity of our method to track 3D objects. Favorable results are presented demonstrating the capabilities of the GNG algorithm for this task. Sergio Orts, José García Rodríguez 0001, Vicente Morell, Miguel Cazorla, Marcelo Saval-Calvo, Jorge Azorín López |
IJCNN | 2 |
| 2015 | Using GNG on 3D Object Recognition in Noisy RGB-D dataabstractThe object recognition task on 3D scenes is a growing research field that faces some problems relative to the use of 3D point clouds. In this work, we focus on dealing with the noise in the clouds through the use of the Growing Neural Gas (GNG) network filtering algorithm. The GNG method is able to represent the input data with a desired amount of neurons while preserving the topology of the input space. The selected recognition pipeline works describing extracted keypoints of the clouds, grouping and comparing it to detect the presence of an object in the scene, through a hypothesis verification algorithm. Experiments show how the GNG method yields better recognitions results that others filtering algorithms when noise is present. José Carlos Rangel, Vicente Morell, Miguel Cazorla, Sergio Orts, José García Rodríguez 0001 |
IJCNN | 5 |
| 2015 | Non-rigid point set registration using color and data downsamplingabstractNowadays, non-rigid registration problem is an active research topic in computer vision. Various proposals exist which face the problem from different perspectives, but it is still a challenging problem. Currently, with the new low-cost RGB-D sensors, the use of both, color and 3D information, is getting more interest in many applications. In this paper, we present a non-rigid registration technique based on CPD, and including color information along with 3D data, to estimate the non-rigid transformation. As the input data size is critical in the processing time, a sampling technique is required. Five sampling techniques are evaluated: a bilinear sampling, a normal-based, a color-based, a combination of the normal and color-based samplings, and a Growing Neural Gas based approach. All of them have been evaluated with the already presented non-rigid registration methods. Results show the performance of each sampling method, obtaining better results for the registration process using color-based sampling techniques. Marcelo Saval-Calvo, Sergio Orts, Jorge Azorín López, José García Rodríguez 0001, Andrés Fuster Guilló, Vicente Morell, Miguel Cazorla |
IJCNN | 4 |
| 2015 | 3D reconstruction of medical images from slices automatically landmarked with growing neural models
Anastassia Angelopoulou, Alexandra Psarrou, José García Rodríguez 0001, Sergio Orts, Jorge Azorín López, Kenneth Revett |
Neurocomputing | 3 |
| 2015 | SARASOM: a supervised architecture based on the recurrent associative SOM
David Gil, José García Rodríguez 0001, Miguel Cazorla, Magnus Johnsson |
Neural Comput. Appl. | 2 |
| 2014 | 3D maps representation using GNGabstractCurrent RGB-D sensors provide a big amount of valuable information for mobile robotics tasks like 3D map reconstruction, but the storage and processing of the incremental data provided by the different sensors through time quickly becomes unmanageable. In this work, we focus on 3D maps representation and we propose the use of a Growing Neural Gas (GNG) network as a 3D representation model of the input data. GNG method is able to represent the input data with a desired amount of neurons while preserving the topology of the input space. Experiments show how GNG method yields better input space adaptation than other state-of-the-art 3D map representation methods. Vicente Morell, Miguel Cazorla, Sergio Orts, José García Rodríguez 0001 |
IJCNN | 4 |
| 2014 | 3D colour object reconstruction based on Growing Neural GasabstractWith the advent of low-cost 3D sensors and 3D printers, surface reconstruction has become an important research topic in the last years. In this work, we propose an automatic method for 3D surface reconstruction from raw unorganized point clouds acquired using low-cost sensors. We have modified the Growing Neural Gas (GNG) network, which is a suitable model because of its flexibility, rapid adaptation and excellent quality of representation, to perform 3D surface reconstruction of different real-world objects. Some improvements have been made on the original algorithm considering colour information during the learning stage and creating complete triangular meshes instead of basic wire-frame representations. The proposed method is able to create 3D faces online, whereas existing 3D reconstruction methods based on Self-Organizing Maps (SOMs) required post-processing steps to close gaps and holes produced during the 3D reconstruction process. Performed experiments validated how the proposed method improves existing techniques removing post-processing steps and including colour information in the final triangular mesh. Sergio Orts, José García Rodríguez 0001, Vicente Morell, Miguel Cazorla, Juan Manuel García Chamizo |
IJCNN | 2 |
| 2014 | Combining visual features and Growing Neural Gas networks for robotic 3D SLAM
Diego Viejo, José García Rodríguez 0001, Miguel Cazorla |
Inf. Sci. | 2 |
| 2014 | Geometric 3D point cloud compression
Vicente Morell, Sergio Orts, Miguel Cazorla, José García Rodríguez 0001 |
Pattern Recognit. Lett. | 4 |
| 2013 | Adaptive learning in motion analysis with self-organising mapsabstractGrowing models have been widely used for clustering or topology learning. Traditionally these models work on stationary environments, grow incrementally and adapt their nodes to a given distribution based on global parameters. In this paper, we present an enhanced unsupervised self-organising network for the modelling of visual objects. We first develop a framework for building non-rigid shapes using the growth mechanism of the self-organising maps, and then we define an optimal number of nodes without overfitting or underfitting the network based on the knowledge obtained from information-theoretic considerations. This model is used to the representation of motion in image sequences by initialising a suitable segmentation. We present experimental results for hands and we quantitatively evaluate the matching capabilities of the proposed method with the topographic product. Anastassia Angelopoulou, José García Rodríguez 0001, Alexandra Psarrou, Gaurav Gupta 0001, Markos Mentzelopoulos |
IJCNN | 2 |
| 2013 | Human behaviour recognition based on trajectory analysis using neural networksabstractAutomated human behaviour analysis has been, and still remains, a challenging problem. It has been dealt from different points of views: from primitive actions to human interaction recognition. This paper is focused on trajectory analysis which allows a simple high level understanding of complex human behaviour. It is proposed a novel representation method of trajectory data, called Activity Description Vector (ADV) based on the number of occurrences of a person is in a specific point of the scenario and the local movements that perform in it. The ADV is calculated for each cell of the scenario in which it is spatially sampled obtaining a cue for different clustering methods. The ADV representation has been tested as the input of several classic classifiers and compared to other approaches using CAVIAR dataset sequences obtaining great accuracy in the recognition of the behaviour of people in a Shopping Centre. Jorge Azorín López, Marcelo Saval-Calvo, Andrés Fuster Guilló, José García Rodríguez 0001 |
IJCNN | 4 |
| 2013 | A semi-parametric approach for football video annotationabstractAutomatic sports video segmentation is a fast growth area of research in the visual information retrieval field. This paper presents a semi-parametric algorithm for parsing football video structures. The approach works on a two interleaved based process that closely collaborate towards a common goal. The core part of the proposed method focus performs a fast automatic football video annotation by looking at the enhance entropy variance within a series of shot frames. The entropy is extracted on the Hue parameter from the HSV color system, not as a global feature but in spatial domain to identify regions within a shot that will characterize a certain activity within the shot period. The second part of the algorithm works towards the identification of dominant color regions that could represent players and playfield for further activity recognition. Experimental results shows that the proposed football video segmentation algorithm performs with high accuracy. Markos Mentzelopoulos, Alexandra Psarrou, Anastassia Angelopoulou, José García Rodríguez 0001 |
IJCNN | 4 |
| 2013 | Point cloud data filtering and downsampling using growing neural gasabstract3D sensors provide valuable information for mobile robotic tasks like scene classification or object recognition, but these sensors often produce noisy data that makes impossible applying classical keypoint detection and feature extraction techniques. Therefore, noise removal and downsampling have become essential steps in 3D data processing. In this work, we propose the use of a 3D filtering and downsampling technique based on a Growing Neural Gas (GNG) network. GNG method is able to deal with outliers presents in the input data. These features allows to represent 3D spaces, obtaining an induced Delaunay Triangulation of the input space. Experiments show how GNG method yields better input space adaptation to noisy data than other filtering and downsampling methods like Voxel Grid. It is also demonstrated how the state-of-the-art keypoint detectors improve their performance using filtered data with GNG network. Descriptors extracted on improved keypoints perform better matching in robotics applications as 3D scene registration. Sergio Orts, Vicente Morell, José García Rodríguez 0001, Miguel Cazorla |
IJCNN | 3 |
| 2013 | Improving drug discovery using a neural networks based parallel scoring functionabstractVirtual Screening (VS) methods can considerably aid clinical research, predicting how ligands interact with drug targets. Most VS methods suppose a unique binding site for the target, but it has been demonstrated that diverse ligands interact with unrelated parts of the target and many VS methods do not take into account this relevant fact. This problem is circumvented by a novel VS methodology named BINDSURF that scans the whole protein surface to find new hotspots, where ligands might potentially interact with, and which is implemented in massively parallel Graphics Processing Units, allowing fast processing of large ligand databases. BINDSURF can thus be used in drug discovery, drug design, drug repurposing and therefore helps considerably in clinical research. However, the accuracy of most VS methods is constrained by limitations in the scoring function that describes biomolecular interactions, and even nowadays these uncertainties are not completely understood. In order to solve this problem, we propose a novel approach where neural networks are trained with databases of known active (drugs) and inactive compounds, and later used to improve VS predictions. Horacio Emilio Pérez Sánchez, Ginés D. Guerrero, José M. García 0001, Jorge Peña-García, José M. Cecilia, Gaspar Cano, Sergio Orts, José García Rodríguez 0001 |
IJCNN | 8 |
| 2013 | 3D gesture recognition with growing neural gasabstractWe propose the design of a real-time system to recognize and interpret hand gestures. The acquisition devices are low cost 3D sensors. 3D hand pose segmentation, characterization and tracking will be implemented using the growing neural gas (GNG) structure. The capacity of the system to obtain information with a high degree of freedom allows the encoding of many gestures and a very accurate motion capture. The use of hand pose models combined with motion information provided with GNG permits to deal with the problem of the hand motion representation. A natural interface applied to a virtual mirror writing system and a module to estimate hand pose have been designed to demonstrate the validity of the system. Jose Antonio Serra-Perez, José García Rodríguez 0001, Sergio Orts, Juan Manuel García Chamizo, Javier Montoyo-Bojo, Anastassia Angelopoulou, Alexandra Psarrou, Markos Mentzelopoulos, Andrew Lewis 0004 |
IJCNN | 2 |
| 2013 | Active Foreground Region Extraction and Tracking for Sports Video Annotation
Markos Mentzelopoulos, Alexandra Psarrou, Anastassia Angelopoulou, José García Rodríguez 0001 |
Neural Process. Lett. | 4 |
| 2012 | Region analysis through close contour transformation using growing neural gasabstractIn our work we aim to explore a general framework that addresses the fundamental problem of universal unsupervised extraction of semantically meaningful visual regions. To this end this paper describes a novel region analysis technique using a self-organising map, the growing neural gas, which is adapted so as to improve modelling speed as well as to ensure a double-linkage chain around all region contours to simplify shape analysis. While the growing neural gas has been extensively applied to shape modelling, it has never explicitly been used for curvature analysis, contour description and region similarity. Once a contour network has been obtained, a transformation is applied that converts the closed contour to an open one, facilitating the use of certain angular descriptors. Discriminative descriptors derived from the properties of regions, their contours and their transformed contours are established and define a feature vector used for the representation of regions based on the appearance and contour information. Gaurav Gupta 0001, Alexandra Psarrou, Anastassia Angelopoulou, José García Rodríguez 0001 |
IJCNN | 4 |
| 2012 | Multi-GPU based camera network system keeps privacy using growing neural gasabstractIn this work we present a multi-camera surveillance system based on the use of self-organizing neural networks to represent events in video. The objectives include: identifying and tracking persons or objects in the scene or the interpretation of user gestures for interaction with services, devices and systems implemented in the digital home. Additionally, the system process several tasks in parallel using GPUs (Graphic Processor Units). Addressing multiple vision tasks of various levels such as segmentation, representation or characterization, analysis and monitoring of the movement to allow the construction of a robust representation of their environment and interpret the elements of the scene. Sergio Orts, José García Rodríguez 0001, Vicente Morell, Jorge Azorín López, Juan Manuel García Chamizo |
IJCNN | 2 |
| 2012 | GPGPU implementation of growing neural gas: Application to 3D scene reconstruction
Sergio Orts, José García Rodríguez 0001, Diego Viejo, Miguel Cazorla, Vicente Morell |
J. Parallel Distributed Comput. | 2 |
| 2012 | Autonomous Growing Neural Gas for applications with time constraint: Optimal parameter estimation
José García Rodríguez 0001, Anastassia Angelopoulou, Juan Manuel García Chamizo, Alexandra Psarrou, Sergio Orts, Vicente Morell |
Neural Networks | 1 |
| 2012 | Using GNG to improve 3D feature extraction - Application to 6DoF egomotion
Diego Viejo, José García Rodríguez 0001, Miguel Cazorla, David Gil, Magnus Johnsson |
Neural Networks | 2 |
| 2011 | Predictions tasks with words and sequences: Comparing a novel recurrent architecture with the Elman networkabstractThe classical connectionist models are not well suited to working with data varying over time. According to this, temporal connectionist models have emerged and constitute a continuously growing research field. In this paper we present a novel supervised recurrent neural network architecture (SARASOM) based on the Associative Self-Organizing Map (A-SOM). The A-SOM is a variant of the Self-Organizing Map (SOM) that develops a representation of its input space as well as learns to associate its activity with an arbitrary number of additional inputs. In this context the A-SOM learns to associate its previous activity with a delay of one iteration. The performance of the SARASOM was evaluated and compared with the Elman network in a number of prediction tasks using sequences of letters (including some experiments with a reduced lexicon of 10 words). The results are very encouraging with SARASOM learning slightly better than the Elman network. David Gil, José García Rodríguez 0001, Miguel Cazorla, Magnus Johnsson |
IJCNN | 2 |
| 2011 | Fast Autonomous Growing Neural GasabstractThis paper aims to address the ability of self-organizing neural network models to manage real-time applications. Specifically, we introduce fAGNG (fast Autonomous Growing Neural Gas), a modified learning algorithm for the incremental model Growing Neural Gas (GNG) network. The Growing Neural Gas network with its attributes of growth, flexibility, rapid adaptation, and excellent quality of representation of the input space makes it a suitable model for real time applications. However, under time constraints GNG fails to produce the optimal topological map for any input data set. In contrast to existing algorithms the proposed fAGNG algorithm introduces multiple neurons per iteration. The number of neurons inserted and input data generated is controlled autonomous and dynamically based on a priory learnt model. Comparative experiments using topological preservation measures are carried out to demonstrate the effectiveness of the new algorithm to represent linear and non-linear input spaces under time restrictions. José García Rodríguez 0001, Anastassia Angelopoulou, Juan Manuel García Chamizo, Alexandra Psarrou, Sergio Orts, Vicente Morell |
IJCNN | 1 |
| 2011 | Using 3D GNG-based reconstruction for 6DoF egomotionabstractSeveral recent works deal with 3D data in mobile robotic problems, e.g. mapping. Data come from any kind of sensor (time of flight cameras and 3D lasers) providing a huge amount of unorganized 3D data. In this paper we detail an efficient method to build complete 3D models from a Growing Neural Gas (GNG). We show that the use of GNG provides better results than other approaches. The GNG obtained is then applied to a sequence. From GNG structure, we propose to calculate planar patches and thus obtaining a fast method to compute the movement performed by a mobile robot by means of a 3D models registration algorithm. Final results of 3D mapping are also shown. Diego Viejo, José García Rodríguez 0001, Miguel Cazorla, David Gil, Magnus Johnsson |
IJCNN | 2 |
| 2010 | Tracking gestures using a probabilistic Self-Organising networkabstractThe Self-Organising Artificial Neural Network Models, of which we have used the Growing Neural Gas (GNG) can be applied to preserve the topology of an input distribution. Traditionally these models neither do include local adaptation of the nodes nor colour information. In this paper, we present an extension to the original growing neural gas network that has probabilistic features and can be applied to preserve the topology of a non-stationary distribution. The network consists of the geometrical position of the nodes, the underline local feature structure of the image, and the distance vector between the modal image and any successive images. Accurate correspondence of the nodes between successive images, is measured through the calculation of the topographic product. The method performs continuously mapping over a distribution that changes over time and works with both smooth and abrupt changes. The method is successfully applied to object modelling and tracking. Anastassia Angelopoulou, Alexandra Psarrou, José García Rodríguez 0001, Gaurav Gupta 0001 |
IJCNN | 3 |
| 2010 | Hand gesture modelling and tracking using a Self-Organising NetworkabstractThe Self-Organising Artificial Neural Network Models, of which we have used the Growing Neural Gas (GNG) can be applied to preserve the topology of an input distribution. Traditionally these models neither do include local adaptation of the nodes nor colour information. In this paper, we extend GNG by presenting an improvement to the network that has both global and local properties and can track in cluttered backgrounds. The method performs continuously mapping over a distribution that changes over time and works with both smooth and abrupt changes. The central mechanism relies on the addition of global and local attributes, and skin colour information to the network which allow us to automatically model and track 2D gestures. Application to hand gesture video tracking is presented. Anastassia Angelopoulou, José García Rodríguez 0001, Alexandra Psarrou, Gaurav Gupta 0001 |
IJCNN | 2 |
| 2010 | GNG based surveillance systemabstractSelf-organising neural networks have shown promise in a variety of applications areas. Their massive and intrinsic parallelism makes those networks suitable to solve hard problems in image-analysis and computer vision applications, especially when non-stationary environments occur. Moreover, this kind of neural networks preserves the topology of an input space by using their inherited competitive learning property. In this work we use a kind of self-organising network, the Growing Neural Gas, to solve some computer vision tasks applied to visual surveillance systems. It has been used their capacity to represent non rigid objects as a result of an adaptive process by a topology-preserving graph that constitutes an induced Delaunay triangulation of their shapes. The neural network is also modified to accelerate the learning algorithm in order to support applications with temporal constraints. This feature has been used to build a system able to track image features in video sequences. The system automatically keeps the correspondence of features among frames in the sequence using its own structure. Information obtained during the tracking process and allocated in the neural network can also be used to analyse the objects motion. José García Rodríguez 0001, Anastassia Angelopoulou, Juan Manuel García Chamizo, Alexandra Psarrou |
IJCNN | 1 |
| 2007 | Robust Modelling and Tracking of NonRigid Objects Using Active-GNGabstractThis paper presents a robust approach to nonrigid modelling and tracking. The contour of the object is described by an active growing neural gas (A-GNG) network which allows the model to re-deform locally. The approach is novel in that the nodes of the network are described by their geometrical position, the underlying local feature structure of the image, and the distance vector between the modal image and any successive images. A second contribution is the correspondence of the nodes which is measured through the calculation of the topographic product, a topology preserving objective function which quantifies the neighbourhood preservation before and after the mapping. As a result, we can achieve the automatic modelling and tracking of objects without using any annotated training sets. Experimental results have shown the superiority of our proposed method over the original growing neural gas (GNG) network. Anastassia Angelopoulou, Alexandra Psarrou, Gaurav Gupta 0001, José García Rodríguez 0001 |
ICCV | 4 |
| 2007 | Image Compression Using Growing Neural GasabstractIn this paper we study the capacities of characterization and synthesis of objects by using a self-organizing neural model, the Growing Neural Gas. These networks, by means of their competitive learning try to preserve the topology of an input space. This feature is being used for the representation of objects and their movement with topology preserving networks. We characterize the object to be represented by means of the obtained maps and kept information solely on the coordinates and the pixel color of the neurons. With this information it is made the synthesis of the original images, applying mathematical morphology and simple filters using the available information. José García Rodríguez 0001, Francisco Flórez-Revuelta, Juan Manuel García Chamizo |
IJCNN | 1 |
| 2006 | Automatically Building 2D Statistical Shapes Using the Topology Preservation Model GNG
José García Rodríguez 0001, Anastassia Angelopoulou, Alexandra Psarrou, Kenneth Revett |
ACCV (1) | 1 |
| 2006 | Learning 2D Hand Shapes Using the Topology Preservation Model GNG
Anastassia Angelopoulou, José García Rodríguez 0001, Alexandra Psarrou |
ECCV (1) | 2 |
| 2006 | Growing Neural Gas for Vision Tasks with Time Restrictions
José García Rodríguez 0001, Francisco Flórez-Revuelta, Juan Manuel García Chamizo |
ICANN (2) | 1 |
| 2006 | Measuring GNG Topology Preservation in Computer Vision Applications
José García Rodríguez 0001, Francisco Flórez-Revuelta, Juan Manuel García Chamizo |
KES (3) | 1 |