Estefanía Talavera

dblp:163/2269 · also Estefania Talavera, Estefania Talavera Martínez · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
14since 2021 · last 2025
0000-0001-5918-8990ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 DADO: A Depth-Attention Framework for Object Discovery
Federico Gonzalez, Estefanía Talavera, Petia Radeva
CAIP (2)2
2025 Vision on the Move: Automated Hazardous Material Plate Detection in Freight Transport
Melissa Tijink, Stanislav Levendeev, Ewaldo Nieuwenhuis, Luuk J. Spreeuwers, Nicola Strisciuglio, Estefanía Talavera
CAIP (1)6
2025 Crime scene classification from skeletal trajectory analysis in surveillance settings
abstract
Video anomaly analysis is a core task to the field of computer vision, particularly for crime detection in surveillance footage. In this work, we address the task of human-related crime classification using skeletal joint trajectories extracted from surveillance video frames. First, we emphasize the need to enhance the ground truth labels for the Human-Related Crime dataset (HR-Crime) and propose supervised and unsupervised methodologies to generate trajectory-level labels. Next, based on the trajectory-level labels, we introduce a trajectory-based crime classification framework, evaluating different architectures and feature fusion strategies for representing human trajectories. Our experiments validate the approach and open avenues for future research on this topic.
Alina Matei, Estefanía Talavera, Maya Aghaei
Eng. Appl. Artif. Intell.2
2024 Body-part Tubelet Transformer for Human-Related Crime Classification
abstract
Detecting human-related crimes from surveillance videos poses an increasingly difficult challenge, especially when confronted with human actions that are relatively similar. In this work, we propose a transformer-based model that induces bias through the incorporation of a Tubelet embedder module-a 3D convolutional layer. The aim is to capture spatiotemporal embeddings from skeletal trajectories extracted from videos using 3D convolutional operations. Our experiments are conducted on the Human-Related Crime dataset, revealing that the use of tubelet embeddings maintains competitive performance (49% accuracy) to the state-of-the-art, while considerably reducing the computational complexity of the model.
Ajay Mathew Joseph, Fath U Min Ullah, Estefanía Talavera
AVSS3
2024 Dual Deep Learning Network for Abnormal Action Detection
abstract
Neural networks have demonstrated remarkable effectiveness in solving distinct real-world vision problems pertaining to activity recognition and violence detection in surveillance scenarios. The broad reliance on practicing a single network for spatial and motion information collection has made them less effective for long-term dependency analysis in video snippets. Our work solves this issue through a multi-network fusion strategy suitable for real-world surveillance. Initially, the spatial information is accessed from a compound coefficient strategy inspired by a robust convolutional neural network (ConvNet). Next, the pyramidal convolutional features from two consecutive frames are obtained through LiteFlowNet. The output from both the networks (ConvNet and LiteFlowNet) is separately passed into a deep-gated recurrent Unit (GRU) that is assembled for a skip connection. The latter obtained from each GRU is fused and further propagated to the dense layer for final decision. The results on the datasets and the ablation study confirm our method’s efficiency, outperforming the state-of-the-art methods. (Code: GitHub)
Fath U Min Ullah, Zulfiqar Ahmad Khan 0002, Sung Wook Baik, Estefanía Talavera, Saeed Anwar, Khan Muhammad 0001
AVSS4
2024 CAST: Clustering self-Attention using Surrogate Tokens for efficient transformers
abstract
The Transformer architecture has shown to be a powerful tool for a wide range of tasks. It is based on the self-attention mechanism, which is an inherently computationally expensive operation with quadratic computational complexity: memory usage and compute time increase quadratically with the length of the input sequences, thus limiting the application of Transformers. In this work, we propose a novel Clustering self-Attention mechanism using Surrogate Tokens (CAST), to optimize the attention computation and achieve efficient transformers. CAST utilizes learnable surrogate tokens to construct a cluster affinity matrix, used to cluster the input sequence and generate novel cluster summaries. The self-attention from within each cluster is then combined with the cluster summaries of other clusters, enabling information flow across the entire input sequence. CAST improves efficiency by reducing the complexity from O ( N 2 ) to O ( α N ) where N is the sequence length, and α is constant according to the number of clusters and samples per cluster. We show that CAST performs better than or comparable to the baseline Transformers on long-range sequence modeling tasks, while also achieving higher results on time and memory efficiency than other efficient transformers. • Computation of self-attention Transformers is limited by the input sequence length. • We propose CAST, an efficient self-attention mechanism with clustered attention. • We propose the use of surrogate tokens to optimize self-attention in transformers. • We observe that cluster summaries enhance training efficiency and results. • CAST reduces complexity of self-attention computation from O ( N 2 ) to O ( α N ) .
Adjorn van Engelenhoven, Nicola Strisciuglio, Estefanía Talavera
Pattern Recognit. Lett.3
2023 Fall Detection with Event-Based Data: A Case Study
Nicoletta Risi, Estefanía Talavera, Elisabetta Chicca, Dimka Karastoyanova, George Azzopardi
CAIP (2)3
2023 Behavioural patterns discovery for lifestyle analysis from egocentric photo-streams
Martín Menchón, Estefanía Talavera, Jose M. Massa, Petia Radeva
Pervasive Mob. Comput.2
2022 Spatial-Temporal Transformer for Crime Recognition in Surveillance Videos
abstract
Human-related crime recognition from surveillance videos becomes an even more challenging task when dealing with relatively similar human actions. We propose a transformer-based model that relies on the spatial-temporal representation of extracted skeletal trajectories for fine-grained classification. We validate the effectiveness of our model on the complex HR-Crime dataset consisting of videos representing 13 categories of human-related crimes. Quantitative and qualitative results suggest that building a transformer architecture with coupled spatial and temporal modules enables the model to compete in performance while improving intrinsic interpretability.
Kayleigh Boekhoudt, Estefanía Talavera
AVSS2
2022 InstaIndoor and multi-modal deep learning for indoor scene recognition
Andreea Glavan, Estefanía Talavera
Neural Comput. Appl.2
2022 Correction to: Playing to distraction: towards a robust training of CNN classifiers through visual explanation techniques
David Morales, Estefanía Talavera, Beatriz Remeseiro
Neural Comput. Appl.2
2021 HR-Crime: Human-Related Anomaly Detection in Surveillance Videos
Kayleigh Boekhoudt, Alina Matei, Maya Aghaei, Estefanía Talavera
CAIP (2)4
2021 Does our social life influence our nutritional behaviour? Understanding nutritional habits from egocentric photo-streams
abstract
Nutrition and social interactions are both key aspects of the daily lives of humans. In this work, we propose a system to evaluate the influence of social interaction in the nutritional habits of a person from a first-person perspective. In order to detect the routine of an individual, we construct a nutritional behaviour pattern discovery model, which outputs routines over a number of days. Our method evaluates similarity of routines with respect to visited food-related scenes over the collected days, making use of Dynamic Time Warping, as well as considering social engagement and its correlation with food-related activities. The nutritional and social descriptors of the collected days are evaluated and encoded using an LSTM Autoencoder. Later, the obtained latent space is clustered to find similar days unaffected by outliers using the Isolation Forest method. Moreover, we introduce a new score metric to evaluate the performance of the proposed algorithm. We validate our method on 104 days and more than 100 k egocentric images gathered by 7 users. Several different visualizations are evaluated for the understanding of the findings. Our results demonstrate good performance and applicability of our proposed model for social-related nutritional behaviour understanding. At the end, relevant applications of the model are discussed by analysing the discovered routine of particular individuals.
Andreea Glavan, Alina Matei, Petia Radeva, Estefanía Talavera
Expert Syst. Appl.4
2021 Playing to distraction: towards a robust training of CNN classifiers through visual explanation techniques
David Morales, Estefanía Talavera, Beatriz Remeseiro
Neural Comput. Appl.2
2020 Topic modelling for routine discovery from egocentric photo-streams
abstract
Developing tools to understand and visualize lifestyle is of high interest when addressing the improvement of habits and well-being of people. Routine, defined as the usual things that a person does daily, helps describe the individuals’ lifestyle. With this paper, we are the first ones to address the development of novel tools for automatic discovery of routine days of an individual from his/her egocentric images. In the proposed model, sequences of images are firstly characterized by semantic labels detected by pre-trained CNNs. Then, these features are organized in temporal-semantic documents to later be embedded into a topic models space. Finally, Dynamic-Time-Warping and Spectral-Clustering methods are used for final day routine/non-routine discrimination. Moreover, we introduce a new EgoRoutine-dataset, a collection of 104 egocentric days with more than 100.000 images recorded by 7 users. Results show that routine can be discovered and behavioural patterns can be observed.
Estefanía Talavera, Carolin Wuerich, Nicolai Petkov, Petia Radeva
Pattern Recognit.1
2020 Hierarchical Approach to Classify Food Scenes in Egocentric Photo-Streams
abstract
Recent studies have shown that the environment where people eat can affect their nutritional behavior [1]. In this paper, we provide automatic tools for personalized analysis of a person's health habits by the examination of daily recorded egocentric photo-streams. Specifically, we propose a new automatic approach for the classification of food-related environments, that is able to classify up to 15 such scenes. In this way, people can monitor the context around their food intake in order to get an objective insight into their daily eating routine. We propose a model that classifies food-related scenes organized in a semantic hierarchy. Additionally, we present and make available a new egocentric dataset composed of more than 33 000 images recorded by a wearable camera, over which our proposed model has been tested. Our approach obtains an accuracy and F-score of 56% and 65%, respectively, clearly outperforming the baseline methods.
Estefanía Talavera, Maria Leyva-Vallina, Md. Mostafa Kamal Sarker, Domenec Puig, Nicolai Petkov, Petia Radeva
IEEE J. Biomed. Health Informatics1
2019 Unsupervised Routine Discovery in Egocentric Photo-Streams
Estefanía Talavera, Nicolai Petkov, Petia Radeva
CAIP (1)1
2017 SR-clustering: Semantic regularized clustering for egocentric photo streams segmentation
abstract
While wearable cameras are becoming increasingly popular, locating relevant information in large unstructured collections of egocentric images is still a tedious and time consuming process. This paper addresses the problem of organizing egocentric photo streams acquired by a wearable camera into semantically meaningful segments, hence making an important step towards the goal of automatically annotating these photos for browsing and retrieval. In the proposed method, first, contextual and semantic information is extracted for each image by employing a Convolutional Neural Networks approach. Later, a vocabulary of concepts is defined in a semantic space by relying on linguistic information. Finally, by exploiting the temporal coherence of concepts in photo streams, images which share contextual and semantic attributes are grouped together. The resulting temporal segmentation is particularly suited for further analysis, ranging from event recognition to semantic indexing and summarization. Experimental results over egocentric set of nearly 31,000 images, show the prominence of the proposed approach over state-of-the-art methods.
Mariella Dimiccoli, Marc Bolaños, Estefanía Talavera, Maedeh Aghaei, Stavri G. Nikolov, Petia Radeva
Comput. Vis. Image Underst.3