VLDB 2026 Research / reviewers in the wild / expert
Jordi Gonzàlez 0001
dblp:24/3310-1
· DBLP profile ↗
80ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0001-8033-0306ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 57 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 6 since 2021Human-computer interaction and ubiquitous computing · 4Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Debiasing CLIP with Neural Interventions
Amelia Gómez Grabowska, Jordi Gonzàlez 0001, Lluís Gómez i Bigorda |
ECIR (3) | 2 |
| 2025 | Facial Expression Generation from Text with FaceCLIP
Wenwen Fu, Wenjuan Gong, Chen-Yang Yu, Wei Wang 0115, Jordi Gonzàlez 0001 |
J. Comput. Sci. Technol. | 5 |
| 2025 | Audio-visual scene recognition using attention-based graph convolutional model
Yikai Wu 0004, Wenjuan Gong, Jordi Gonzàlez 0001 |
Multim. Tools Appl. | 5 |
| 2024 | FedSOKD-TFA: Federated Learning with Stage-Optimal Knowledge Distillation and Three-Factor Aggregation
Jianhao Liu, Wenjuan Gong, Tingbo Shi, Kechen Li, Jordi Gonzàlez 0001 |
ICPR (2) | 6 |
| 2024 | A Generative Multi-Resolution Pyramid and Normal-Conditioning 3D Cloth DrapingabstractRGB cloth generation has been deeply studied in the related literature, however, 3D garment generation remains an open problem. In this paper, we build a conditional variational autoencoder for 3D garment generation and draping. We propose a pyramid network to add garment details progressively in a canonical space, i.e. unposing and unshaping the garments w.r.t. the body. We study conditioning the network on surface normal UV maps, as an intermediate representation, which is an easier problem to optimize than 3D coordinates. Our results on two public datasets, CLOTH3D and CAPE, show that our model is robust, controllable in terms of detail generation by the use of multi-resolution pyramids, and achieves state-of-the-art results that can highly generalize to unseen garments, poses, and shapes even when training with small amounts of data. The code can be found at: https://github.com/HunorLaczko/pyramid-drape Hunor Laczkó, Meysam Madadi, Sergio Escalera, Jordi Gonzàlez 0001 |
WACV | 4 |
| 2024 | MCLEMCD: multimodal collaborative learning encoder for enhanced music classification from dances
Wenjuan Gong, Qingshuang Yu, Wendong Huang, Peng Cheng 0008, Jordi Gonzàlez 0001 |
Multim. Syst. | 6 |
| 2024 | Meta-MMFNet: Meta-learning-based Multi-model Fusion Network for Micro-expression RecognitionabstractDespite its wide applications in criminal investigations and clinical communications with patients suffering from autism, automatic micro-expression recognition remains a challenging problem because of the lack of training data and imbalanced classes problems. In this study, we proposed a meta-learning-based multi-model fusion network (Meta-MMFNet) to solve the existing problems. The proposed method is based on the metric-based meta-learning pipeline, which is specifically designed for few-shot learning and is suitable for model-level fusion. The frame difference and optical flow features were fused, deep features were extracted from the fused feature, and finally in the meta-learning-based framework, weighted sum model fusion method was applied for micro-expression classification. Meta-MMFNet achieved better results than state-of-the-art methods on four datasets. The code is available at https://github.com/wenjgong/meta-fusion-based-method . Wenjuan Gong, Yue Zhang 0087, Wei Wang 0115, Peng Cheng 0008, Jordi Gonzàlez 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Single image super-resolution based on directional variance attention network
Parichehr Behjati, Pau Rodríguez, Carles Fernández, Isabelle Hupont, Armin Mehri, Jordi Gonzàlez 0001 |
Pattern Recognit. | 6 |
| 2022 | End-to-end global to local convolutional neural network learning for hand pose recovery in depth dataabstractAbstract Despite recent advances in 3‐D pose estimation of human hands, thanks to the advent of convolutional neural networks (CNNs) and depth cameras, this task is still far from being solved in uncontrolled setups. This is mainly due to the highly non‐linear dynamics of fingers and self‐occlusions, which make hand model training a challenging task. In this study, a novel hierarchical tree‐like structured CNN is exploited, in which branches are trained to become specialised in predefined subsets of hand joints called local poses. Further, local pose features, extracted from hierarchical CNN branches, are fused to learn higher order dependencies among joints in the final pose by end‐to‐end training. Lastly, the loss function used is also defined to incorporate appearance and physical constraints about doable hand motions and deformations. Finally, a non‐rigid data augmentation approach is introduced to increase the amount of training depth data. Experimental results suggest that feeding a tree‐shaped CNN, specialised in local poses, into a fusion network for modelling joints' correlations and dependencies, helps to increase the precision of final estimations, showing competitive results on NYU, MSRA, Hands17 and SyntheticHand datasets. Meysam Madadi, Sergio Escalera, Xavier Baró, Jordi Gonzàlez 0001 |
IET Comput. Vis. | 4 |
| 2022 | A Closer Look at Embedding Propagation for Manifold SmoothingabstractSupervised training of neural networks requires a large amount of manually annotated data and the resulting networks tend to be sensitive to out-of-distribution (OOD) data. Self- and semi-supervised training schemes reduce the amount of annotated data required during the training process. However, OOD generalization remains a major challenge for most methods. Strategies that promote smoother decision boundaries play an important role in out-of-distribution generalization. For example, embedding propagation (EP) for manifold smoothing has recently shown to considerably improve the OOD performance for few-shot classification. EP achieves smoother class manifolds by building a graph from sample embeddings and propagating information through the nodes in an unsupervised manner. In this work, we extend the original EP paper providing additional evidence and experiments showing that it attains smoother class embedding manifolds and improves results in settings beyond few-shot classification. Concretely, we show that EP improves the robustness of neural networks against multiple adversarial attacks as well as semi- and self-supervised learning performance. Diego A. Velázquez, Pau Rodríguez, Josep M. Gonfaus, F. Xavier Roca, Jordi Gonzàlez 0001 |
J. Mach. Learn. Res. | 5 |
| 2022 | Image rain removal and illumination enhancement done in one go
Yecong Wan, Yuanshuo Cheng, Ming-Wen Shao, Jordi Gonzàlez 0001 |
Knowl. Based Syst. | 4 |
| 2022 | Deep Pain: Exploiting Long Short-Term Memory Networks for Facial Expression ClassificationabstractPain is an unpleasant feeling that has been shown to be an important factor for the recovery of patients. Since this is costly in human resources and difficult to do objectively, there is the need for automatic systems to measure it. In this paper, contrary to current state-of-the-art techniques in pain assessment, which are based on facial features only, we suggest that the performance can be enhanced by feeding the raw frames to deep learning models, outperforming the latest state-of-the-art results while also directly facing the problem of imbalanced data. As a baseline, our approach first uses convolutional neural networks (CNNs) to learn facial features from VGG_Faces, which are then linked to a long short-term memory to exploit the temporal relation between video frames. We further compare the performances of using the so popular schema based on the canonically normalized appearance versus taking into account the whole image. As a result, we outperform current state-of-the-art area under the curve performance in the UNBC-McMaster Shoulder Pain Expression Archive Database. In addition, to evaluate the generalization properties of our proposed methodology on facial motion recognition, we also report competitive results in the Cohn Kanade+ facial expression database. Pau Rodríguez, Guillem Cucurull, Jordi Gonzàlez 0001, Josep M. Gonfaus, Kamal Nasrollahi, Thomas B. Moeslund, F. Xavier Roca |
IEEE Trans. Cybern. | 3 |
| 2021 | OverNet: Lightweight Multi-Scale Super-Resolution with Overscaling NetworkabstractSuper-resolution (SR) has achieved great success due to the development of deep convolutional neural networks (CNNs). However, as the depth and width of the networks increase, CNN-based SR methods have been faced with the challenge of computational complexity in practice. More-over, most SR methods train a dedicated model for each target resolution, losing generality and increasing memory requirements. To address these limitations we introduce OverNet, a deep but lightweight convolutional network to solve SISR at arbitrary scale factors with a single model. We make the following contributions: first, we introduce a lightweight feature extractor that enforces efficient reuse of information through a novel recursive structure of skip and dense connections. Second, to maximize the performance of the feature extractor, we propose a model agnostic reconstruction module that generates accurate high-resolution images from overscaled feature maps obtained from any SR architecture. Third, we introduce a multi-scale loss function to achieve generalization across scales. Experiments show that our proposal outperforms previous state-of-the-art approaches in standard benchmarks, while maintaining relatively low computation and memory requirements. Parichehr Behjati, Pau Rodríguez, Armin Mehri, Isabelle Hupont, Carles Fernández Tena, Jordi Gonzàlez 0001 |
WACV | 6 |
| 2020 | Pay Attention to the Activations: A Modular Attention Mechanism for Fine-Grained Image RecognitionabstractFine-grained image recognition is central to many multimedia tasks such as search, retrieval, and captioning. Unfortunately, these tasks are still challenging since the appearance of samples of the same class can be more different than those from different classes. This issue is mainly due to changes in deformation, pose, and the presence of clutter. In the literature, attention has been one of the most successful strategies to handle the aforementioned problems. Attention has been typically implemented in neural networks by selecting the most informative regions of the image that improve classification. In contrast, in this paper, attention is not applied at the image level but to the convolutional feature activations. In essence, with our approach, the neural model learns to attend to lower-level feature activations without requiring part annotations and uses those activations to update and rectify the output likelihood distribution. The proposed mechanism is modular, architecture-independent, and efficient in terms of both parameters and computation required. Experiments demonstrate that well-known networks such as wide residual networks and ResNeXt, when augmented with our approach, systematically improve their classification accuracy and become more robust to changes in deformation and pose and to the presence of clutter. As a result, our proposal reaches state-of-the-art classification accuracies in CIFAR-10, the Adience gender recognition task, Stanford Dogs, and UEC-Food100 while obtaining competitive performance in ImageNet, CIFAR-100, CUB200 Birds, and Stanford Cars. In addition, we analyze the different components of our model, showing that the proposed attention modules succeed in finding the most discriminative regions of the image. Finally, as a proof of concept, we demonstrate that with only local predictions, an augmented neural network can successfully classify an image before reaching any fully connected layer, thus reducing the computational amount up to 10%. Pau Rodríguez, Diego Velazquez Dorta, Guillem Cucurull, Josep M. Gonfaus, F. Xavier Roca, Jordi Gonzàlez 0001 |
IEEE Trans. Multim. | 6 |
| 2019 | From 2D to 3D geodesic-based garment matching
Egils Avots, Meysam Madadi, Sergio Escalera, Jordi Gonzàlez 0001, Xavier Baró, Paul Pällin, Gholamreza Anbarjafari |
Multim. Tools Appl. | 4 |
| 2018 | Attend and Rectify: A Gated Attention Mechanism for Fine-Grained Recovery
Pau Rodríguez, Josep M. Gonfaus, Guillem Cucurull, F. Xavier Roca, Jordi Gonzàlez 0001 |
ECCV (8) | 5 |
| 2018 | PRAXIS: Towards automatic cognitive assessment using gesture recognition
Farhood Negin, Pau Rodríguez, Michal Koperski, Adlen Kerboua, Jordi Gonzàlez 0001, Jeremy Bourgeois, Emmanuelle Chapoulie, Philippe Robert, François Brémond |
Expert Syst. Appl. | 5 |
| 2018 | Looking at People Special Issue
Sergio Escalera, Jordi Gonzàlez 0001, Hugo Jair Escalante, Xavier Baró, Isabelle Guyon |
Int. J. Comput. Vis. | 2 |
| 2018 | Top-down model fitting for hand pose recovery in sequences of depth images
Meysam Madadi, Sergio Escalera, Alex Carruesco, Carlos Andújar, Xavier Baró, Jordi Gonzàlez 0001 |
Image Vis. Comput. | 6 |
| 2018 | Beyond one-hot encoding: Lower dimensional target embedding
Pau Rodríguez, Miguel Ángel Bautista 0001, Jordi Gonzàlez 0001, Sergio Escalera |
Image Vis. Comput. | 3 |
| 2017 | Occlusion Aware Hand Pose Recovery from Sequences of Depth ImagesabstractState-of-the-art approaches on hand pose estimation from depth images have reported promising results under quite controlled considerations. In this paper we propose a two-step pipeline for recovering the hand pose from a sequence of depth images. The pipeline has been designed to deal with images taken from any viewpoint and exhibiting a high degree of finger occlusion. In a first step we initialize the hand pose using a part-based model, fitting a set of hand components in the depth images. In a second step we consider temporal data and estimate the parameters of a trained bilinear model consisting of shape and trajectory bases. Results on a synthetic, highly-occluded dataset demonstrate that the proposed method outperforms most recent pose recovering approaches, including those based on CNNs. Meysam Madadi, Sergio Escalera, Alex Carruesco, Carlos Andújar, Xavier Baró, Jordi Gonzàlez 0001 |
FG | 6 |
| 2017 | Regularizing CNNs with Locally Constrained Decorrelations
Pau Rodríguez, Jordi Gonzàlez 0001, Guillem Cucurull, Josep M. Gonfaus, F. Xavier Roca |
ICLR (Poster) | 2 |
| 2017 | Age and gender recognition in the wild with deep attention
Pau Rodríguez, Guillem Cucurull, Josep M. Gonfaus, F. Xavier Roca, Jordi Gonzàlez 0001 |
Pattern Recognit. | 5 |
| 2016 | Guest Editorial: Analysis and Retrieval of Events/Actions and Workflows in Video Streams
Anastasios Doulamis, Nikolaos D. Doulamis, Marco Bertini 0001, Jordi Gonzàlez 0001, Thomas B. Moeslund |
Multim. Tools Appl. | 4 |
| 2016 | Guest Editors' Introduction to the Special Issue on Multimodal Human Pose Recovery and Behavior AnalysisabstractThe sixteen papers in this special section focus on human pose recovery and behavior analysis (HuPBA). This is one of the most challenging topics in computer vision, pattern analysis, and machine learning. It is of critical importance for application areas that include gaming, computer interaction, human robot interaction, security, commerce, assistive technologies and rehabilitation, sports, sign language recognition, and driver assistance technology, to mention just a few. In essence, HuPBA requires dealing with the articulated nature of the human body, changes in appearance due to clothing, and the inherent problems of clutter scenes, such as background artifacts, occlusions, and illumination changes. These papers represent the most recent research in this field, including new methods considering still images, image sequences, depth data, stereo vision, 3D vision, audio, and IMUs, among others. Sergio Escalera, Jordi Gonzàlez 0001, Xavier Baró, Jamie Shotton |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Segmentation of RGB-D indoor scenes by stacking random forests and conditional random fields
Mikkel Thøgersen, Sergio Escalera, Jordi Gonzàlez 0001, Thomas B. Moeslund |
Pattern Recognit. Lett. | 3 |
| 2015 | ChaLearn looking at people 2015 new competitions: Age estimation and cultural event recognitionabstractFollowing previous series on Looking at People (LAP) challenges [1], [2], [3], in 2015 ChaLearn runs two new competitions within the field of Looking at People: age and cultural event recognition in still images. We propose the first crowd-sourcing application to collect and label data about apparent age of people instead of the real age. In terms of cultural event recognition, tens of categories have to be recognized. This involves scene understanding and human analysis. This paper summarizes both challenges and data, providing some initial baselines. The results of the first round of the competition were presented at ChaLearn LAP 2015 IJCNN special session on computer vision and robotics http://www.dtic.ua.es/~jgarcia/IJCNN2015. Details of the ChaLearn LAP competitions can be found at http://gesture.chalearn.org/. Sergio Escalera, Jordi Gonzàlez 0001, Xavier Baró, Pablo Pardo, Junior Fabian, Marc Oliu, Hugo Jair Escalante, Ivan Huerta Casado, Isabelle Guyon |
IJCNN | 2 |
| 2015 | Factorized appearances for object detection
Josep M. Gonfaus, Marco Pedersoli, Jordi Gonzàlez 0001, Andrea Vedaldi, F. Xavier Roca |
Comput. Vis. Image Underst. | 3 |
| 2015 | Chromatic shadow detection and tracking for moving foreground segmentation
Ivan Huerta Casado, Michael B. Holte, Thomas B. Moeslund, Jordi Gonzàlez 0001 |
Image Vis. Comput. | 4 |
| 2015 | Combining where and what in change detection for unsupervised foreground learning in surveillance
Ivan Huerta Casado, Marco Pedersoli, Jordi Gonzàlez 0001, Alberto Sanfeliu |
Pattern Recognit. | 3 |
| 2015 | A coarse-to-fine approach for fast deformable object detection
Marco Pedersoli, Andrea Vedaldi, Jordi Gonzàlez 0001, F. Xavier Roca |
Pattern Recognit. | 3 |
| 2015 | Multi-part body segmentation based on depth maps for soft biometry analysis
Meysam Madadi, Sergio Escalera, Jordi Gonzàlez 0001, F. Xavier Roca, Felipe Lumbreras |
Pattern Recognit. Lett. | 3 |
| 2014 | Corrigendum to "Hierarchical On-line Appearance-Based Tracking for 3D Head Pose, Eyebrows, Lips, Eyelids and Irises" [Image Vision Comput. (2013) 322-340]
Javier Orozco, Ognjen Rudovic, Jordi Gonzàlez 0001, Maja Pantic |
Image Vis. Comput. | 3 |
| 2014 | Special issue on background modeling for foreground detection in real-world dynamic scenes
Thierry Bouwmans, Jordi Gonzàlez 0001, Caifeng Shan, Massimo Piccardi, Larry Davis 0001 |
Mach. Vis. Appl. | 2 |
| 2014 | Spherical Blurred Shape Model for 3-D Object and Pose Recognition: Quantitative Analysis and HCI Applications in Smart EnvironmentsabstractThe use of depth maps is of increasing interest after the advent of cheap multisensor devices based on structured light, such as Kinect. In this context, there is a strong need of powerful 3-D shape descriptors able to generate rich object representations. Although several 3-D descriptors have been already proposed in the literature, the research of discriminative and computationally efficient descriptors is still an open issue. In this paper, we propose a novel point cloud descriptor called spherical blurred shape model (SBSM) that successfully encodes the structure density and local variabilities of an object based on shape voxel distances and a neighborhood propagation strategy. The proposed SBSM is proven to be rotation and scale invariant, robust to noise and occlusions, highly discriminative for multiple categories of complex objects like the human hand, and computationally efficient since the SBSM complexity is linear to the number of object voxels. Experimental evaluation in public depth multiclass object data, 3-D facial expressions data, and a novel hand poses data sets show significant performance improvements in relation to state-of-the-art approaches. Moreover, the effectiveness of the proposal is also proved for object spotting in 3-D scenes and for real-time automatic hand pose recognition in human computer interaction scenarios. Oscar Lopes, Miguel Reyes, Sergio Escalera, Jordi Gonzàlez 0001 |
IEEE Trans. Cybern. | 4 |
| 2014 | Color Constancy Using 3D Scene Geometry Derived From a Single ImageabstractThe aim of color constancy is to remove the effect of the color of the light source. As color constancy is inherently an ill-posed problem, most of the existing color constancy algorithms are based on specific imaging assumptions (e.g., gray-world and white patch assumption). In this paper, 3D geometry models are used to determine which color constancy method to use for the different geometrical regions (depth/layer) found in images. The aim is to classify images into stages (rough 3D geometry models). According to stage models, images are divided into stage regions using hard and soft segmentation. After that, the best color constancy methods are selected for each geometry depth. To this end, we propose a method to combine color constancy algorithms by investigating the relation between depth, local image statistics, and color constancy. Image statistics are then exploited per depth to select the proper color constancy method. Our approach opens the possibility to estimate multiple illuminations by distinguishing nearby light source from distant illuminations. Experiments on state-of-the-art data sets show that the proposed algorithm outperforms state-of-the-art single color constancy algorithms with an improvement of almost 50% of median angular error. When using a perfect classifier (i.e, all of the test images are correctly classified into stages); the performance of the proposed method achieves an improvement of 52% of the median angular error compared with the best-performing single color constancy algorithm. Noha M. Elfiky, Theo Gevers, Arjan Gijsenij, Jordi Gonzàlez 0001 |
IEEE Trans. Image Process. | 4 |
| 2014 | Toward Real-Time Pedestrian Detection Based on a Deformable Template ModelabstractMost advanced driving assistance systems already include pedestrian detection systems. Unfortunately, there is still a tradeoff between precision and real time. For a reliable detection, excellent precision-recall such a tradeoff is needed to detect as many pedestrians as possible while, at the same time, avoiding too many false alarms; in addition, a very fast computation is needed for fast reactions to dangerous situations. Recently, novel approaches based on deformable templates have been proposed since these show a reasonable detection performance although they are computationally too expensive for real-time performance. In this paper, we present a system for pedestrian detection based on a hierarchical multiresolution part-based model. The proposed system is able to achieve state-of-the-art detection accuracy due to the local deformations of the parts while exhibiting a speedup of more than one order of magnitude due to a fast coarse-to-fine inference technique. Moreover, our system explicitly infers the level of resolution available so that the detection of small examples is feasible with a very reduced computational cost. We conclude this contribution by presenting how a graphics processing unit-optimized implementation of our proposed system is suitable for real-time pedestrian detection in terms of both accuracy and speed. Marco Pedersoli, Jordi Gonzàlez 0001, F. Xavier Roca |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2013 | Multi-modal descriptors for multi-class hand pose recognition in human computer interaction systemsabstractHand pose recognition in advanced Human Computer Interaction systems (HCI) is becoming more feasible thanks to the use of affordable multi-modal RGB-Depth cameras. Depth data generated by these sensors is a very valuable input information, although the representation of 3D descriptors is still a critical step to obtain robust object representations. This paper presents an overview of different multi-modal descriptors, and provides a comparative study of two feature descriptors called Multi-modal Hand Shape (MHS) and Fourier-based Hand Shape (FHS), which compute local and global 2D-3D hand shape statistics to robustly describe hand poses. A new dataset of 38K hand poses has been created for real-time hand pose and gesture recognition, corresponding to five hand shape categories recorded from eight users. Experimental results show good performance of the fused MHS and FHS descriptors, improving recognition accuracy while assuring real-time computation in HCI scenarios. Jordi Abella, Raúl Alcaide, Anna Sabaté, Joan Mas Romeu, Sergio Escalera, Jordi Gonzàlez 0001, Coen Antens |
ICMI | 6 |
| 2013 | ChaLearn multi-modal gesture recognition 2013: grand challenge and workshop summaryabstractWe organized a Grand Challenge and Workshop on Multi-Modal Gesture Recognition. Sergio Escalera, Jordi Gonzàlez 0001, Xavier Baró, Miguel Reyes, Isabelle Guyon, Vassilis Athitsos, Hugo Jair Escalante, Leonid Sigal, Antonis A. Argyros, Cristian Sminchisescu, Richard Bowden, Stan Sclaroff |
ICMI | 2 |
| 2013 | Multi-modal gesture recognition challenge 2013: dataset and resultsabstractThe recognition of continuous natural gestures is a complex and challenging problem due to the multi-modal nature of involved visual cues (e.g. fingers and lips movements, subtle facial expressions, body pose, etc.), as well as technical limitations such as spatial and temporal resolution and unreliable depth cues. In order to promote the research advance on this field, we organized a challenge on multi-modal gesture recognition. We made available a large video database of 13,858 gestures from a lexicon of 20 Italian gesture categories recorded with a Kinect™ camera, providing the audio, skeletal model, user mask, RGB and depth images. The focus of the challenge was on user independent multiple gesture learning. There are no resting positions and the gestures are performed in continuous sequences lasting 1-2 minutes, containing between 8 and 20 gesture instances in each sequence. As a result, the dataset contains around 1.720.800 frames. In addition to the 20 main gesture categories, "distracter" gestures are included, meaning that additional audio and gestures out of the vocabulary are included. The final evaluation of the challenge was defined in terms of the Levenshtein edit distance, where the goal was to indicate the real order of gestures within the sequence. 54 international teams participated in the challenge, and outstanding results were obtained by the first ranked participants. Sergio Escalera, Jordi Gonzàlez 0001, Xavier Baró, Miguel Reyes, Oscar Lopes, Isabelle Guyon, Vassilis Athitsos, Hugo Jair Escalante |
ICMI | 2 |
| 2013 | 4th ACM/IEEE ARTEMIS 2013 international workshop on analysis and retrieval of tracked events and motion in imagery streamsabstractIn this paper, we give a short summary of the papers proposed in ACMARTEMIS 2013 which is held in Barcelona Spain in conjunction with ACM Multimedia. The workshop handles the areas of features analysis both at low and high level for efficient events detection, retrieval of multimedia events and objects and video synchronization issues and also events and behavior recognition from visual data. All papers were classified into three session of a single track workshop. The first session named "Video Features and Scene Analysis" includes articles that handle low level and high level visual analysis appropriate for event detection. The second session entitled "Retrieval of Multimedia Objects/Events" applies schemes for media data retrieval and video synchronization. Finally the third session "Analysis of Visual Events" describes algorithms for detecting actions, behaviors and events in complex visual scenes. Anastasios Doulamis, Nikolaos D. Doulamis, Marco Bertini 0001, Jordi Gonzàlez 0001, Thomas B. Moeslund |
ACM Multimedia | 4 |
| 2013 | Large scale continuous visual event recognition using max-margin Hough transformation framework
Bhaskar Chakraborty, Jordi Gonzàlez 0001, F. Xavier Roca |
Comput. Vis. Image Underst. | 2 |
| 2013 | Human action recognition using an ensemble of body-part detectorsabstractAbstract This paper describes an approach to human action recognition based on a probabilistic optimization model of body parts using hidden Markov model (HMM). Our method is able to distinguish between similar actions by only considering the body parts having major contribution to the actions, for example, legs for walking, jogging and running; arms for boxing, waving and clapping. We apply HMMs to model the stochastic movement of the body parts for action recognition. The HMM construction uses an ensemble of body‐part detectors, followed by grouping of part detections, to perform human identification. Three example‐based body‐part detectors are trained to detect three components of the human body: the head, legs and arms. These detectors cope with viewpoint changes and self‐occlusions through the use of ten sub‐classifiers that detect body parts over a specific range of viewpoints. Each sub‐classifier is a support vector machine trained on features selected for the discriminative power for each particular part/viewpoint combination. Grouping of these detections is performed using a simple geometric constraint model that yields a viewpoint‐invariant human detector. We test our approach on three publicly available action datasets: the KTH dataset, Weizmann dataset and HumanEva dataset. Our results illustrate that with a simple and compact representation we can achieve robust recognition of human actions comparable to the most complex, state‐of‐the‐art methods. Bhaskar Chakraborty, Andrew D. Bagdanov, Jordi Gonzàlez 0001, F. Xavier Roca |
Expert Syst. J. Knowl. Eng. | 3 |
| 2013 | Exploiting multiple cues in motion segmentation based on background subtraction
Ivan Huerta Casado, Ariel Amato, F. Xavier Roca, Jordi Gonzàlez 0001 |
Neurocomputing | 4 |
| 2013 | Hierarchical On-line Appearance-Based Tracking for 3D head pose, eyebrows, lips, eyelids and irises
Javier Orozco, Ognjen Rudovic, Jordi Gonzàlez 0001, Maja Pantic |
Image Vis. Comput. | 3 |
| 2012 | On partial least squares in head pose estimation: How to simultaneously deal with misalignmentabstractHead pose estimation is a critical problem in many computer vision applications. These include human computer interaction, video surveillance, face and expression recognition. In most prior work on heads pose estimation, the positions of the faces on which the pose is to be estimated are specified manually. Therefore, the results are reported without studying the effect of misalignment. We propose a method based on partial least squares (PLS) regression to estimate pose and solve the alignment problem simultaneously. The contributions of this paper are two-fold: 1) we show that the kernel version of PLS (kPLS) achieves better than state-of-the-art results on the estimation problem and 2) we develop a technique to reduce misalignment based on the learned PLS factors. Murad Al Haj, Jordi Gonzàlez 0001, Larry Davis 0001 |
CVPR | 2 |
| 2012 | 3D human pose estimation using 2D body part detectors
Adela Barbulescu, Wenjuan Gong, Jordi Gonzàlez 0001, Thomas B. Moeslund, F. Xavier Roca |
ICPR | 3 |
| 2012 | Edge classification using photo-geometric features
Josep M. Gonfaus, Theo Gevers, Arjan Gijsenij, F. Xavier Roca, Jordi Gonzàlez 0001 |
ICPR | 5 |
| 2012 | Selective spatio-temporal interest points
Bhaskar Chakraborty, Michael B. Holte, Thomas B. Moeslund, Jordi Gonzàlez 0001 |
Comput. Vis. Image Underst. | 4 |
| 2012 | Semantic Understanding of Human Behaviors in Image Sequences: From video-surveillance to video-hermeneutics
Jordi Gonzàlez 0001, Thomas B. Moeslund, Liang Wang 0001 |
Comput. Vis. Image Underst. | 1 |
| 2012 | Harmony Potentials - Fusing Global and Local Scale for Semantic Image Segmentation
Xavier Boix, Josep M. Gonfaus, Joost van de Weijer 0001, Andrew D. Bagdanov, Joan Serrat 0002, Jordi Gonzàlez 0001 |
Int. J. Comput. Vis. | 6 |
| 2012 | Compact and adaptive spatial pyramids for scene recognition
Noha M. Elfiky, Jordi Gonzàlez 0001, F. Xavier Roca |
Image Vis. Comput. | 2 |
| 2012 | Discriminative compact pyramids for object and scene recognition
Noha M. Elfiky, Fahad Shahbaz Khan, Joost van de Weijer 0001, Jordi Gonzàlez 0001 |
Pattern Recognit. | 4 |
| 2011 | A coarse-to-fine approach for fast deformable object detectionabstractWe present a method that can dramatically accelerate object detection with part based models. The method is based on the observation that the cost of detection is likely to be dominated by the cost of matching each part to the image, and not by the cost of computing the optimal configuration of the parts as commonly assumed. Therefore accelerating detection requires minimizing the number of part-to-image comparisons. To this end we propose a multiple-resolutions hierarchical part based model and a corresponding coarse-to-fine inference procedure that recursively eliminates from the search space unpromising part placements. The method yields a ten-fold speedup over the standard dynamic programming approach and is complementary to the cascade-of-parts approach of. Compared to the latter, our method does not have parameters to be determined empirically, which simplifies its use during the training of the model. Most importantly, the two techniques can be combined to obtain a very significant speedup, of two orders of magnitude in some cases. We evaluate our method extensively on the PASCAL VOC and INRIA datasets, demonstrating a very high increase in the detection speed with little degradation of the accuracy. Marco Pedersoli, Andrea Vedaldi, Jordi Gonzàlez 0001 |
CVPR | 3 |
| 2011 | A selective spatio-temporal interest point detector for human action recognition in complex scenesabstractRecent progress in the field of human action recognition points towards the use of Spatio-Temporal Interest Points (STIPs) for local descriptor-based recognition strategies. In this paper we present a new approach for STIP detection by applying surround suppression combined with local and temporal constraints. Our method is significantly different from existing STIP detectors and improves the performance by detecting more repeatable, stable and distinctive STIPs for human actors, while suppressing unwanted background STIPs. For action representation we use a bag-of-visual words (BoV) model of local N-jet features to build a vocabulary of visual-words. To this end, we introduce a novel vocabulary building strategy by combining spatial pyramid and vocabulary compression techniques, resulting in improved performance and efficiency. Action class specific Support Vector Machine (SVM) classifiers are trained for categorization of human actions. A comprehensive set of experiments on existing benchmark datasets, and more challenging datasets of complex scenes, validate our approach and show state-of-the-art performance. Bhaskar Chakraborty, Michael B. Holte, Thomas B. Moeslund, Jordi Gonzàlez 0001, F. Xavier Roca |
ICCV | 4 |
| 2011 | Determining the best suited semantic events for cognitive surveillance
Carles Fernández, Pau Baiget, F. Xavier Roca, Jordi Gonzàlez 0001 |
Expert Syst. Appl. | 4 |
| 2011 | Augmenting video surveillance footage with virtual agents for incremental event evaluation
Carles Fernández, Pau Baiget, F. Xavier Roca, Jordi Gonzàlez 0001 |
Pattern Recognit. Lett. | 4 |
| 2011 | Efficient discriminative multiresolution cascade for real-time human detection applications
Marco Pedersoli, Jordi Gonzàlez 0001, Andrew D. Bagdanov, F. Xavier Roca |
Pattern Recognit. Lett. | 2 |
| 2011 | Accurate Moving Cast Shadow Suppression Based on Local Color Constancy DetectionabstractThis paper describes a novel framework for detection and suppression of properly shadowed regions for most possible scenarios occurring in real video sequences. Our approach requires no prior knowledge about the scene, nor is it restricted to specific scene structures. Furthermore, the technique can detect both achromatic and chromatic shadows even in the presence of camouflage that occurs when foreground regions are very similar in color to shadowed regions. The method exploits local color constancy properties due to reflectance suppression over shadowed regions. To detect shadowed regions in a scene, the values of the background image are divided by values of the current frame in the RGB color space. We show how this luminance ratio can be used to identify segments with low gradient constancy, which in turn distinguish shadows from foreground. Experimental results on a collection of publicly available datasets illustrate the superior performance of our method compared with the most sophisticated, state-of-the-art shadow detection algorithms. These results show that our approach is robust and accurate over a broad range of shadow types and challenging video conditions. Ariel Amato, Mikhail G. Mozerov, Andrew D. Bagdanov, Jordi Gonzàlez 0001 |
IEEE Trans. Image Process. | 4 |
| 2010 | Harmony potentials for joint classification and segmentationabstractHierarchical conditional random fields have been successfully applied to object segmentation. One reason is their ability to incorporate contextual information at different scales. However, these models do not allow multiple labels to be assigned to a single node. At higher scales in the image, this yields an oversimplified model, since multiple classes can be reasonable expected to appear within one region. This simplified model especially limits the impact that observations at larger scales may have on the CRF model. Neglecting the information at larger scales is undesirable since class-label estimates based on these scales are more reliable than at smaller, noisier scales. To address this problem, we propose a new potential, called harmony potential, which can encode any possible combination of class labels. We propose an effective sampling strategy that renders tractable the underlying optimization problem. Results show that our approach obtains state-of-the-art results on two challenging datasets: Pascal VOC 2009 and MSRC-21. Josep M. Gonfaus, Xavier Boix, Joost van de Weijer 0001, Andrew D. Bagdanov, Joan Serrat 0002, Jordi Gonzàlez 0001 |
CVPR | 6 |
| 2010 | Automatic Learning of Background Semantics in Generic Surveilled Scenes
Carles Fernández, Jordi Gonzàlez 0001, F. Xavier Roca |
ECCV (2) | 2 |
| 2010 | Recursive Coarse-to-Fine Localization for Fast Object Detection
Marco Pedersoli, Jordi Gonzàlez 0001, Andrew D. Bagdanov, Juan José Villanueva |
ECCV (6) | 2 |
| 2010 | Reactive Object Tracking with a Single PTZ CameraabstractIn this paper we describe a novel approach to reactive tracking of moving targets with a pan-tilt-zoom camera. The approach uses an extended Kalman filter to jointly track the object position in the real world, its velocity in 3D and the camera intrinsics, in addition to the rate of change of these parameters. The filter outputs are used as inputs to PID controllers which continuously adjust the camera motion in order to reactively track the object at a constant image velocity while simultaneously maintaining a desirable target scale in the image plane. We provide experimental results on simulated and real tracking sequences to show how our tracker is able to accurately estimate both 3D object position and camera intrinsics with very high precision over a wide range of focal lengths. Murad Al Haj, Andrew D. Bagdanov, Jordi Gonzàlez 0001, F. Xavier Roca |
ICPR | 3 |
| 2010 | First ACM international workshop on analysis and retrieval of tracked events and motion in imagery streams (ARTEMI 2010)abstractThe advancement of novel capabilities for video understanding does increase the cross-fertilization between multiple computer vision and pattern recognition research topics. ARTEMIS2010 provides the forum for discussing a holistic view on the interpretation and description of human behaviors in multimedia content such as sports, news, documentaries, movies and surveillance footage. Anastasios Doulamis, Jordi Gonzàlez 0001 |
ACM Multimedia | 2 |
| 2010 | On tracking inside groups
Daniel Rowe, Jordi Gonzàlez 0001, Marco Pedersoli, Juan José Villanueva |
Mach. Vis. Appl. | 2 |
| 2009 | Detection and removal of chromatic moving shadows in surveillance scenariosabstractSegmentation in the surveillance domain has to deal with shadows to avoid distortions when detecting moving objects. Most segmentation approaches dealing with shadow detection are typically restricted to penumbra shadows. Therefore, such techniques cannot cope well with umbra shadows. Consequently, umbra shadows are usually detected as part of moving objects. In this paper we present a novel technique based on gradient and colour models for separating chromatic moving cast shadows from detected moving objects. Firstly, both a chromatic invariant colour cone model and an invariant gradient model are built to perform automatic segmentation while detecting potential shadows. In a second step, regions corresponding to potential shadows are grouped by considering “a bluish effect” and an edge partitioning. Lastly, (i) temporal similarities between textures and (ii) spatial similarities between chrominance angle and brightness distortions are analysed for all potential shadow regions in order to finally identify umbra shadows. Unlike other approaches, our method does not make any a-priori assumptions about camera location, surface geometries, surface textures, shapes and types of shadows, objects, and background. Experimental results show the performance and accuracy of our approach in different shadowed materials and illumination conditions. Ivan Huerta Casado, Michael B. Holte, Thomas B. Moeslund, Jordi Gonzàlez 0001 |
ICCV | 4 |
| 2009 | Trinocular stereo matching with composite disparity space imageabstractIn this paper we propose a method that smartly improves occlusion handling in stereo matching using trinocular stereo. The main idea is based on the assumption that any occluded region in a matched stereo pair (middle-left images) in general is not occluded in the opposite matched pair (middle-right images). Then two disparity space images (DSI) are merged in one composite DSI. The proposed integration differs from the known approach that uses a cumulative cost. The experimental results are evaluated on the Middlebury data set, showing high performance of the proposed algorithm especially in the occluded regions. Our method solves the problem on the base of a real matching cost, in such a way a global optimization problem is solved just once, and the resultant solution does not have to be corrected in the occluded regions. In contrast, the traditional methods that use two images approach have to complicate a lot their algorithms by additional add hog or heuristic techniques to reach competitive results in occluded regions. Mikhail G. Mozerov, Jordi Gonzàlez 0001, F. Xavier Roca, Juan José Villanueva |
ICIP | 2 |
| 2009 | Editorial
Liang Wang 0001, Qiang Wu 0001, Ming Li 0010, Jordi Gonzàlez 0001, Xin Geng 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2009 | Understanding dynamic scenes based on human sequence evaluation
Jordi Gonzàlez 0001, Daniel Rowe, Javier Varona, F. Xavier Roca |
Image Vis. Comput. | 1 |
| 2009 | Toward natural interaction through visual recognition of body gestures in real-timeabstractIn most of the existing human–computer interfaces, enactive knowledge as new natural interaction paradigm has not been fully exploited yet. Recent technological advances have created the possibility to enhance naturally and significantly the interface perception by means of visual inputs, the so-called Vision-Based Interfaces (VBI). In the present paper, we explore the recovery of the user’s body posture by means of combining robust computer vision techniques and a well known inverse kinematics algorithm in real-time. Specifically, we focus on recognizing the user’s motions with a particular mean, that is, a body gesture. Defining an appropriate representation of the user’s body posture based on a temporal parameterization, we apply non-parametric techniques to learn and recognize the user’s body gestures. This scheme of recognition has been applied to control a computer videogame in real-time to show the viability of the presented approach. Javier Varona, Antoni Jaume-i-Capó, Jordi Gonzàlez 0001, Francisco José Perales López |
Interact. Comput. | 3 |
| 2009 | Generation of augmented video sequences combining behavioral animation and multi-object trackingabstractAbstract In this paper we present a novel approach to generate augmented video sequences in real‐time, involving interactions between virtual and real agents in real scenarios. On the one hand, real agent motion is estimated by means of a multi‐object tracking algorithm, which determines real objects' position over the scenario for each time step. On the other hand, virtual agents are provided with behavior models considering their interaction with the environment and with other agents. The resulting framework allows to generate video sequences involving behavior‐based virtual agents that react to real agent behavior and has applications in education, simulation, and in the game and movie industries. We show the performance of the proposed approach in an indoor and outdoor scenario simulating human and vehicle agents. Copyright © 2009 John Wiley & Sons, Ltd. Pau Baiget, Carles Fernández, F. Xavier Roca, Jordi Gonzàlez 0001 |
Comput. Animat. Virtual Worlds | 4 |
| 2009 | Real-time gaze tracking with appearance-based models
Javier Orozco, F. Xavier Roca, Jordi Gonzàlez 0001 |
Mach. Vis. Appl. | 3 |
| 2009 | Action-specific motion prior for efficient Bayesian 3D human body tracking
Ignasi Rius, Jordi Gonzàlez 0001, Javier Varona, F. Xavier Roca |
Pattern Recognit. | 2 |
| 2008 | View-invariant human-body detection with extension to human action recognition using component-wise HMM of body partsabstractThis paper presents a technique for view invariant human detection and extending this idea to recognize basic human actions like walking, jogging, hand waving and boxing etc. To achieve this goal we detect the human in its body parts and then learn the changes of those body parts for action recognition. Human-body part detection in different views is an extremely challenging problem due to drastic change of 3D-pose of human body, self occlusions etc while performing actions. In order to cope with these problems we have designed three example-based detectors that are trained to find separately three components of the human body, namely the head, legs and arms. We incorporate 10 sub-classifiers for the head, arms and the leg detection. Each sub-classifier detects body parts under a specific range of viewpoints. Then, view-invariance is fulfilled by combining the results of these sub classifiers. Subsequently, we extend this approach to recognize actions based on component-wise hidden Markov models (HMM). This is achieved by designing a HMM for each action, which is trained based on the detected body parts. Consequently, we are able to distinguish between similar actions by only considering the body parts which has major contributions to those actions e.g. legs for walking, running etc; hands for boxing, waving etc. Bhaskar Chakraborty, Ognjen Rudovic, Jordi Gonzàlez 0001 |
FG | 3 |
| 2008 | Confidence assessment on eyelid and eyebrow expression recognitionabstractIn this paper, we address the recognition of subtle facial expressions by reasoning on the classification confidence. Psychological evidences have determined that eyelids and eyebrows are significant for the recognition of subtle facial expressions and the early perception of human emotions. This early perception results in a more complex problem, which requires a confidence assessment for any provided solution. Thus, traditional score-based classifiers (e.g. k-NN and NN) are not able to produce confident estimates. Instead, we first present five confidence estimators and a confidence classification assessment for Case-Based Reasoning (CBR). Second, we improve the expression retrieval from the database by learning the neighbourhood's dimensions for the expected classification confidences. Third, we reuse the previous classified expressions and the confidence assessment to improve the classification achieved by k-NN. Fourth, we improve the database for generalization with new subjects by learning thresholds to minimize misclassification with low confidence, maximize correct classifications with high confidence and re-arrange misclassification with high confidence. The proposed system represents an effective contribution for both subtle expression recognition and CBR methodology. It achieves an average recognition of 97% plusmn 1% with a confidence of 96% plusmn 2% for expressiveness between 20% and 100%. Javier Orozco, Ognjen Rudovic, F. Xavier Roca, Jordi Gonzàlez 0001 |
FG | 4 |
| 2008 | Background subtraction technique based on chromaticity and intensity patternsabstractThis paper presents an efficient real-time method for detecting moving objects in unconstrained environments, using a background subtraction technique. A new background model that combines spatial and temporal information based on similarity measure in angles and intensity between two color vectors is introduced. The comparison is done in RGB color space. A new feature based on chromaticity and intensity pattern is extracted in order to improve the accuracy in the ambiguity region where there is a strong similarity between background and foreground and to cope with cast shadows. The effectiveness of the proposed method is demonstrated in the experimental results and comparison with others approaches is also shown. Ariel Amato, Mikhail G. Mozerov, Ivan Huerta Casado, Jordi Gonzàlez 0001, Juan José Villanueva |
ICPR | 4 |
| 2008 | Automatic face and facial features initialization for robust and accurate trackingabstractFace detection and tracking, through image sequences, are primary steps in many applications such as video surveillance, human computer interface, and expression analysis. Many currently existing techniques donpsilat perform well due to pose variations, appearance changes, illumination changes, complex backgrounds, and inaccurate initialization. The last short coming, which is the difficulty to initialize motion regions, is a problem facing any tracker. In this paper, we present an automatic and robust face detection and tracking system for color image sequences. Face detection is done using skin color segmentation and connected components analysis. Later, facial features are detected by active shape models and a face mesh is initialized. Finally, the tracking is done by active appearance models. Experimental detection and tracking results on a pose varying face video are given. Murad Al Haj, Javier Orozco, Jordi Gonzàlez 0001, Juan José Villanueva |
ICPR | 3 |
| 2008 | Interpretation of complex situations in a semantic-based surveillance framework
Carles Fernández, Pau Baiget, F. Xavier Roca, Jordi Gonzàlez 0001 |
Signal Process. Image Commun. | 4 |
| 2007 | Deterministic and Stochastic Methods for Gaze Tracking in Real-Time
Javier Orozco, F. Xavier Roca, Jordi Gonzàlez 0001 |
CAIP | 3 |
| 2000 | iTrack: Image-Based Probabilistic Tracking of PeopleabstractReal applications on people tracking are usually based on image heuristics. Real approaches do not use to apply recent prediction-estimation theoretical frameworks. These require the definition of complex dynamical and shape object models before the tracking process. We present a probablistic framework that takes profit of these theories adapting them to real applications. The key idea of this work is to estimate the shape model and dynamical objects parameters using only image data. The flexibility of our algorithm makes it suitable to be used on different real applications. Some experiments have been done in order to test our method in outdoor scenes people tracking. Javier Varona, Jordi Gonzàlez 0001, F. Xavier Roca, Juan José Villanueva |
ICPR | 2 |