Abel Gonzalez-Garcia

dblp:156/0213 · also Abel González-García · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
3since 2021 · last 2024
0000-0002-9334-6646ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Video understanding and tracking · 20% Generative modeling · 18% Image recognition and object detection · 15%
Computer graphics and multimedia
3 papers
Visual content generation and editing · 72% Multimedia analysis and retrieval · 28%

Topics — the 22 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing
image-to-image translation
1.132020
Semi-Supervised Learning for Few-Shot Image-to-Image Translation · CVPR 2020
SDIT: Scalable and Diverse Cross-domain Image Translation · ACM Multimedia 2019
Image-to-image translation for cross-domain disentanglement · NeurIPS 2018
Computer vision › Video understanding and tracking › object tracking
thermal infrared tracking
0.922021
Unsupervised Cross-Modal Distillation for Thermal Infrared Tracking · ACM Multimedia 2021
Synthetic Data Generation for End-to-End Thermal Infrared Tracking · IEEE Trans. Image Process. 2019
Machine learning › Generative modeling
generative adversarial network
0.932020
MineGAN: Effective Knowledge Transfer From GANs to Target Domains With Few Images · CVPR 2020
Image-to-image translation for cross-domain disentanglement · NeurIPS 2018
Transferring GANs: Generating Images from Limited Data · ECCV (6) 2018
Computer vision › Video understanding and tracking
object tracking
0.822019
Synthetic Data Generation for End-to-End Thermal Infrared Tracking · IEEE Trans. Image Process. 2019
Learning the Model Update for Siamese Trackers · ICCV 2019
Machine learning › Transfer learning and domain adaptation
knowledge transfer
0.722024
MineGAN: Effective Knowledge Transfer From GANs to Target Domains With Few Images · CVPR 2020
MineGAN++: Mining Generative Models for Efficient Knowledge Transfer to Limited Data Domains · Int. J. Comput. Vis. 2024
Computer vision › Image recognition and object detection
object detection
0.622019
Active Learning for Deep Detection Neural Networks · ICCV 2019
An active search strategy for efficient object class detection · CVPR 2015
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
cross-modal distillation
0.512021
Unsupervised Cross-Modal Distillation for Thermal Infrared Tracking · ACM Multimedia 2021
Machine learning › Efficient and distributed learning › distillation
unsupervised distillation
0.512021
Unsupervised Cross-Modal Distillation for Thermal Infrared Tracking · ACM Multimedia 2021
Machine learning › Deep learning architectures and training › recurrent neural network
LSTM
0.412020
Orderless Recurrent Models for Multi-Label Classification · CVPR 2020
Machine learning › Learning paradigms
multi-label classification
0.412020
Orderless Recurrent Models for Multi-Label Classification · CVPR 2020
Machine learning › Deep learning architectures and training
recurrent neural network
0.412020
Orderless Recurrent Models for Multi-Label Classification · CVPR 2020
Multimedia analysis and retrieval
semi-supervised learning
0.412020
Semi-Supervised Learning for Few-Shot Image-to-Image Translation · CVPR 2020
Computer vision › Image recognition and object detection › object detection › detector training
active learning for object detection
0.412019
Active Learning for Deep Detection Neural Networks · ICCV 2019
Computer vision › Video understanding and tracking › object tracking › deep tracking
siamese tracking
0.412019
Learning the Model Update for Siamese Trackers · ICCV 2019
Machine learning › Representation and self-supervised learning › representation learning › disentangled representation learning
cross-domain disentanglement
0.312018
Image-to-image translation for cross-domain disentanglement · NeurIPS 2018
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.312018
Image-to-image translation for cross-domain disentanglement · NeurIPS 2018
Machine learning › Generative modeling › generative adversarial network
image-to-image translation
0.312018
Image-to-image translation for cross-domain disentanglement · NeurIPS 2018
Machine learning › Trustworthy machine learning › interpretability › explainable AI
interpretable convolutional neural network
0.312018
Do Semantic Parts Emerge in Convolutional Neural Networks? · Int. J. Comput. Vis. 2018
Computer vision › Image recognition and object detection › object detection
contextual reasoning
0.212015
An active search strategy for efficient object class detection · CVPR 2015
Computer vision › Video understanding and tracking › object tracking
target representation
0.112021
Unsupervised Cross-Modal Distillation for Thermal Infrared Tracking · ACM Multimedia 2021
Machine learning › Efficient and distributed learning
active learning
0.112019
Active Learning for Deep Detection Neural Networks · ICCV 2019
Machine learning › Efficient and distributed learning
data-efficient learning
0.112019
Active Learning for Deep Detection Neural Networks · ICCV 2019

Methods — techniques the papers use, named apart from their topics

convolutional neural network · 2.0generative model mining · 0.8unsupervised cross-modal distillation · 0.5pseudo-labeling · 0.4latent space mining · 0.4fine-tuning · 0.4dynamic label ordering · 0.4cycle consistency · 0.4CNN-RNN · 0.4updatenet · 0.4temporal selection · 0.4latent variable modeling · 0.4image-level scoring · 0.4attention mechanism · 0.4generative adversarial network · 0.3cross-domain autoencoder · 0.3
YearPublicationVenuePosition
2024 MineGAN++: Mining Generative Models for Efficient Knowledge Transfer to Limited Data Domains
Yaxing Wang, Abel Gonzalez-Garcia, Chenshen Wu, Luis Herranz, Fahad Shahbaz Khan, Shangling Jui, Jian Yang 0003, Joost van de Weijer 0001
Int. J. Comput. Vis.2
2021 Unsupervised Cross-Modal Distillation for Thermal Infrared Tracking
abstract
The target representation learned by convolutional neural networks plays an important role in Thermal Infrared (TIR) tracking. Currently, most of the top-performing TIR trackers are still employing representations learned by the model trained on the RGB data. However, this representation does not take into account the information in the TIR modality itself, limiting the performance of TIR tracking.
Jingxian Sun 0003, Lichao Zhang 0001, Yufei Zha, Abel Gonzalez-Garcia, Peng Zhang 0005, Wei Huang 0013, Yanning Zhang 0001
ACM Multimedia4
2021 Controlling biases and diversity in diverse image-to-image translation
Yaxing Wang, Abel Gonzalez-Garcia, Luis Herranz, Joost van de Weijer 0001
Comput. Vis. Image Underst.2
2020 MineGAN: Effective Knowledge Transfer From GANs to Target Domains With Few Images
abstract
One of the attractive characteristics of deep neural networks is their ability to transfer knowledge obtained in one domain to other related domains. As a result, high-quality networks can be trained in domains with relatively little training data. This property has been extensively studied for discriminative networks but has received significantly less attention for generative models. Given the often enormous effort required to train GANs, both computationally as well as in the dataset collection, the re-use of pretrained GANs is a desirable objective. We propose a novel knowledge transfer method for generative models based on mining the knowledge that is most beneficial to a specific target domain, either from a single or multiple pretrained GANs. This is done using a miner network that identifies which part of the generative distribution of each pretrained GAN outputs samples closest to the target domain. Mining effectively steers GAN sampling towards suitable regions of the latent space, which facilitates the posterior finetuning and avoids pathologies of other methods such as mode collapse and lack of flexibility. We perform experiments on several complex datasets using various GAN architectures (BigGAN, Progressive GAN) and show that the proposed method, called MineGAN, effectively transfers knowledge to domains with few target images, outperforming existing methods. In addition, MineGAN can successfully transfer knowledge from multiple pretrained GANs. Our code is available at: \url{https://github.com/yaxingwang/MineGAN}.
Yaxing Wang, Abel Gonzalez-Garcia, David Berga, Luis Herranz, Fahad Shahbaz Khan, Joost van de Weijer 0001
CVPR2
2020 Semi-Supervised Learning for Few-Shot Image-to-Image Translation
abstract
In the last few years, unpaired image-to-image translation has witnessed Remarkable progress. Although the latest methods are able to generate realistic images, they crucially rely on a large number of labeled images. Recently, some methods have tackled the challenging setting of few-shot image-to-image ranslation, reducing the labeled data requirements for the target domain during inference. In this work, we go one step further and reduce the amount of required labeled data also from the source domain during training. To do so, we propose applying semi-supervised learning via a noise-tolerant pseudo-labeling procedure. We also apply a cycle consistency constraint to further exploit the information from unlabeled images, either from the same dataset or external. Additionally, we propose several structural modifications to facilitate the image translation task under these circumstances. Our semi-supervised method for few-shot image translation, called \emph{SEMIT}, achieves excellent results on four different datasets using as little as 10\% of the source labels, and matches the performance of the main fully-supervised competitor using only 20\% labeled data. Our code and models are made public at: \url{https://github.com/yaxingwang/SEMIT}.
Yaxing Wang, Salman Khan 0001, Abel Gonzalez-Garcia, Joost van de Weijer 0001, Fahad Shahbaz Khan
CVPR3
2020 Orderless Recurrent Models for Multi-Label Classification
abstract
Recurrent neural networks (RNN) are popular for many computer vision tasks, including multi-label classification. Since RNNs produce sequential outputs, labels need to be ordered for the multi-label classification task. Current approaches sort labels according to their frequency, typically ordering them in either rare-first or frequent-first. These imposed orderings do not take into account that the natural order to generate the labels can change for each image, e.g. first the dominant object before summing up the smaller objects in the image. Therefore, in this paper, we propose ways to dynamically order the ground truth labels with the predicted label sequence. This allows for the faster training of more optimal LSTM models for multi-label classification. Analysis evidences that our method does not suffer from duplicate generation, something which is common for other models. Furthermore, it outperforms other CNN-RNN models, and we show that a standard architecture of an image encoder and language decoder trained with our proposed loss obtains the state-of-the-art results on the challenging MS-COCO, WIDER Attribute and PA-100K and competitive results on NUS-WIDE.
Vacit Oguz Yazici, Abel Gonzalez-Garcia, Arnau Ramisa, Bartlomiej Twardowski, Joost van de Weijer 0001
CVPR2
2020 Validation of instruments to measure social entrepreneurship competence. The OpenSocialLab project
abstract
Education within universities should consider the promotion of training activities aimed at training people who are creative, innovative, enterprising and aware of their environment and needs. The purpose of this paper is to present the preliminary results of the piloting of three instruments for a methodological proposal aimed at measuring the level of mastery scaled by students of undergraduate and graduate courses, in terms of social entrepreneurship skills. Thus, the instruments were validated through various strategies such as expert judgement, non-participating observation, statistical validity and reliability. Furthermore, the piloting takes place within the framework of a mixed method, since the data collection instruments were the focus group, the questionnaire and the semi-structured interview. The sample participating in the validation was different, depending on the instrument piloted: focus group (n =5), questionnaire (n =98) and interview (n =4). Finally, the contributions of this work can be of value in studying social entrepreneurship.
Abel Gonzalez-Garcia, Luis M. Romero-Rodriguez, José-María Romero-Rodríguez, María Soledad Ramírez-Montoya
EDUCON1
2020 Emerging technologies for the proposal and design of a MOOC on social entrepreneurship
abstract
Emerging technologies such as virtual reality and mixed reality are an educational resource that attracts students' attention and increases their interest in the content presented. The application of this technology to a MOOC course makes the proposal to motivate students really interesting. This paper presents the proposal and design of a MOOC course on social entrepreneurship that integrates virtual reality and open educational resources. Furthermore, the methodology used in the course is based on the Learning Environment Modeling Language (LEML). This makes it really attractive to the public, while improving the traditional content of a training course. Finally, the main findings and implications of the training course on competence in social entrepreneurship and its relation to the development of Sustainable Development Objectives are discussed.
María Soledad Ramírez-Montoya, José-Guadalupe González-Padrón, Marlene Muzquiz-Flores, Abel Gonzalez-Garcia, José-María Romero-Rodríguez, Inmaculada Aznar-Díaz
EDUCON4
2019 Active Learning for Deep Detection Neural Networks
abstract
The cost of drawing object bounding boxes (i.e. labeling) for millions of images is prohibitively high. For instance, labeling pedestrians in a regular urban image could take 35 seconds on average. Active learning aims to reduce the cost of labeling by selecting only those images that are informative to improve the detection network accuracy. In this paper, we propose a method to perform active learning of object detectors based on convolutional neural networks. We propose a new image-level scoring process to rank unlabeled images for their automatic selection, which clearly outperforms classical scores. The proposed method can be applied to videos and sets of still images. In the former case, temporal selection rules can complement our scoring process. As a relevant use case, we extensively study the performance of our method on the task of pedestrian detection. Overall, the experiments show that the proposed method performs better than random selection.
Hamed H. Aghdam, Abel Gonzalez-Garcia, Antonio M. López 0001, Joost van de Weijer 0001
ICCV2
2019 Learning the Model Update for Siamese Trackers
abstract
Siamese approaches address the visual tracking problem by extracting an appearance template from the current frame, which is used to localize the target in the next frame. In general, this template is linearly combined with the accumulated template from the previous frame, resulting in an exponential decay of information over time. While such an approach to updating has led to improved results, its simplicity limits the potential gain likely to be obtained by learning to update. Therefore, we propose to replace the handcrafted update function with a method which learns to update. We use a convolutional neural network, called UpdateNet, which given the initial template, the accumulated template and the template of the current frame aims to estimate the optimal template for the next frame. The UpdateNet is compact and can easily be integrated into existing Siamese trackers. We demonstrate the generality of the proposed approach by applying it to two Siamese trackers, SiamFC and DaSiamRPN. Extensive experiments on VOT2016, VOT2018, LaSOT, and TrackingNet datasets demonstrate that our UpdateNet effectively predicts the new target template, outperforming the standard linear update. On the large-scale TrackingNet dataset, our UpdateNet improves the results of DaSiamRPN with an absolute gain of 3.9% in terms of success score.
Lichao Zhang 0001, Abel Gonzalez-Garcia, Joost van de Weijer 0001, Martin Danelljan, Fahad Shahbaz Khan
ICCV2
2019 SDIT: Scalable and Diverse Cross-domain Image Translation
abstract
Recently, image-to-image translation research has witnessed remarkable progress. Although current approaches successfully generate diverse outputs or perform scalable image transfer, these properties have not been combined into a single method. To address this limitation, we propose SDIT: Scalable and Diverse image-to-image translation. These properties are combined into a single generator. The diversity is determined by a latent variable which is randomly sampled from a normal distribution. The scalability is obtained by conditioning the network on the domain attributes. Additionally, we also exploit an attention mechanism that permits the generator to focus on the domain-specific attribute. We empirically demonstrate the performance of the proposed method on face mapping and other datasets beyond faces.
Yaxing Wang, Abel Gonzalez-Garcia, Joost van de Weijer 0001, Luis Herranz
ACM Multimedia2
2019 Saliency for fine-grained object recognition in domains with scarce training data
Carola Figueroa Flores, Abel Gonzalez-Garcia, Joost van de Weijer 0001, Bogdan Raducanu
Pattern Recognit.2
2019 Synthetic Data Generation for End-to-End Thermal Infrared Tracking
abstract
The usage of both off-the-shelf and end-to-end trained deep networks have significantly improved the performance of visual tracking on RGB videos. However, the lack of large labeled datasets hampers the usage of convolutional neural networks for tracking in thermal infrared (TIR) images. Therefore, most state-of-the-art methods on tracking for TIR data are still based on handcrafted features. To address this problem, we propose to use image-to-image translation models. These models allow us to translate the abundantly available labeled RGB data to synthetic TIR data. We explore both the usage of paired and unpaired image translation models for this purpose. These methods provide us with a large labeled dataset of synthetic TIR sequences, on which we can train end-to-end optimal features for tracking. To the best of our knowledge, we are the first to train end-to-end features for TIR tracking. We perform extensive experiments on the VOT-TIR2017 dataset. We show that a network trained on a large dataset of synthetic TIR data obtains better performance than one trained on the available real TIR data. Combining both data sources leads to further improvement. In addition, when we combine the network with motion features, we outperform the state of the art with a relative gain of over 10%, clearly showing the efficiency of using synthetic data to train end-to-end TIR trackers.
Lichao Zhang 0001, Abel Gonzalez-Garcia, Joost van de Weijer 0001, Martin Danelljan, Fahad Shahbaz Khan
IEEE Trans. Image Process.2
2018 Objects as Context for Detecting Their Semantic Parts
abstract
We present a semantic part detection approach that effectively leverages object information. We use the object appearance and its class as indicators of what parts to expect. We also model the expected relative location of parts inside the objects based on their appearance. We achieve this with a new network module, called OffsetNet, that efficiently predicts a variable number of part locations within a given object. Our model incorporates all these cues to detect parts in the context of their objects. This leads to considerably higher performance for the challenging task of part detection compared to using part appearance alone (+5 mAP on the PASCAL-Part dataset). We also compare to other part detection methods on both PASCAL-Part and CUB200-2011 datasets.
Abel Gonzalez-Garcia, Davide Modolo, Vittorio Ferrari
CVPR1
2018 Transferring GANs: Generating Images from Limited Data
Yaxing Wang, Chenshen Wu, Luis Herranz, Joost van de Weijer 0001, Abel Gonzalez-Garcia, Bogdan Raducanu
ECCV (6)5
2018 Image-to-image translation for cross-domain disentanglement
abstract
Deep image translation methods have recently shown excellent results, outputting high-quality images covering multiple modes of the data distribution. There has also been increased interest in disentangling the internal representations learned by deep methods to further improve their performance and achieve a finer control. In this paper, we bridge these two objectives and introduce the concept of cross-domain disentanglement. We aim to separate the internal representation into three parts. The shared part contains information for both domains. The exclusive parts, on the other hand, contain only factors of variation that are particular to each domain. We achieve this through bidirectional image translation based on Generative Adversarial Networks and cross-domain autoencoders, a novel network component. Our model offers multiple advantages. We can output diverse samples covering multiple modes of the distributions of both domains, perform domain- specific image transfer and interpolation, and cross-domain retrieval without the need of labeled data, only paired images. We compare our model to the state-of-the-art in multi-modal image translation and achieve better results for translation on challenging datasets as well as for cross-domain retrieval on realistic datasets.
Abel Gonzalez-Garcia, Joost van de Weijer 0001, Yoshua Bengio
NeurIPS1
2018 Do Semantic Parts Emerge in Convolutional Neural Networks?
abstract
Semantic object parts can be useful for several visual recognition tasks. Lately, these tasks have been addressed using Convolutional Neural Networks (CNN), achieving outstanding results. In this work we study whether CNNs learn semantic parts in their internal representation. We investigate the responses of convolutional filters and try to associate their stimuli with semantic parts. We perform two extensive quantitative analyses. First, we use ground-truth part bounding-boxes from the PASCAL-Part dataset to determine how many of those semantic parts emerge in the CNN. We explore this emergence for different layers, network depths, and supervision levels. Second, we collect human judgements in order to study what fraction of all filters systematically fire on any semantic part, even if not annotated in PASCAL-Part. Moreover, we explore several connections between discriminative power and semantics. We find out which are the most discriminative filters for object recognition, and analyze whether they respond to semantic parts or to other image patches. We also investigate the other direction: we determine which semantic parts are the most discriminative and whether they correspond to those parts emerging in the network. This enables to gain an even deeper understanding of the role of semantic parts in the network.
Abel Gonzalez-Garcia, Davide Modolo, Vittorio Ferrari
Int. J. Comput. Vis.1
2015 An active search strategy for efficient object class detection
abstract
Object class detectors typically apply a window classifier to all the windows in a large set, either in a sliding window manner or using object proposals. In this paper, we develop an active search strategy that sequentially chooses the next window to evaluate based on all the information gathered before. This results in a substantial reduction in the number of classifier evaluations and in a more elegant approach in general. Our search strategy is guided by two forces. First, we exploit context as the statistical relation between the appearance of a window and its location relative to the object, as observed in the training set. This enables to jump across distant regions in the image (e.g. observing a sky region suggests that cars might be far below) and is done efficiently in a Random Forest framework. Second, we exploit the score of the classifier to attract the search to promising areas surrounding a highly scored window, and to keep away from areas near low scored ones. Our search strategy can be applied on top of any classifier as it treats it as a black-box. In experiments with R-CNN on the challenging SUN2012 dataset, our method matches the detection accuracy of evaluating all windows independently, while evaluating 9× fewer windows.
Abel Gonzalez-Garcia, Alexander Vezhnevets, Vittorio Ferrari
CVPR1