Alvaro Soto

dblp:25/3682 · also Álvaro Soto · DBLP profile ↗
← Back
47ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0001-9378-397XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
17 papers
Learning paradigms · 19% Representation and self-supervised learning · 16% Language models and text generation · 14%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 50% Knowledge graphs · 50%

Topics — the 30 heaviest of 39, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning paradigms
continual learning
1.222023
PIVOT: Prompting for Video Continual Learning · CVPR 2023
Optimizing Reusable Knowledge for Continual Learning via Metalearning · NeurIPS 2021
Computer vision › Video understanding and tracking
action recognition
1.042018
End-to-End Joint Semantic Segmentation of Actors and Actions in Video · ECCV (4) 2018
A Hierarchical Pose-Based Approach to Complex Action Understanding Using Dictionaries of Actionlets and Motion Poselets · CVPR 2016
Action Recognition in Video Using Sparse Coding and Relative Features · CVPR 2016
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.912025
Data Distributional Properties As Inductive Bias for Systematic Generalization · CVPR 2025
Machine learning › Representation and self-supervised learning
systematic generalization
0.912025
Data Distributional Properties As Inductive Bias for Systematic Generalization · CVPR 2025
Natural language and speech › Language models and text generation
prompting
0.712023
PIVOT: Prompting for Video Continual Learning · CVPR 2023
Machine learning › Learning paradigms › continual learning
video continual learning
0.712023
PIVOT: Prompting for Video Continual Learning · CVPR 2023
Natural language and speech › Language models and text generation › instruction tuning
instruction generation
0.612022
Bridging the Visual Semantic Gap in VLN via Semantically Richer Instructions · ECCV (37) 2022
Computer vision › Vision and language
vision-and-language navigation
0.612022
Bridging the Visual Semantic Gap in VLN via Semantically Richer Instructions · ECCV (37) 2022
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.512021
Optimizing Reusable Knowledge for Continual Learning via Metalearning · NeurIPS 2021
Machine learning › Efficient and distributed learning
parameter-efficient learning
0.512021
Optimizing Reusable Knowledge for Continual Learning via Metalearning · NeurIPS 2021
Machine learning › Efficient and distributed learning
adaptive computation
0.412020
Differentiable Adaptive Computation Time for Visual Reasoning · CVPR 2020
Bioinformatics and computational biology › functional genomics
gene function prediction
0.422017
GENIUS: web server to predict local gene networks and key genes for biological functions · Bioinform. 2017
Discriminative local subspaces in gene expression data for effective gene function prediction · Bioinform. 2012
Computer vision › Image recognition and object detection
visual recognition
0.422015
Learning Shared, Discriminative, and Compact Representations for Visual Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Hierarchical Joint Max-Margin Learning of Mid and Top Level Representations for Visual Recognition · ICCV 2013
Robotics › Robot navigation and mapping › mobile robot navigation
indoor navigation
0.312018
A Deep Learning Based Behavioral Approach to Indoor Autonomous Navigation · ICRA 2018
Natural language and speech › Language models and text generation
instruction following
0.312018
Translating Navigation Instructions in Natural Language to a High-Level Plan for Behavioral Robot Navigation · EMNLP 2018
Robotics › Robot navigation and mapping › embodied navigation
natural language navigation instructions
0.312018
Translating Navigation Instructions in Natural Language to a High-Level Plan for Behavioral Robot Navigation · EMNLP 2018
Robotics › Robot navigation and mapping › visual navigation
semantic navigation
0.312018
A Deep Learning Based Behavioral Approach to Indoor Autonomous Navigation · ICRA 2018
Computer vision › Segmentation and scene understanding
semantic segmentation
0.312018
End-to-End Joint Semantic Segmentation of Actors and Actions in Video · ECCV (4) 2018
Bioinformatics and computational biology › network bioinformatics › biological network analysis › biological network inference
gene network inference
0.312017
GENIUS: web server to predict local gene networks and key genes for biological functions · Bioinform. 2017
Knowledge graphs
commonsense knowledge
0.312017
How a General-Purpose Commonsense Ontology can Improve Performance of Learning-Based Image Retrieval · IJCAI 2017
Information retrieval
cross-modal retrieval
0.312017
How a General-Purpose Commonsense Ontology can Improve Performance of Learning-Based Image Retrieval · IJCAI 2017
Machine learning › Representation and self-supervised learning › visual representation › image representation › mid-level representation
mid-level representation learning
0.212015
Learning Shared, Discriminative, and Compact Representations for Visual Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Computer vision › Video understanding and tracking › activity recognition
complex activity recognition
0.212014
Discriminative Hierarchical Modeling of Spatio-temporally Composable Human Activities · CVPR 2014
Machine learning › Representation and self-supervised learning › visual representation › image representation
bag of visual words
0.212013
Hierarchical Joint Max-Margin Learning of Mid and Top Level Representations for Visual Recognition · ICCV 2013
Machine learning › Representation and self-supervised learning › hierarchical representation
hierarchical representation learning
0.212013
Hierarchical Joint Max-Margin Learning of Mid and Top Level Representations for Visual Recognition · ICCV 2013
Natural language and speech › Language models and text generation › pre-trained language model
BERT
0.112021
Augmenting BERT-style Models with Predictive Coding to Improve Discourse-level Representations · EMNLP (1) 2021
Bioinformatics and computational biology › network bioinformatics › biological network analysis
gene co-expression network analysis
0.112012
Discriminative local subspaces in gene expression data for effective gene function prediction · Bioinform. 2012
Computer vision › Face, body and person analysis
human pose estimation
0.122016
A Hierarchical Pose-Based Approach to Complex Action Understanding Using Dictionaries of Actionlets and Motion Poselets · CVPR 2016
Discriminative Hierarchical Modeling of Spatio-temporally Composable Human Activities · CVPR 2014
Computer vision › Vision and language
visual reasoning
0.112020
Differentiable Adaptive Computation Time for Visual Reasoning · CVPR 2020
Computer vision › Image recognition and object detection
object detection
0.122017
How a General-Purpose Commonsense Ontology can Improve Performance of Learning-Based Image Retrieval · IJCAI 2017
Indoor scene recognition through object detection · ICRA 2010

Methods — techniques the papers use, named apart from their topics

deep learning · 0.9normalized mutual information · 0.9prompting · 0.7pre-trained model · 0.7semantic instruction enrichment · 0.6trainable mask · 0.5self-supervised learning · 0.5predictive coding · 0.5meta-learning · 0.5knowledge base · 0.5support vector machine · 0.4object detectors · 0.3machine learning · 0.3conceptnet · 0.3supervised machine learning · 0.1
YearPublicationVenuePosition
2025 Data Distributional Properties As Inductive Bias for Systematic Generalization
abstract
Deep neural networks (DNNs) struggle at systematic generalization (SG). Several studies have evaluated the possibility of promoting SG through the proposal of novel architectures, loss functions, or training methodologies. Few studies, however, have focused on the role of training data properties in promoting SG. In this work, we investigate the impact of certain data distributional properties, as inductive biases for the SG ability of a multi-modal language model. To this end, we study three different properties. First, data diversity, instantiated as an increase in the possible values a latent property in the training distribution may take. Second, burstiness, where we probabilistically restrict the number of possible values of latent factors on particular inputs during training. Third, latent intervention, where a particular latent factor is altered randomly during training. We find that all three factors significantly enhance SG, with diversity contributing an 89% absolute increase in accuracy in the most affected property. Through a series of experiments, we test various hypotheses to understand why these properties promote SG. Finally, we find that Normalized Mutual Information (NMI) between latent attributes in the training distribution is strongly predictive of out-of-distribution generalization. We find that a mechanism by which lower NMI induces SG is in the geometry of representations. In particular, we find that NMI induces more parallelism in neural representations (i.e., input features coded in parallel neural vectors) of the model, a property related to the capacity of reasoning by analogy. Our code is available at: https://github.com/fdelrio89/data-systematic
Felipe del Río, Alain Raymond-Saez, Daniel Florea, Rodrigo Toro Icarte, Julio Hurtado, Cristian Buc Calderon, Alvaro Soto
CVPR7
2025 CXR-LT 2024: A MICCAI challenge on long-tailed, multi-label, and zero-shot disease classification from chest X-ray
Mingquan Lin, Gregory Holste, Song Wang 0026, Yiliang Zhou, Yishu Wei, Imon Banerjee, Pengyi Chen, Tianjie Dai, Yuexi Du, Nicha C. Dvornek, Yuyan Ge, Zuwei Guo, Shohei Hanaoka, Dongkyun Kim, Pablo Messina, Yang Lu 0009, Denis Parra, Donghyun Son, Alvaro Soto, Aisha Urooj Khan, René Vidal, Yosuke Yamagishi, Pingkun Yan, Zefan Yang, Ruichi Zhang, Yang Zhou 0019, Leo A. Celi, Ronald M. Summers, Zhiyong Lu, Hao Chen 0011, Adam E. Flanders, George Shih, Zhangyang Wang, Yifan Peng 0002
Medical Image Anal.19
2023 PIVOT: Prompting for Video Continual Learning
abstract
Modern machine learning pipelines are limited due to data availability, storage quotas, privacy regulations, and expensive annotation processes. These constraints make it difficult or impossible to train and update large-scale models on such dynamic annotated sets. Continual learning directly approaches this problem, with the ultimate goal of devising methods where a deep neural network effectively learns relevant patterns for new (unseen) classes, without significantly altering its performance on previously learned ones. In this paper, we address the problem of continual learning for video data. We introduce PIVOT, a novel method that leverages extensive knowledge in pre-trained models from the image domain, thereby reducing the number of trainable parameters and the associated forgetting. Unlike previous methods, ours is the first approach that effectively uses prompting mechanisms for continual learning without any in-domain pre-training. Our experiments show that PIVOT improves state-of-the-art methods by a significant 27% on the 20-task ActivityNet setup.
Andrés Villa, Juan Leon Alcazar, Motasem Alfarra, Kumail Alhamoud, Julio Hurtado, Fabian Caba Heilbron, Alvaro Soto, Bernard Ghanem
CVPR7
2022 Bridging the Visual Semantic Gap in VLN via Semantically Richer Instructions
Joaquín Ossandón, Benjamín Earle, Alvaro Soto
ECCV (37)3
2022 Evaluation Benchmarks for Spanish Sentence Representations
abstract
Due to the success of pre-trained language models, versions of languages other than English have been released in recent years. This fact implies the need for resources to evaluate these models. In the case of Spanish, there are few ways to systematically assess the models’ quality. In this paper, we narrow the gap by building two evaluation benchmarks. Inspired by previous work (Conneau and Kiela, 2018; Chen et al., 2019), we introduce Spanish SentEval and Spanish DiscoEval, aiming to assess the capabilities of stand-alone and discourse-aware sentence representations, respectively. Our benchmarks include considerable pre-existing and newly constructed datasets that address different tasks from various domains. In addition, we evaluate and analyze the most recent pre-trained Spanish language models to exhibit their capabilities and limitations. As an example, we discover that for the case of discourse evaluation tasks, mBERT, a language model trained on multiple languages, usually provides a richer latent representation than models trained only with documents in Spanish. We hope our contribution will motivate a fairer, more comparable, and less cumbersome way to evaluate future Spanish language models.
Vladimir Araujo, Andres Carvallo, Souvik Kundu 0008, José Cañete, Marcelo Mendoza, Robert E. Mercer, Felipe Bravo-Marquez, Marie-Francine Moens, Alvaro Soto
LREC9
2021 TNT: Text-Conditioned Network with Transductive Inference for Few-Shot Video Classification
Andrés Villa, Juan-Manuel Pérez-Rúa, Vladimir Araujo, Juan Carlos Niebles, Victor Escorcia, Alvaro Soto
BMVC6
2021 Augmenting BERT-style Models with Predictive Coding to Improve Discourse-level Representations
abstract
Current language models are usually trained using a self-supervised scheme, where the main focus is learning representations at the word or sentence level.However, there has been limited progress in generating useful discourse-level representations.In this work, we propose to use ideas from predictive coding theory to augment BERT-style language models with a mechanism that allows them to learn suitable discourse-level representations.As a result, our proposed approach is able to predict future sentences using explicit top-down connections that operate at the intermediate layers of the network.By experimenting with benchmarks designed to evaluate discourse-related knowledge using pre-trained sentence representations, we demonstrate that our approach improves performance in 6 out of 11 tasks by excelling in discourse relationship detection.
Vladimir Araujo, Andrés Villa, Marcelo Mendoza, Marie-Francine Moens, Alvaro Soto
EMNLP (1)5
2021 Optimizing Reusable Knowledge for Continual Learning via Metalearning
abstract
When learning tasks over time, artificial neural networks suffer from a problem known as Catastrophic Forgetting (CF). This happens when the weights of a network are overwritten during the training of a new task causing forgetting of old information. To address this issue, we propose MetA Reusable Knowledge or MARK, a new method that fosters weight reusability instead of overwriting when learning a new task. Specifically, MARK keeps a set of shared weights among tasks. We envision these shared weights as a common Knowledge Base (KB) that is not only used to learn new tasks, but also enriched with new knowledge as the model learns new tasks. Key components behind MARK are two-fold. On the one hand, a metalearning approach provides the key mechanism to incrementally enrich the KB with new knowledge and to foster weight reusability among tasks. On the other hand, a set of trainable masks provides the key mechanism to selectively choose from the KB relevant weights to solve each task. By using MARK, we achieve state of the art results in several popular benchmarks, surpassing the best performing methods in terms of average accuracy by over 10% on the 20-Split-MiniImageNet dataset, while achieving almost zero forgetfulness using 55% of the number of parameters. Furthermore, an ablation study provides evidence that, indeed, MARK is learning reusable knowledge that is selectively used by each task.
Julio Hurtado, Alain Raymond-Saez, Alvaro Soto
NeurIPS3
2020 Differentiable Adaptive Computation Time for Visual Reasoning
abstract
This paper presents a novel attention-based algorithm for achieving adaptive computation called DACT, which, unlike existing ones, is end-to-end differentiable. Our method can be used in conjunction with many networks; in particular, we study its application to the widely know MAC architecture, obtaining a significant reduction in the number of recurrent steps needed to achieve similar accuracies, therefore improving its performance to computation ratio. Furthermore, we show that by increasing the maximum number of steps used, we surpass the accuracy of even our best non-adaptive MAC in the CLEVR dataset, demonstrating that our approach is able to control the number of steps without significant loss of performance. Additional advantages provided by our approach include considerably improving interpretability by discarding useless steps and providing more insights into the underlying reasoning process. Finally, we present adaptive computation as an equivalent to an ensemble of models, similar to a mixture of expert formulation. Both the code and the configuration files for our experiments are made available to support further research in this area.
Cristóbal Eyzaguirre, Alvaro Soto
CVPR2
2020 CompactNets: Compact Hierarchical Compositional Networks for Visual Recognition
Hans Lobel, René Vidal, Alvaro Soto
Comput. Vis. Image Underst.3
2020 GENE: Graph generation conditioned on named entities for polarity and controversy detection in social media
Marcelo Mendoza, Denis Parra, Alvaro Soto
Inf. Process. Manag.3
2020 Explaining VQA predictions using visual grounding and a knowledge base
Felipe Riquelme, Alfredo De Goyeneche, Yundong Zhang 0001, Juan Carlos Niebles, Alvaro Soto
Image Vis. Comput.5
2019 Interpretable Visual Question Answering by Visual Grounding From Attention Supervision Mining
abstract
A key aspect of visual question answering (VQA) models that are interpretable is their ability to ground their answers to relevant regions in the image. Current approaches with this capability rely on supervised learning and human annotated groundings to train attention mechanisms inside the VQA architecture. Unfortunately, obtaining human annotations specific for visual grounding is difficult and expensive. In this work, we demonstrate that we can effectively train a VQA architecture with grounding supervision that can be automatically obtained from available region descriptions and object annotations. We also show that our model trained with this mined supervision generates visual groundings that achieve a higher correlation with respect to manually-annotated groundings, meanwhile achieving state-of-the-art VQA accuracy.
Yundong Zhang 0001, Juan Carlos Niebles, Alvaro Soto
WACV3
2019 Content-based artwork recommendation: integrating painting metadata with neural and manually-engineered visual features
Pablo Messina, Vicente Dominguez, Denis Parra, Christoph Trattner, Alvaro Soto
User Model. User Adapt. Interact.5
2018 End-to-End Joint Semantic Segmentation of Actors and Actions in Video
Jingwei Ji, Shyamal Buch, Alvaro Soto, Juan Carlos Niebles
ECCV (4)3
2018 Translating Navigation Instructions in Natural Language to a High-Level Plan for Behavioral Robot Navigation
abstract
Xiaoxue Zang, Ashwini Pokle, Marynel Vázquez, Kevin Chen, Juan Carlos Niebles, Alvaro Soto, Silvio Savarese. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
Xiaoxue Zang, Ashwini Pokle, Marynel Vázquez, Kevin Chen 0001, Juan Carlos Niebles, Alvaro Soto, Silvio Savarese
EMNLP6
2018 A Deep Learning Based Behavioral Approach to Indoor Autonomous Navigation
abstract
We present a semantically rich graph representation for indoor robotic navigation. Our graph representation encodes: semantic locations such as offices or corridors as nodes, and navigational behaviors such as enter office or cross a corridor as edges. In particular, our navigational behaviors operate directly from visual inputs to produce motor controls and are implemented with deep learning architectures. This enables the robot to avoid explicit computation of its precise location or the geometry of the environment, and enables navigation at a higher level of semantic abstraction. We evaluate the effectiveness of our representation by simulating navigation tasks in a large number of virtual environments. Our results show that using a simple sets of perceptual and navigational behaviors, the proposed approach can successfully guide the way of the robot as it completes navigational missions such as going to a specific office. Furthermore, our implementation shows to be effective to control the selection and switching of behaviors.
Gabriel Sepulveda, Juan Carlos Niebles, Alvaro Soto
ICRA3
2017 Unsupervised Local Regressive Attributes for Pedestrian Re-identification
Billy Peralta, Luis Alberto Caro, Alvaro Soto
CIARP3
2017 How a General-Purpose Commonsense Ontology can Improve Performance of Learning-Based Image Retrieval
abstract
The knowledge representation community has built general-purpose ontologies which contain large amounts of commonsense knowledge over relevant aspects of the world, including useful visual information, e.g.: "a ball is used by a football player", "a tennis player is located at a tennis court". Current state-of-the-art approaches for visual recognition do not exploit these rule-based knowledge sources. Instead, they learn recognition models directly from training examples. In this paper, we study how general-purpose ontologies—specifically, MIT's ConceptNet ontology—can improve the performance of state-of-the-art vision systems. As a testbed, we tackle the problem of sentence-based image retrieval. Our retrieval approach incorporates knowledge from ConceptNet on top of a large pool of object detectors derived from a deep learning technique. In our experiments, we show that ConceptNet can improve performance on a common benchmark dataset. Key to our performance is the use of the ESPGAME dataset to select visually relevant relations from ConceptNet. Consequently, a main conclusion of this work is that general-purpose commonsense ontologies improve performance on visual reasoning tasks when properly filtered to select meaningful visual relations.
Rodrigo Toro Icarte, Jorge A. Baier, Cristian Ruz, Alvaro Soto
IJCAI4
2017 GENIUS: web server to predict local gene networks and key genes for biological functions
abstract
Summary: GENIUS is a user-friendly web server that uses a novel machine learning algorithm to infer functional gene networks focused on specific genes and experimental conditions that are relevant to biological functions of interest. These functions may have different levels of complexity, from specific biological processes to complex traits that involve several interacting processes. GENIUS also enriches the network with new genes related to the biological function of interest, with accuracies comparable to highly discriminative Support Vector Machine methods. Availability and Implementation: GENIUS currently supports eight model organisms and is freely available for public use at http://networks.bio.puc.cl/genius . Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Tomas Puelma, Viviana Araus, Javier Canales, Elena A. Vidal, Juan M. Cabello, Alvaro Soto, Rodrigo A. Gutiérrez
Bioinform.6
2017 Sparse composition of body poses and atomic actions for human activity recognition in RGB-D videos
Ivan Lillo, Juan Carlos Niebles, Alvaro Soto
Image Vis. Comput.3
2017 Corrigendum to "Sparse Composition of Body Poses and Atomic Actions for Human Activity Recognition in RGB-D Videos" [Image Vis. Comput. 59 (2017) 63-75]
Ivan Lillo, Juan Carlos Niebles, Alvaro Soto
Image Vis. Comput.3
2016 Action Recognition in Video Using Sparse Coding and Relative Features
abstract
This work presents an approach to category-based action recognition in video using sparse coding techniques. The proposed approach includes two main contributions: i) A new method to handle intra-class variations by decomposing each video into a reduced set of representative atomic action acts or key-sequences, and ii) A new video descriptor, ITRA: Inter-Temporal Relational Act Descriptor, that exploits the power of comparative reasoning to capture relative similarity relations among key-sequences. In terms of the method to obtain key-sequences, we introduce a loss function that, for each video, leads to the identification of a sparse set of representative key-frames capturing both, relevant particularities arising in the input video, as well as relevant generalities arising in the complete class collection. In terms of the method to obtain the ITRA descriptor, we introduce a novel scheme to quantify relative intra and inter-class similarities among local temporal patterns arising in the videos. The resulting ITRA descriptor demonstrates to be highly effective to discriminate among action categories. As a result, the proposed approach reaches remarkable action recognition performance on several popular benchmark datasets, outperforming alternative state-of the-art techniques by a large margin.
Analí Alfaro, Domingo Mery, Alvaro Soto
CVPR3
2016 A Hierarchical Pose-Based Approach to Complex Action Understanding Using Dictionaries of Actionlets and Motion Poselets
abstract
In this paper, we introduce a new hierarchical model for human action recognition using body joint locations. Our model can categorize complex actions in videos, and perform spatio-temporal annotations of the atomic actions that compose the complex action being performed. That is, for each atomic action, the model generates temporal action annotations by estimating its starting and ending times, as well as, spatial annotations by inferring the human body parts that are involved in executing the action. Our model includes three key novel properties: (i) it can be trained with no spatial supervision, as it can automatically discover active body parts from temporal action annotations only, (ii) it jointly learns flexible representations for motion poselets and actionlets that encode the visual variability of body parts and atomic actions, (iii) a mechanism to discard idle or non-informative body parts which increases its robustness to common pose estimation errors. We evaluate the performance of our method using multiple action recognition benchmarks. Our model consistently outperforms baselines and state-of-the-art action recognition methods.
Ivan Lillo, Juan Carlos Niebles, Alvaro Soto
CVPR3
2016 A proposal for supervised clustering with Dirichlet Process using labels
Billy Peralta, Luis Alberto Caro, Alvaro Soto
Pattern Recognit. Lett.3
2015 Visual Recognition to Access and Analyze People Density and Flow Patterns in Indoor Environments
abstract
This work describes our experience developing a system to access density and flow of people in large indoor spaces using a network of RGB cameras. The proposed system is based on a set of overlapped and calibrated cameras. This facilitates the use of geometric constraints that help to reduce visual ambiguities. These constraints are combined with classifiers based on visual appearance to produce an efficient and robust method to detect and track humans. In this work, we argue that flow and density of people are low level measurements that need to be complemented with suitable analytic tools to bridge semantic gaps and become useful information for a target application. Consequently, we also propose a set of analytic tools that help a human user to effectively take advantage of the measurements provided by the system. Finally, we report results that demonstrate the relevance of the proposed ideas.
Cristian Ruz, Christian Pieringer, Billy Peralta, Ivan Lillo, Pablo Espinace, R. Gonzalez, B. Wendt, Domingo Mery, Alvaro Soto
WACV9
2015 Learning Shared, Discriminative, and Compact Representations for Visual Recognition
abstract
Dictionary-based and part-based methods are among the most popular approaches to visual recognition. In both methods, a mid-level representation is built on top of low-level image descriptors and high-level classifiers are trained on top of the mid-level representation. While earlier methods built the mid-level representation without supervision, there is currently great interest in learning both representations jointly to make the mid-level representation more discriminative. In this work we propose a new approach to visual recognition that jointly learns a shared, discriminative, and compact mid-level representation and a compact high-level representation. By using a structured output learning framework, our approach directly handles the multiclass case at both levels of abstraction. Moreover, by using a group-sparse prior in the structured output learning framework, our approach encourages sharing of visual words and thus reduces the number of words used to represent each class. We test our proposed method on several popular benchmarks. Our results show that, by jointly learning mid- and high-level representations, and fostering the sharing of discriminative visual words among target classes, we are able to achieve state-of-the-art recognition performance using far less visual words than previous approaches.
Hans Lobel, René Vidal, Alvaro Soto
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Discriminative Hierarchical Modeling of Spatio-temporally Composable Human Activities
abstract
This paper proposes a framework for recognizing complex human activities in videos. Our method describes human activities in a hierarchical discriminative model that operates at three semantic levels. At the lower level, body poses are encoded in a representative but discriminative pose dictionary. At the intermediate level, encoded poses span a space where simple human actions are composed. At the highest level, our model captures temporal and spatial compositions of actions into complex human activities. Our human activity classifier simultaneously models which body parts are relevant to the action of interest as well as their appearance and composition using a discriminative approach. By formulating model learning in a max-margin framework, our approach achieves powerful multi-class discrimination while providing useful annotations at the intermediate semantic level. We show how our hierarchical compositional model provides natural handling of occlusions. To evaluate the effectiveness of our proposed framework, we introduce a new dataset of composed human activities. We provide empirical evidence that our method achieves state-of-the-art activity classification performance on several benchmark datasets.
Ivan Lillo, Alvaro Soto, Juan Carlos Niebles
CVPR2
2014 Local feature selection using Gaussian process regression
abstract
Most feature selection methods determine a global subset of features, where all data instances are projected in order to improve classification accuracy. An attractive alternative solution is to adaptively find a local subset of features for each data instance, such that, the classification of each instance is performed according to its own selective subspace. This paper presents a novel application of Gaussian Processes (GPs) that improves classification performance by learning a set of functions that quantify the discriminative power of each feature. Specifically, GP regressions are used to build for each available feature a function that estimates its discriminative properties over all its input space. Afterwards, by locally joining these regressions it is possible to obtain a discriminative subspace for any position of the input space. New instances are then classified by using a K-NN classifier that operates in the local subspaces. Experimental results show that by using local discriminative subspaces, we are able to reach higher levels of classification accuracy than alternative state-of-the-art feature selection approaches.
Karim Pichara, Alvaro Soto
Intell. Data Anal.2
2014 Embedded local feature selection within mixture of experts
Billy Peralta, Alvaro Soto
Inf. Sci.2
2013 Hierarchical Joint Max-Margin Learning of Mid and Top Level Representations for Visual Recognition
abstract
Currently, Bag-of-Visual-Words (BoVW) and part-based methods are the most popular approaches for visual recognition. In both cases, a mid-level representation is built on top of low-level image descriptors and top-level classifiers use this mid-level representation to achieve visual recognition. While in current part-based approaches, mid- and top-level representations are usually jointly trained, this is not the usual case for BoVW schemes. A main reason for this is the complex data association problem related to the usual large dictionary size needed by BoVW approaches. As a further observation, typical solutions based on BoVW and part-based representations are usually limited to extensions of binary classification schemes, a strategy that ignores relevant correlations among classes. In this work we propose a novel hierarchical approach to visual recognition based on a BoVW scheme that jointly learns suitable mid- and top-level representations. Furthermore, using a max-margin learning framework, the proposed approach directly handles the multiclass case at both levels of abstraction. We test our proposed method using several popular benchmark datasets. As our main result, we demonstrate that, by coupling learning of mid- and top-level representations, the proposed approach fosters sharing of discriminative visual words among target classes, being able to achieve state-of-the-art recognition performance using far less visual words than previous approaches.
Hans Lobel, René Vidal, Alvaro Soto
ICCV3
2013 Human Action Recognition from Inter-temporal Dictionaries of Key-Sequences
Analí Alfaro, Domingo Mery, Alvaro Soto
PSIVT3
2013 Joint Dictionary and Classifier Learning for Categorization of Images Using a Max-margin Framework
Hans Lobel, René Vidal, Domingo Mery, Alvaro Soto
PSIVT4
2013 Enhancing K-Means using class labels
abstract
Clustering is a relevant problem in machine learning where the main goal is to locate meaningful partitions of unlabeled data. In the case of labeled data, a related problem is supervised clustering, where the objective is to locate class-uniform clusters. Most current approaches to supervised clus tering optimize a score related to cluster purity with respect to class labels. In particular, we present Labeled K-Means (LK-Means), an algorithm for supervised clustering based on a variant of K-Means that incorporates information about class labels. LK-Means replaces the classical cost function of K-Means by a convex combination of the joint cost associated to: (i) A discriminative score based on class labels, and (ii) A generative score based on a traditional metric for unsupervised clustering. We test the performance of LK-Means using standard real datasets and an application for object recognition. Moreover, we also compare its performance against classical K-Means and a popular K-Medoids-based supervised clustering method. Our experiments show that, in most cases, LK-Means outperforms the alternative techniques by a considerable margin. Furthermore, LK-Means presents execution times considerably lower than the alternative supervised clustering method under evaluation.
Billy Peralta, Pablo Espinace, Alvaro Soto
Intell. Data Anal.3
2012 Adaptive hierarchical contexts for object recognition with conditional mixture of trees
abstract
Robust category-level object recognition is currently a major goal for the computer vision community. Intra-class and pose variations, as well as, background clutter and partial occlusions are some of the main difficulties to achieve this goal. Contextual information, in the form of object co-occurrences and spatial constraints, has been successfully applied to improve object recognition performance, however, previous work considers only fixed contextual relations that do not depend of the type of scene under inspection. In this work, we present a method that learns adaptive conditional relationships that depend on the type of scene being analyzed. In particular, we propose a model based on a conditional mixture of trees that is able to capture contextual relationships among objects using global information about a scene. Our experiments show that the adaptive specialization of contextual relationships improves object recognition accuracy outperforming previous state-of-the-art approaches.
Billy Peralta, Pablo Espinace, Alvaro Soto
BMVC3
2012 Discriminative local subspaces in gene expression data for effective gene function prediction
abstract
MOTIVATION: Massive amounts of genome-wide gene expression data have become available, motivating the development of computational approaches that leverage this information to predict gene function. Among successful approaches, supervised machine learning methods, such as Support Vector Machines (SVMs), have shown superior prediction accuracy. However, these methods lack the simple biological intuition provided by co-expression networks (CNs), limiting their practical usefulness. RESULTS: In this work, we present Discriminative Local Subspaces (DLS), a novel method that combines supervised machine learning and co-expression techniques with the goal of systematically predict genes involved in specific biological processes of interest. Unlike traditional CNs, DLS uses the knowledge available in Gene Ontology (GO) to generate informative training sets that guide the discovery of expression signatures: expression patterns that are discriminative for genes involved in the biological process of interest. By linking genes co-expressed with these signatures, DLS is able to construct a discriminative CN that links both, known and previously uncharacterized genes, for the selected biological process. This article focuses on the algorithm behind DLS and shows its predictive power using an Arabidopsis thaliana dataset and a representative set of 101 GO terms from the Biological Process Ontology. Our results show that DLS has a superior average accuracy than both SVMs and CNs. Thus, DLS is able to provide the prediction accuracy of supervised learning methods while maintaining the intuitive understanding of CNs. AVAILABILITY: A MATLAB® implementation of DLS is available at http://virtualplant.bio.puc.cl/cgi-bin/Lab/tools.cgi.
Tomas Puelma, Rodrigo A. Gutiérrez, Alvaro Soto
Bioinform.3
2011 Mixing Hierarchical Contexts for Object Recognition
Billy Peralta, Alvaro Soto
CIARP2
2011 Learning discriminative local binary patterns for face recognition
abstract
Histograms of Local Binary Patterns (LBPs) and variations thereof are a popular local visual descriptor for face recognition. So far, most variations of LBP are designed by hand or are learned with non-supervised methods. In this work we propose a simple method to learn discriminative LBPs in a supervised manner. The method represents an LBP-like descriptor as a set of pixel comparisons within a neighborhood and heuristically seeks for a set of pixel comparisons so as to maximize a Fisher separability criterion for the resulting histograms. Tests on standard face recognition datasets show that this method can create compact yet discriminative descriptors.
Daniel Maturana, Domingo Mery, Alvaro Soto
FG3
2011 Active learning and subspace clustering for anomaly detection
abstract
Today, anomaly detection is a highly valuable application in the analysis of current huge datasets. Insurance companies, banks and many manufacturing industries need systems to help humans to detect anomalies in their daily information. In general, anomalies are a very small fraction of the data, t herefore their detection is not an easy task. Usually real sources of an anomaly are given by specific values expressed on selective dimensions of datasets, furthermore, many anomalies are not really interesting for humans, due to the fact that interestingness of anomalies is categorized subjectively by the human user. In this paper we propose a new semi-supervised algorithm that actively learns to detect relevant anomalies by interacting with an expert user in order to obtain semantic information about user preferences. Our approach is based on 3 main steps. First, a Bayes network identifies an initial set of candidate anomalies. Afterwards, a subspace clustering technique identifies relevant subsets of dimensions. Finally, a probabilistic active learning scheme, based on properties of Dirichlet distribution, uses the feedback from an expert user to efficiently search for relevant anomalies. Our results, using synthetic and real datasets, indicate that, under noisy data and anomalies presenting regular patterns, our approach correctly identifies relevant anomalies.
Karim Pichara, Alvaro Soto
Intell. Data Anal.2
2010 Face Recognition with Decision Tree-Based Local Binary Patterns
Daniel Maturana, Domingo Mery, Alvaro Soto
ACCV (4)3
2010 Indoor scene recognition through object detection
abstract
Scene recognition is a highly valuable perceptual ability for an indoor mobile robot, however, current approaches for scene recognition present a significant drop in performance for the case of indoor scenes. We believe that this can be explained by the high appearance variability of indoor environments. This stresses the need to include high-level semantic information in the recognition process. In this work we propose a new approach for indoor scene recognition based on a generative probabilistic hierarchical model that uses common objects as an intermediate semantic representation. Under this model, we use object classifiers to associate low-level visual features to objects, and at the same time, we use contextual relations to associate objects to scenes. As a further contribution, we improve the performance of current state-of-the-art category-level object classifiers by including geometrical information obtained from a 3D range sensor that facilitates the implementation of a focus of attention mechanism within a Monte Carlo sampling scheme. We test our approach using real data, showing significant advantages with respect to previous state-of-the-art methods.
Pablo Espinace, Thomas Kollar, Alvaro Soto, Nicholas Roy
ICRA3
2010 Human detection using a mobile platform and novel features derived from a visual saliency mechanism
Sebastian Montabone, Alvaro Soto
Image Vis. Comput.2
2007 Human Detection in Indoor Environments Using Multiple Visual Cues and a Mobile Robot
Stefan Pszczólkowski, Alvaro Soto
CIARP2
2007 An Accelerated Algorithm for Density Estimation in Large Databases Using Gaussian Mixtures
abstract
Today, with the advances of computer storage and technology, there are huge datasets available, offering an opportunity to extract valuable information. Probabilistic approaches are specially suited to learn from data by representing knowledge as density functions. In this paper, we choose Gaussian mixture models (GMMs) to represent densities, as they possess great flexibility to adequate to a wide class of problems. The classical estimation approach for GMMs corresponds to the iterative algorithm of expectation maximization (EM). This approach, however, does not scale properly to meet the high demanding processing requirements of large databases. In this paper we introduce an EM-based algorithm, that solves the scalability problem. Our approach is based on the concept of data condensation which, in addition to substantially diminishing the computational load, provides sound starting values that allow the algorithm to reach convergence faster. We also focus on the model selection problem. We test our algorithm using synthetic and real databases, and find several advantages, when compared to other standard existing procedures.
Alvaro Soto, Felipe Zavala, Anita Araneda
Cybern. Syst.1
2006 Automatic Selection and Detection of Visual Landmarks Using Multiple Segmentations
Daniel Langdon, Alvaro Soto, Domingo Mery
PSIVT2
2005 Self Adaptive Particle Filter
Alvaro Soto
IJCAI1
2001 Probabilistic Adaptive Agent Based System for Dynamic State Estimation using Multiple Visual Cues
Alvaro Soto, Pradeep K. Khosla
ISRR1