Jordi Vitrià

dblp:41/75 · also Jordi Vitrià Marca · DBLP profile ↗
← Back
81ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0003-1484-539XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 51 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6Databases, data management, data science and information retrieval · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Solving the Contextual Pure Cold-Start Problem under Uncertainty
abstract
Part of the success of an online media platform depends on its ability to convert first-time users into recurring ones. However, this task is often challenging due to the Pure Cold-start problem, which refers to the difficulty of providing useful recommendations to users without historical data. Although presenting a list of popular items may seem like an easy fix, it can lead to the ”popularity bias” problem. In this article, we address a specific issue: the Contextual Pure Cold-start Problem in Public Service Media (PSM) scenarios, characterized by the absence of user-specific data and the presence of only anonymous and limited contextual information. We propose a novel approach aligned with PSM values in recommendations, as ensuring these values is crucial for their task. Our approach seeks to enhance the equity of recommendation rankings by incorporating various types of uncertainty, providing a solution that ensures a fairer exposure to items and broader coverage of the PSM catalog. Using a substantial dataset of real interactions from a public television network, our approach successfully mitigates the ”popularity bias” issue through the use of an uncertainty-based stochastic ranker. Consequently, we achieve a 64% improvement in fair exposure and a 42% increase in coverage metrics, with only an 8% reduction in Hit Rate accuracy metrics.
Paula Gómez Duran, Axel Brando, Jordi Vitrià
Trans. Recomm. Syst.3
2025 Optimizing Reachability in Graph-Based Recommender Systems
abstract
While accuracy has long been prioritized as the primary metric for Recommender Systems (RSs), it is increasingly accepted that the system’s overall quality is not solely determined by this factor. Reachability, the ease with which users can navigate the whole content catalog through recommendations, emerges as a pivotal yet under-explored concept: not only it ensures a smooth experience for users, but it also provides more equitable exposure for the items, avoiding that only a small fraction of popular items get the bulk of the attention. Despite its importance, the few existing studies analyze reachability without attempting a proper optimization. In this article, we study the problem of optimizing the overall reachability of a RS while maintaining high-quality recommendations. We model a user browsing session as a random walk on a recommendation graph, where the links and the transition probabilities are defined based on the relevance score of the recommendation list that the user gets at every step. In this setting, reachability is modeled as the expected length of a path to reach a given item. We introduce two optimization problems, one discrete and one continuous, and characterize their theoretical properties. We then devise two algorithms that outperform non-trivial baseline methods in enhancing reachability while maintaining a high Normalized Discounted Cumulative Gain (nDCG) score. Our experimental results show that, in some settings, our methods are able to improve the reachability metric by 80% while only compromising nDCG by 5%. Moreover, our empirical analysis shows that optimizing for reachability provides positive effects also on other prevalent “beyond-accuracy” metrics.
Alex Martinez, Federico Cinus, Francesco Bonchi, Jordi Vitrià
ACM Trans. Intell. Syst. Technol.4
2024 Overcoming Diverse Undesired Effects in Recommender Systems: A Deontological Approach
abstract
In today’s digital landscape, recommender systems have gained ubiquity as a means of directing users toward personalized products, services, and content. However, despite their widespread adoption and a long track of research, these systems are not immune to shortcomings. A significant challenge faced by recommender systems is the presence of biases, which produces various undesirable effects, prominently the popularity bias. This bias hampers the diversity of recommended items, thus restricting users’ exposure to less popular or niche content. Furthermore, this issue is compounded when multiple stakeholders are considered, requiring the balance of multiple, potentially conflicting objectives. In this article, we present a new approach to address a wide range of undesired consequences in recommender systems that involve various stakeholders. Instead of adopting a consequentialist perspective that aims to mitigate the repercussions of a recommendation policy, we propose a deontological approach centered around a minimal set of ethical principles. More precisely, we introduce two distinct principles aimed at avoiding overconfidence in predictions and accurately modeling the genuine interests of users. The proposed approach circumvents the need for defining a multi-objective system, which has been identified as one of the main limitations when developing complex recommenders. Through extensive experimentation, we show the efficacy of our approach in mitigating the adverse impact of the recommender from both user and item perspectives, ultimately enhancing various beyond accuracy metrics. This study underscores the significance of responsible and equitable recommendations and proposes a strategy that can be easily deployed in real-world scenarios.
Paula Gómez Duran, Pere Gilabert, Santi Seguí, Jordi Vitrià
ACM Trans. Intell. Syst. Technol.4
2023 Self-supervised out-of-distribution detection in wireless capsule endoscopy images
abstract
While deep learning has displayed excellent performance in a broad spectrum of application areas, neural networks still struggle to recognize what they have not seen, i.e., out-of-distribution (OOD) inputs. In the medical field, building robust models that are able to detect OOD images is highly critical, as these rare images could show diseases or anomalies that should be detected. In this study, we use wireless capsule endoscopy (WCE) images to present a novel patch-based self-supervised approach comprising three stages. First, we train a triplet network to learn vector representations of WCE image patches. Second, we cluster the patch embeddings to group patches in terms of visual similarity. Third, we use the cluster assignments as pseudolabels to train a patch classifier and use the Out-of-Distribution Detector for Neural Networks (ODIN) for OOD detection. The system has been tested on the Kvasir-capsule, a publicly released WCE dataset. Empirical results show an OOD detection improvement compared to baseline methods. Our method can detect unseen pathologies and anomalies such as lymphangiectasia, foreign bodies and blood with AUROC>0.6. This work presents an effective solution for OOD detection models without needing labeled images.
Arnau Quindós, Pablo Laiz, Jordi Vitrià, Santi Seguí
Artif. Intell. Medicine3
2022 Deep Non-crossing Quantiles through the Partial Derivative
abstract
Quantile Regression (QR) provides a way to approximate a single conditional quantile. To have a more informative description of the conditional distribution, QR can be merged with deep learning techniques to simultaneously estimate multiple quantiles. However, the minimisation of the QR-loss function does not guarantee non-crossing quantiles, which affects the validity of such predictions and introduces a critical issue in certain scenarios. In this article, we propose a generic deep learning algorithm for predicting an arbitrary number of quantiles that ensures the quantile monotonicity constraint up to the machine precision and maintains its modelling performance with respect to alternative models. The presented method is evaluated over several real-world datasets obtaining state-of-the-art results as well as showing that it scales to large-size data sets.
Axel Brando, Joan Gimeno, José A. Rodríguez-Serrano, Jordi Vitrià
AISTATS4
2022 Time-coherent embeddings for Wireless Capsule Endoscopy
abstract
Deep learning models thrive with high amounts of data where the classes are, usually, appropriately balanced. In medical imaging, however, we often encounter the opposite case. Wireless Capsule Endoscopy is not an exception; even if huge amounts of data could be obtained, labeling each frame of a video could take up to twelve hours for an expert physician. Those videos would show no pathologies for most patients, while a minority would have a few frames with associated pathology. Overall, there would be low amounts of data and a great unbalance. Self-supervised learning provides means to use unlabelled data to initialize models that can perform better even under the described circumstance. We propose a novel contrastive loss derived from Triplet Loss, crafted to leverage temporal information in endoscopy videos. We show that our model outperforms existing models and other contrastive methods in several tasks.
Guillem Pascual, Jordi Vitrià, Santi Seguí
ICPR2
2019 Modelling heterogeneous distributions with an Uncountable Mixture of Asymmetric Laplacians
abstract
In regression tasks, aleatoric uncertainty is commonly addressed by considering a parametric distribution of the output variable, which is based on strong assumptions such as symmetry, unimodality or by supposing a restricted shape. These assumptions are too limited in scenarios where complex shapes, strong skews or multiple modes are present. In this paper, we propose a generic deep learning framework that learns an Uncountable Mixture of Asymmetric Laplacians (UMAL), which will allow us to estimate heterogeneous distributions of the output variable and shows its connections to quantile regression. Despite having a fixed number of parameters, the model can be interpreted as an infinite mixture of components, which yields a flexible approximation for heterogeneous distributions. Apart from synthetic cases, we apply this model to room price forecasting and to predict financial operations in personal bank accounts. We demonstrate that UMAL produces proper distributions, which allows us to extract richer insights and to sharpen decision-making.
Axel Brando, José A. Rodríguez-Serrano, Jordi Vitrià, Alberto Rubio
NeurIPS3
2018 Uncertainty Modelling in Deep Networks: Forecasting Short and Noisy Series
Axel Brando, José A. Rodríguez-Serrano, Mauricio Ciprian, Roberto Maestre, Jordi Vitrià
ECML/PKDD (3)5
2016 Deep Learning Features for Wireless Capsule Endoscopy Analysis
Santi Seguí, Michal Drozdzal, Guillem Pascual, Petia Radeva, Carolina Malagelada, Fernando Azpiroz, Jordi Vitrià
CIARP7
2014 Intestinal event segmentation for endoluminal video analysis
abstract
In this paper, we tackle the problem of unsupervised segmentation of intestinal events using various motility descriptors extracted from the endoluminal video. The segments of constant intestinal activity are detected with a robust statistical test that is based on Hoeffding's inequality. The qualitative analysis of the results shows that the segments adjust well to the motility information and thus, this segmentation can constitute basic units for higher-level motility description.
Michal Drozdzal, Jordi Vitrià, Santi Seguí, Carolina Malagelada, Fernando Azpiroz, Petia Radeva
ICIP2
2014 Detection of Wrinkle Frames in Endoluminal Videos Using Betweenness Centrality Measures for Images
abstract
Intestinal contractions are one of the most important events to diagnose motility pathologies of the small intestine. When visualized by wireless capsule endoscopy (WCE), the sequence of frames that represents a contraction is characterized by a clear wrinkle structure in the central frames that corresponds to the folding of the intestinal wall. In this paper, we present a new method to robustly detect wrinkle frames in full WCE videos by using a new mid-level image descriptor that is based on a centrality measure proposed for graphs. We present an extended validation, carried out in a very large database, that shows that the proposed method achieves state-of-the-art performance for this task.
Santi Seguí, Michal Drozdzal, Ekaterina Zaytseva, Carolina Malagelada, Fernando Azpiroz, Petia Radeva, Jordi Vitrià
IEEE J. Biomed. Health Informatics7
2013 Bagged One-Class Classifiers in the Presence of Outliers
abstract
The problem of training classifiers only with target data arises in many applications where nontarget data are too costly, difficult to obtain, or not available at all. Several one-class classification methods have been presented to solve this problem, but most of the methods are highly sensitive to the presence of outliers in the target class. Ensemble methods have therefore been proposed as a powerful way to improve the classification performance of binary/multi-class learning algorithms by introducing diversity into classifiers. However, their application to one-class classification has been rather limited. In this paper, we present a new ensemble method based on a nonparametric weighted bagging strategy for one-class classification, to improve accuracy in the presence of outliers. While the standard bagging strategy assumes a uniform data distribution, the method we propose here estimates a probability density based on a forest structure of the data. This assumption allows the estimation of data distribution from the computation of simple univariate and bivariate kernel densities. Experiments using original and noisy versions of 20 different datasets show that bagging ensemble methods applied to different one-class classifiers outperform base one-class classification methods. Moreover, we show that, in noisy versions of the datasets, the nonparametric weighted bagging strategy we propose outperforms the classical bagging strategy in a statistically significant way.
Santi Seguí, Laura Igual, Jordi Vitrià
Int. J. Pattern Recognit. Artif. Intell.3
2012 Human Relative Position Detection Based on Mutual Occlusion
Víctor Borjas, Michal Drozdzal, Petia Radeva, Jordi Vitrià
CIARP4
2012 Sketchable Histograms of Oriented Gradients for Object Detection
Ekaterina Zaytseva, Santi Seguí, Jordi Vitrià
CIARP3
2012 A search based approach to non maximum suppression in face detection
abstract
Face detectors typically produce a large number of false positives and this leads to the need to have a further non maximum suppression stage to eliminate multiple and spurious responses. This stage is based on considering spatial heuristics: true positive responses are selected by implicitly considering several restrictions on the spatial distribution of detector responses in natural images. In this paper we analyze the limitations of this approach and propose an efficient search method to overcome them. Results show how the application of this new non-maximum suppression approach to a simple face detector boosts its performance to state of the art results.
Ekaterina Zaytseva, Jordi Vitrià
ICIP2
2012 An Integrated Approach to Contextual Face Detection
Santi Seguí, Michal Drozdzal, Petia Radeva, Jordi Vitrià
ICPRAM (2)4
2012 Selected papers from Iberian Conference on Pattern Recognition and Image Analysis
Mario Hernández-Tejera, J. Miguel Sanches, Jordi Vitrià
Pattern Recognit.3
2012 Minimal design of error-correcting output codes
Miguel Ángel Bautista 0001, Sergio Escalera, Xavier Baró, Petia Radeva, Jordi Vitrià, Oriol Pujol
Pattern Recognit. Lett.5
2012 Categorization and Segmentation of Intestinal Content Frames for Wireless Capsule Endoscopy
abstract
Wireless capsule endoscopy (WCE) is a device that allows the direct visualization of gastrointestinal tract with minimal discomfort for the patient, but at the price of a large amount of time for screening. In order to reduce this time, several works have proposed to automatically remove all the frames showing intestinal content. These methods label frames as {intestinal content- clear} without discriminating between types of content (with different physiological meaning) or the portion of image covered. In addition, since the presence of intestinal content has been identified as an indicator of intestinal motility, its accurate quantification can show a potential clinical relevance. In this paper, we present a method for the robust detection and segmentation of intestinal content in WCE images, together with its further discrimination between turbid liquid and bubbles. Our proposal is based on a twofold system. First, frames presenting intestinal content are detected by a support vector machine classifier using color and textural information. Second, intestinal content frames are segmented into {turbid, bubbles, and clear} regions. We show a detailed validation using a large dataset. Our system outperforms previous methods and, for the first time, discriminates between turbid from bubbles media.
Santi Seguí, Michal Drozdzal, Fernando Vilariño, Carolina Malagelada, Fernando Azpiroz, Petia Radeva, Jordi Vitrià
IEEE Trans. Inf. Technol. Biomed.7
2011 Predicting dominance judgements automatically: A machine learning approach
abstract
The amount of multimodal devices that surround us is growing everyday. In this context, human interaction and communication have become a focus of attention and a hot topic of research. A crucial element in human relations is the evaluation of individuals with respect to facial traits, what is called a first impression. Studies based on appearance have suggested that personality can be expressed by appearance and the observer may use such information to form judgments. In the context of rapid facial evaluation, certain personality traits seem to have a more pronounced effect on the relations and perceptions inside groups. The perception of dominance has been shown to be an active part of social roles at different stages of life, and even play a part in mate selection. The aim of this paper is to study to what extent this information is learnable from the point of view of computer science. Specifically we intend to determine if judgments of dominance can be learned by machine learning techniques. We implement two different descriptors in order to assess this. The first is the histogram of oriented gradients (HOG), and the second is a probabilistic appearance descriptor based on the frequencies of grouped binary tests. State of the art classification rules validate the performance of both descriptors, with respect to the prediction task. Experimental results show that machine learning techniques can predict judgments of dominance rather accurately (accuracies up to 90%) and that the HOG descriptor may characterize appropriately the information necessary for such task.
Mario Rojas Quiñones, David Masip, Jordi Vitrià
FG3
2010 Automatic point-based facial trait judgments evaluation
abstract
Humans constantly evaluate the personalities of other people using their faces. Facial trait judgments have been studied in the psychological field, and have been determined to influence important social outcomes of our lives, such as elections outcomes and social relationships. Recent work on textual descriptions of faces has shown that trait judgments are highly correlated. Further, behavioral studies suggest that two orthogonal dimensions, valence and dominance, can describe the basis of the human judgments from faces. In this paper, we used a corpus of behavioral data of judgments on different trait dimensions to automatically learn a trait predictor from facial pixel images. We study whether trait evaluations performed by humans can be learned using machine learning classifiers, and used later in automatic evaluations of new facial images. The experiments performed using local point-based descriptors show promising results in the evaluation of the main traits.
Mario Rojas Quiñones, David Masip, Alexander Todorov, Jordi Vitrià
CVPR4
2010 Online pattern recognition and machine learning techniques for computer-vision: Theory and applications
Bogdan Raducanu, Jordi Vitrià, Ales Leonardis
Image Vis. Comput.2
2010 Intestinal Motility Assessment With Video Capsule Endoscopy: Automatic Annotation of Phasic Intestinal Contractions
abstract
Intestinal motility assessment with video capsule endoscopy arises as a novel and challenging clinical fieldwork. This technique is based on the analysis of the patterns of intestinal contractions shown in a video provided by an ingestible capsule with a wireless micro-camera. The manual labeling of all the motility events requires large amount of time for offline screening in search of findings with low prevalence, which turns this procedure currently unpractical. In this paper, we propose a machine learning system to automatically detect the phasic intestinal contractions in video capsule endoscopy, driving a useful but not feasible clinical routine into a feasible clinical procedure. Our proposal is based on a sequential design which involves the analysis of textural, color, and blob features together with SVM classifiers. Our approach tackles the reduction of the imbalance rate of data and allows the inclusion of domain knowledge as new stages in the cascade. We present a detailed analysis, both in a quantitative and a qualitative way, by providing several measures of performance and the assessment study of interobserver variability. Our system performs at 70% of sensitivity for individual detection, whilst obtaining equivalent patterns to those of the experts for density of contractions.
Fernando Vilariño, Panagiota Spyridonos, Fosca De Iorio, Jordi Vitrià, Fernando Azpiroz, Petia Radeva
IEEE Trans. Medical Imaging4
2009 You are fired! Nonverbal role analysis in competitive meetings
abstract
This paper addresses the problem of social interaction analysis in competitive meetings, using nonverbal cues. For our study, we made use of ldquoThe Apprenticerdquo reality TV show, which features a competition for a real, highly paid corporate job. Our analysis is centered around two tasks regarding a person's role in a meeting: predicting the person with the highest status and predicting the fired candidates. The current study was carried out using nonverbal audio cues. Results obtained from the analysis of a full season of the show, representing around 90 minutes of audio data, are very promising (up to 85.7% of accuracy in the first case and up to 92.8% in the second case). Our approach is based only on the nonverbal interaction dynamics during the meeting without relying on the spoken words.
Bogdan Raducanu, Jordi Vitrià, Daniel Gatica-Perez
ICASSP2
2009 Visual content layer for scalable object recognition in urban image databases
abstract
Rich online map interaction represents a useful tool to get multimedia information related to physical places. With this type of systems, users can automatically compute the optimal route for a trip or to look for entertainment places or hotels near their actual position. Standard maps are defined as a fusion of layers, where each one contains specific data such height, streets, or a particular business location. In this paper we propose the construction of a visual content layer which describes the visual appearance of geographic locations in a city. We captured, by means of a mobile mapping system, a huge set of georeferenced images (> 500 K) which cover the whole city of Barcelona. For each image, hundreds of region descriptions are computed off-line and described as a hash code. This allows an efficient and scalable way of accessing maps by visual content.
Xavier Baró, Sergio Escalera, Petia Radeva, Jordi Vitrià
ICME4
2009 Traffic Sign Recognition Using Evolutionary Adaboost Detection and Forest-ECOC Classification
abstract
The high variability of sign appearance in uncontrolled environments has made the detection and classification of road signs a challenging problem in computer vision. In this paper, we introduce a novel approach for the detection and classification of traffic signs. Detection is based on a boosted detectors cascade, trained with a novel evolutionary version of Adaboost, which allows the use of large feature spaces. Classification is defined as a multiclass categorization problem. A battery of classifiers is trained to split classes in an Error-Correcting Output Code (ECOC) framework. We propose an ECOC design through a forest of optimal tree structures that are embedded in the ECOC matrix. The novel system offers high performance and better accuracy than the state-of-the-art strategies and is potentially better in terms of noise, affine deformation, partial occlusions, and reduced illumination.
Xavier Baró, Sergio Escalera, Jordi Vitrià, Oriol Pujol, Petia Radeva
IEEE Trans. Intell. Transp. Syst.3
2009 Boosted Online Learning for Face Recognition
abstract
Face recognition applications commonly suffer from three main drawbacks: a reduced training set, information lying in high-dimensional subspaces, and the need to incorporate new people to recognize. In the recent literature, the extension of a face classifier in order to include new people in the model has been solved using online feature extraction techniques. The most successful approaches of those are the extensions of the principal component analysis or the linear discriminant analysis. In the current paper, a new online boosting algorithm is introduced: a face recognition method that extends a boosting-based classifier by adding new classes while avoiding the need of retraining the classifier each time a new person joins the system. The classifier is learned using the multitask learning principle where multiple verification tasks are trained together sharing the same feature space. The new classes are added taking advantage of the structure learned previously, being the addition of new classes not computationally demanding. The present proposal has been (experimentally) validated with two different facial data sets by comparing our approach with the current state-of-the-art techniques. The results show that the proposed online boosting algorithm fares better in terms of final accuracy. In addition, the global performance does not decrease drastically even when the number of classes of the base problem is multiplied by eight.
David Masip, Àgata Lapedriza, Jordi Vitrià
IEEE Trans. Syst. Man Cybern. Part B3
2008 On the use of independent tasks for face recognition
abstract
We present a method for learning discriminative linear feature extraction using independent tasks. More concretely, given a target classification task, we consider a complementary classification task that is independent of the target one. For example, in face classification field, subject recognition can be a target task while facial expression classification can be a complementary task. Then, we use labels of the complementary task in order to obtain a more robust feature extraction, being the new feature space less sensitive to the complementary classification. To learn the proposed feature extraction we use the mutual information measure between the projected data and both labels from the target and the complementary tasks. In our experiments, this framework has been applied to a face recognition problem, in order to inhibit this classification task from environmental artifacts, and to mitigate the effects of the small sample size problem. Our classification experiments show an improved feature extraction process using the proposed method.
Àgata Lapedriza, David Masip, Jordi Vitrià
CVPR3
2008 Weighted Dissociated Dipoles: An Extended Visual Feature Set
Xavier Baró, Jordi Vitrià
ICVS2
2008 Diagnostic System for Intestinal Motility Disfunctions Using Video Capsule Endoscopy
Santi Seguí, Laura Igual, Fernando Vilariño, Petia Radeva, Carolina Malagelada, Fernando Azpiroz, Jordi Vitrià
ICVS7
2008 Face Recognition by Artificial Vision Systems: a Cognitive Perspective
abstract
Cognitive development refers to the ability of a system to gradually acquire knowledge through experiences during its existence. As a consequence, the learning strategy should be represented as an integrated, online process that aims to build a model of the "world" and a continuous update of this model. Considering as reference the Modal Model of Memory introduced by Atkinson and Schiffrin, we propose an online learning algorithm for cognitive systems design. The incremental part of the algorithm is responsible of updating existing information or creating new data categories and the decremental part, to efficiently evaluate the system's performance facing partial or total loss of data. The proposed algorithm has been applied to the face recognition problem. More generally, the current approach can be extended to large-scale classification problems, to limit the memory requirements for optimal data representation and storage.
Bogdan Raducanu, Jordi Vitrià
Int. J. Pattern Recognit. Artif. Intell.2
2008 A sparse Bayesian approach for joint feature selection and classifier learning
Àgata Lapedriza, Santi Seguí, David Masip, Jordi Vitrià
Pattern Anal. Appl.4
2008 Non-parametric distance-based classification techniques and their applications
Filiberto Pla, Petia Radeva, Jordi Vitrià
Pattern Anal. Appl.3
2008 Online nonparametric discriminant analysis for incremental subspace learning and recognition
Bogdan Raducanu, Jordi Vitrià
Pattern Anal. Appl.2
2008 Learning to learn: From smart machines to intelligent machines
Bogdan Raducanu, Jordi Vitrià
Pattern Recognit. Lett.2
2008 Shared Feature Extraction for Nearest Neighbor Face Recognition
abstract
In this paper, we propose a new supervised linear feature extraction technique for multiclass classification problems that is specially suited to the nearest neighbor classifier (NN). The problem of finding the optimal linear projection matrix is defined as a classification problem and the Adaboost algorithm is used to compute it in an iterative way. This strategy allows the introduction of a multitask learning (MTL) criterion in the method and results in a solution that makes no assumptions about the data distribution and that is specially appropriated to solve the small sample size problem. The performance of the method is illustrated by an application to the face recognition problem. The experiments show that the representation obtained following the multitask approach improves the classic feature extraction algorithms when using the NN classifier, especially when we have a few examples from each class.
David Masip, Jordi Vitrià
IEEE Trans. Neural Networks2
2007 Eigenmotion-Based Detection of Intestinal Contractions
Laura Igual, Santi Seguí, Jordi Vitrià, Fernando Azpiroz, Petia Radeva
CAIP3
2007 A Semi-supervised Learning Method for Motility Disease Diagnostic
Santi Seguí, Laura Igual, Petia Radeva, Carolina Malagelada, Fernando Azpiroz, Jordi Vitrià
CIARP6
2007 Online Learning for Human-Robot Interaction
abstract
This paper presents a novel approach for incremental subspace learning based on an online version of the non-parametric discriminant analysis (NDA). For many real-world applications (like the study of visual processes, for instance) there is impossible to know beforehand the number of total classes or the exact number of instances per class. This motivated us to propose a new algorithm, in which new samples can be added asynchronously, at different time stamps, as soon as they become available. The proposed technique for NDA-eigenspace representation has been applied to the problem of online face recognition for human-robot interaction scenario.
Bogdan Raducanu, Jordi Vitrià
CVPR2
2007 Incremental On-Line Topological Map Learning for A Visual Homing Application
abstract
In this paper we propose an on-line incremental vision-based topological map learning for AIBO robots. The topological map is represented through a graph, where the vertices encode views from the robot's environment and the edges the spatial relationship between these views. The views are represented through their SIFT keypoints. The proposed map learning method has been successfully applied to a homing application.
Elvina Motard, Bogdan Raducanu, Viviane Cadenat, Jordi Vitrià
ICRA4
2007 Bayesian Classification of Cork Stoppers Using Class-Conditional Independent Component Analysis
abstract
In this paper, a real-time application for visual inspection and classification of cork stoppers is presented. The process of cork inspection and quality grading is based on analyzing a large set of characteristics corresponding to visual features that are related to cork porosity. We have applied a set of nonparametric and parametric classification methods for comparing and evaluating their performance in this real problem. The best results have been achieved using Bayesian classification through probabilistic modeling in a high-dimensional space. In this context, it is well known that high dimensionality represents a serious problem for density estimation. We propose a class-conditional independent component analysis representation of the data that allows an accurate estimation of the data probability density function by factorizing it. The method has achieved a success of 98% of correct classification
Jordi Vitrià, Marco Bressan 0001, Petia Radeva
IEEE Trans. Syst. Man Cybern. Part C1
2006 Linear Radial Patterns Characterization for Automatic Detection of Tonic Intestinal Contractions
Fernando Vilariño, Panagiota Spyridonos, Jordi Vitrià, Carolina Malagelada, Petia Radeva
CIARP3
2006 A Machine Learning Framework Using SOMs: Applications in the Intestinal Motility Assessment
Fernando Vilariño, Panagiota Spyridonos, Jordi Vitrià, Carolina Malagelada, Petia Radeva
CIARP3
2006 Anisotropic Feature Extraction from Endoluminal Images for Detection of Intestinal Contractions
Panagiota Spyridonos, Fernando Vilariño, Jordi Vitrià, Fernando Azpiroz, Petia Radeva
MICCAI (2)3
2006 Discriminant ECOC: A Heuristic Method for Application Dependent Design of Error Correcting Output Codes
abstract
We present a heuristic method for learning error correcting output codes matrices based on a hierarchical partition of the class space that maximizes a discriminative criterion. To achieve this goal, the optimal codeword separation is sacrificed in favor of a maximum class discrimination in the partitions. The creation of the hierarchical partition set is performed using a binary tree. As a result, a compact matrix with high discrimination power is obtained. Our method is validated using the UCI database and applied to a real problem, the classification of traffic sign images.
Oriol Pujol, Petia Radeva, Jordi Vitrià
IEEE Trans. Pattern Anal. Mach. Intell.3
2006 Boosted discriminant projections for nearest neighbor classification
David Masip, Jordi Vitrià
Pattern Recognit.2
2005 Identification of Intestinal Motility Events of Capsule Endoscopy Video Analysis
Panagiota Spyridonos, Fernando Vilariño, Jordi Vitrià, Petia Radeva
ACIVS3
2005 Feature extraction for nearest neighbor classification: Application to gender recognition
abstract
In this article, we perform an extended analysis of different face-processing techniques for gender recognition problems. Prior research works show that support vector machines (SVM) achieve the best classification results. We will show that a nearest neighbor classification approach can reach a similar performance or improve the SVM results, given an adequate selection of features of the input data. This selection is performed using a dimensionality reduction technique based on a modification of nonparametric discriminant analysis, designed to improve the nearest neighbor classification. The choice of nearest neighbor is especially justified by the use of a large database. We also analyze a nonlinear algorithm, locally linear embedding, and its supervised version. Given that this technique is focused on preserving the local configuration of the neighborhood of each point, it should be a priori a good dimensionality reduction technique for extracting good features for nearest neighbor classification. A complete comparative study with the most classical face-processing techniques is also performed. © 2005 Wiley Periodicals, Inc. Int J Int Syst 20: 561–576, 2005.
David Masip, Jordi Vitrià
Int. J. Intell. Syst.2
2005 An ensemble-based method for linear feature extraction for two-class problems
David Masip, Ludmila I. Kuncheva, Jordi Vitrià
Pattern Anal. Appl.3
2004 Adaboost to Classify Plaque Appearance in IVUS Images
Oriol Pujol, Petia Radeva, Jordi Vitrià, Josepa Mauri
CIARP3
2004 Discriminant Projections Embedding for Nearest Neighbor Classification
Petia Radeva, Jordi Vitrià
CIARP2
2004 Multiclass Object Recognition Using Class-Conditional Independent Component Analysis
abstract
A new modeling technique, based on independent component analysis (ICA), is proposed to represent and recognize high-dimensional samples from a large set of classes. The model is constructed via density estimation techniques, and recognition is performed in the Bayesian decision framework. We show that the technique can be successfully used for automatic object identification in environments where a visual observer is faced with a classification problem in high-dimensional spaces with a large number of classes. A first experiment illustrates that classification using an ICA representation is a technique that, even in low dimensions, performs comparably to standard classification techniques. The second experiment tests the ICA classification model on high-dimensional data. Recognition was performed using local color histograms of images corresponding to 400 different objects. It is also shown how our approach outperforms other techniques commonly used in the context of appearance-based recognition.
Marco Bressan 0001, David Guillamet, Jordi Vitrià
Cybern. Syst.3
2003 Local Appearance-Based Models using High-Order Statistics of Image Features
abstract
We propose a novel local appearance modeling method for object detection and recognition in cluttered scenes. The approach is based on the joint distribution of local feature vectors at multiple salient points and factorization with the independent component analysis (ICA). The resulting densities are simple multiplicative distributions modeled through adaptive Gaussian mixture models. This leads to computationally tractable joint probability densities, which can model high-order dependencies. Furthermore, different models are compared based on appearance, color and geometry information. Also, the combination of all of them results in a hybrid model, which obtains the best results using the COIL-100 object database. Our technique has been tested under different natural and cluttered scenes with different degrees of occlusions with promising results. Finally, a large statistical test with the MNIST digit database is used to demonstrate the improved performance obtained by explicit modeling of high-order dependencies.
Baback Moghaddam, David Guillamet, Jordi Vitrià
CVPR (1)3
2003 Higher-order dependencies in local appearance models
abstract
A novel local appearance modeling method for object detection and recognition in cluttered scenes. The approach is based on the joint distribution of local feature vectors at multiple salient points and their factorization with independent component analysis (ICA). The resulting densities are simple multiplicative distributions modeled through adaptative Gaussian mixture models. This leads to computationally tractable joint probability densities which can model high-order dependencies. Our technique has been initially tested under different natural and cluttered scenes with different degrees of occlusions yielding promising results. In this work, we provide a large statistical test with the MNIST digit database in order to demonstrate the improved performance obtained by explicit modeling of higher-order dependencies.
David Guillamet, Baback Moghaddam, Jordi Vitrià
ICIP (1)3
2003 On the Selection and Classification of Independent Features
abstract
This paper is focused on the problems of feature selection and classification when classes are modeled by statistically independent features. We show that, under the assumption of class-conditional independence, the class separability measure of divergence is greatly simplified, becoming a sum of unidimensional divergences, providing a feature selection criterion where no exhaustive search is required. Since the hypothesis of independence is infrequently met in practice, we also provide a framework making use of class-conditional Independent Component Analyzers where this assumption can be held on stronger grounds. Divergence and the Bayes decision scheme are adapted to this class-conditional representation. An algorithm that integrates the proposed representation, feature selection technique, and classifier is presented. Experiments on artificial, benchmark, and real-world data illustrate our technique and evaluate its performance.
Marco Bressan 0001, Jordi Vitrià
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Using an ICA representation of local color histograms for object recognition
Marco Bressan 0001, David Guillamet, Jordi Vitrià
Pattern Recognit.3
2003 Nonparametric discriminant analysis and nearest neighbor classification
Marco Bressan 0001, Jordi Vitrià
Pattern Recognit. Lett.2
2003 Evaluation of distance metrics for recognition based on non-negative matrix factorization
David Guillamet, Jordi Vitrià
Pattern Recognit. Lett.2
2003 Introducing a weighted non-negative matrix factorization for image classification
David Guillamet, Jordi Vitrià, Bernt Schiele
Pattern Recognit. Lett.2
2002 Shot Partitioning Based Recognition of TV Commercials
Juan María Sánchez, Xavier Binefa, Jordi Vitrià
Multim. Tools Appl.3
2001 Unsupervised Learning of Part-Based Representations
David Guillamet, Jordi Vitrià
CAIP2
2001 Using an ICA Representation of High Dimensional Data for Object Recognition and Classification
abstract
This paper applies a Bayesian classification scheme to the problem of recognition through probabilistic modeling of high dimensional data. In this context, high dimensionality does not allow precision in the density estimation. We propose a local independent component analysis (ICA) representation of the data. The components can be assumed statistically independent and, in many cases, sparsity is observed. We show how these two characteristics can be used to simplify and add accuracy to the density estimation and develop bayesian decision within this representation. A first experiment illustrates that classification using an ICA representation is a technique that, even in low dimensions, performs comparably to standard classification techniques. The second experiment tests the ICA classification model on high dimensional data. Recognition was performed using local color histograms as salient features. It is also shown how our approach outperforms other techniques commonly used in the context of appearance-based recognition.
Marco Bressan 0001, David Guillamet, Jordi Vitrià
CVPR (1)3
2001 A Weighted Non-Negative Matrix Factorization for Local Representations
abstract
The paper presents an improvement of the classical Non-negative Matrix Factorization (NMF) approach for dealing with local representations of image objects. NHF, when applied to global data representations such as faces presents a high ability to represent local features of the original data in an unsupervised way. However, when applied to local representations, NMF generates redundant basis. This work implements an improvement on the original NMF approach by incorporating prior knowledge in the form of a weight matrix extracted from the training data. A detailed mathematical description of the inclusion of this weight matrix is provided, and results demonstrating its advantages are included. Furthermore, the original NMF approach lacks a hierarchy of the elements of the estimated basis. A technique to determine an ordered set of discriminant basis is also presented. Finally, the effectiveness of the weighted approach with respect to the classical approach is experimentally compared. This is done by implementing a clustering algorithm that automatically extracts object parts from the NMF representation of an image database corresponding to newspapers.
David Guillamet, Marco Bressan 0001, Jordi Vitrià
CVPR (1)3
2001 Region-based approach for discriminant snakes
abstract
This paper proposes a statistical framework for segmenting textured areas over real images by discriminant snakes. Our active contour model has the ability to learn different texture prototypes and generate a global statistical model from a multi-valued function. This function is generated by means of filter responses over the texture regions. Linear discriminant analysis is performed to obtain a statistical classifier embodied into the snake scheme. Given an input image composed of different texture types, a likelihood map is built and the discriminant snake deforms on it to delineate regions with similar texture descriptions according to the learned texture patterns. Our method is tested on two different image applications: aerial images and medical (ultrasound) images, and the results are very encouraging.
Jordi Vitrià, Petia Radeva
ICIP (2)1
2001 Topological principal component analysis for face encoding and recognition
Albert Pujol, Jordi Vitrià, Felipe Lumbreras, Juan José Villanueva
Pattern Recognit. Lett.2
2001 Clustering in image space for place recognition and visual annotations for human-robot interaction
abstract
The most classical way of attempting to solve the vision-guided navigation problem for autonomous robots corresponds to the use of three-dimensional (3-D) geometrical descriptions of the scene; what is known as model-based approaches. However, these approaches do not facilitate the user's task because they require that geometrically precise models of the 3-D environment be given by the user. In this paper, we propose the use of "annotations" posted on some type of blackboard or "descriptive" map to facilitate this user-robot interaction. We show that, by using this technique, user commands can be as simple as "go to label 5." To build such a mechanism, new approaches for vision-guided mobile robot navigation have to be found. We show that this can be achieved by using mixture models within an appearance-based paradigm. Mixture models are more useful in practice than other pattern recognition methods such as principal component analysis (PCA) or Fisher discriminant analysis (FDA)-also known as linear discriminant analysis (LDA), because they can represent nonlinear subspaces. However, given the fact that mixture models are usually learned using the expectation-maximization (EM) algorithm which is a gradient ascent technique, the system cannot always converge to a desired final solution, due to the local maxima problem. To resolve this, a genetic version of the EM algorithm is used. We then show the capabilities of this latest approach on a navigation task that uses the above described "annotations."
Aleix Martinez, Jordi Vitrià
IEEE Trans. Syst. Man Cybern. Part B2
2000 Tracking of Elongated Structures Using Statistical Snakes
abstract
In this paper we introduce a statistic snake that learns and tracks image features by means of statistic learning techniques. Using probabilistic principal component analysis a feature description is obtained from a training set of object profiles. In our approach a sound statistical model is introduced to define a likelihood estimate of the grey-level local image profiles together with their local orientation. This likelihood estimate allows to define a probabilistic potential field of the snake where the elastic curve deforms to maximise the overall probability of detecting learned image features. To improve the convergence of snake deformation, we enhance the likelihood map by a physics-based model simulating a dipole-dipole interaction. A new extended local coherent interaction is introduced defined in terms of extended structure tensor of the image to give priority to parallel coherence vectors.
Ricardo Toledo, Xavier Orriols, Xavier Binefa, Petia Radeva, Jordi Vitrià, Juan José Villanueva
CVPR5
2000 A Comparison of Global versus Local Color Histograms for Object Recognition
abstract
Global color distributions have been efficiently used as signatures for object recognition. However, these methods are very sensitive to partial occlusions and to background regions. Our approach is directed to minimize these effects by working with small neighborhoods. We compare global and local color representations on an automatic object recognition system. Local representations significantly outperformed global representations in terms of recognition rates. Local color distributions are a strong constraint when objects consist of distinctive local regions. Eigenspace techniques are applied to detect discriminant local representations and support vector machines are used during the recognition process in order to maximize the recognition rate.
David Guillamet, Jordi Vitrià
ICPR2
2000 Probabilistic Saliency Approach for Elongated Structure Detection Using Deformable Models
abstract
We address the object recognition problem in a probabilistic framework to detect and describe object appearance through image features organized by means of active contour models. We consider the formulation of saliency in terms of visual similarity embedded in the probabilistic principal component analysis framework. A likelihood of object structure detection is obtained using the relation between the visual field and the internal object representation. Deformable models are employed introducing a computational methodology for a perceptual organisation of image features as an abstract understanding of the integration between structure and constraints of the visual information-processing problem. A specific application of the integrated approach for vessels segmentation in angiography is considered and the results are encouraging.
Xavier Orriols, Ricardo Toledo, Xavier Binefa, Petia Radeva, Jordi Vitrià, Juan José Villanueva
ICPR5
2000 Eigensnakes for Vessel Segmentation in Angiography
abstract
We introduce a new deformable model, called eigensnake, for segmentation of elongated structures in a probabilistic framework. Instead of snake attraction by specific image features extracted independently of the snake, our eigensnake learns an optimal object description and searches for such image feature in the target image. This is achieved applying principal component analysis on image responses of a bank of Gaussian derivative filters. Therefore, attraction by eigensnakes is defined in terms of classification of image features. The potential energy for the snake is defined in terms of likelihood in the feature space and incorporated into a new energy minimising scheme. Hence, the snake deforms to minimise the mahalanobis distance in the feature space. A real application of segmenting and tracking coronary vessels in angiography is considered and the results are very encouraging.
Ricardo Toledo, Xavier Orriols, Petia Radeva, Xavier Binefa, Jordi Vitrià, Cristina Cañero Morales, Juan José Villanueva
ICPR5
2000 Eigenfiltering for Flexible Eigentracking (EFE)
abstract
Traditional techniques for tracking nonrigid objects such as optical flow, correlation, active contours or color, cannot deal with situations where image changes are not due to motion but appearance (e.g. tracking the lips when the teeth appear). Two main contributions for appearance tracking of flexible objects are proposed. The first one is a flexible generalization of eigentracking within the same robust continuous optimization framework. The second one is a generalization of traditional graylevel eigenspaces, constructing a multiple channel "eigenspace" using filter responses to give robustness against variations in the training conditions such as illumination changes. Additionally, 3D geometric transformations are incorporated, a regularization term is added for numerical stability reasons and the optimization problem is solved in closed form. Experiments on lip tracking are reported.
Fernando De la Torre, Javier Melenchón, Jordi Vitrià, Petia Radeva
ICPR3
2000 Learning mixture models using a genetic version of the EM algorithm
Aleix Martinez, Jordi Vitrià
Pattern Recognit. Lett.2
1999 EigenHistograms: Using Low Dimensional Models of Color Distribution for Real Time Object Recognition
Jordi Vitrià, Petia Radeva, Xavier Binefa
CAIP1
1997 Optimal 3 x 3 decomposable disks for morphological transformations
María Vanrell 0001, Jordi Vitrià
Image Vis. Comput.2
1997 A multidimensional scaling approach to explore the behavior of a texture perception algorithm
María Vanrell 0001, Jordi Vitrià, F. Xavier Roca
Mach. Vis. Appl.2
1996 3×3 decomposition of circular structuring elements
abstract
This paper presents some results to decompose circular structuring elements into 3/spl times/3 elements. Decomposition allows one to improve the expended time in computing morphological operations. Generally, the shape of the structuring element determines the image transformation. Morphological operations with disks can be used as shape and size descriptors. The optimal discrete approximation of a disk can not be decomposed into 3/spl times/3 factors. Therefore, for a given radius, we give a hexadecagon that can be decomposed, and which optimally fits a disk. Afterwards, we present the decomposition of the disk in terms of the hexadecagon parameters. The decomposition prime factors can be given by different families of basic structuring elements.
María Vanrell 0001, Jordi Vitrià
ICIP (3)2
1996 A contrast-based focusing criterium
abstract
In this paper a new criterion for determining the optimal focus setting during the image acquisition process is presented. It is based upon a new representation of the focal variation of contrast between images: the dynamic contrast histogram (DCH). The DCH of a point is a scaled joint probability density function between the gray levels of two images of the same scene, acquired with different focus settings in a neighbourhood of that point. The dynamic focusing criterion (DFC) developed here computes the slope of the principal axis of the DCH as a focusing criterion. The results of this technique are compared to other methods based on a quality measure of its focusing ability, its sensitivity to noise, and its computational complexity. Results are presented from microscopic images.
Xavier Binefa, Jordi Vitrià
ICPR2
1996 ViLi (Vision LISP): a software environment for teaching image processing and analysis
Jordi Vitrià
ITiCSE2
1996 Reconstructing 3D light microscopic images using the EM algorithm
Jordi Vitrià, Jorge Llacer
Pattern Recognit. Lett.1
1992 Morphological algorithms for visual analysis of integrated circuits
Jordi Vitrià, Xavier Binefa, Juan José Villanueva
J. Vis. Commun. Image Represent.1
1990 Analysis of x-ray hand images for bone age assessment
abstract
In this paper we describe a model-based system for the assessment of skeletal maturity on hand radiographs by the TW2 method. The problem consists in classifiying a set of bones appearing in an image, in one of several stages described in an atlas. A first approach consisting in pre-processing, segmentation and classification independent phases is also presented. However, it is only well suited for well contrasted, low noise images, without superimposed bones, were the edge detection by zero crossing of second directional derivatives is able to extract all bone contours, maybe with little gaps, and few false edges on the background. Hence, the use of all available knowledge about the problem domain is needed to build a rather general system. We have designed a rule-based system for narrow down the rank of possible stages for each bone, and guide the analysis process. It calls procedures written in conventional languages for matching stage models against the image, and getting features needed in the classification process.
Joan Serrat 0002, Jordi Vitrià, Juan José Villanueva
VCIP2