Xavier Boix

dblp:87/8652 · also Xavier Boix Bosch · DBLP profile ↗
← Back
26ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0003-4656-3485ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
15 papers
Vision and language · 25% Representation and self-supervised learning · 14% Image recognition and object detection · 13%
Computer graphics and multimedia
3 papers
Image and video processing · 56% Multimedia analysis and retrieval · 44%

Topics — the 28 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
neural module network
1.322024
Transformer Module Networks for Systematic Generalization in Visual Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2024
How Modular should Neural Module Networks Be for Systematic Generalization? · NeurIPS 2021
Machine learning › Representation and self-supervised learning
systematic generalization
1.322024
Transformer Module Networks for Systematic Generalization in Visual Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2024
How Modular should Neural Module Networks Be for Systematic Generalization? · NeurIPS 2021
Computer vision › Vision and language
visual question answering
1.322024
Transformer Module Networks for Systematic Generalization in Visual Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2024
How Modular should Neural Module Networks Be for Systematic Generalization? · NeurIPS 2021
Machine learning › Optimization for machine learning
implicit regularization
0.512021
Frivolous Units: Wider Networks Are Not Really That Wide · AAAI 2021
Machine learning › Efficient and distributed learning
model compression
0.512021
Frivolous Units: Wider Networks Are Not Really That Wide · AAAI 2021
Machine learning › Deep learning architectures and training
neural network width
0.512021
Frivolous Units: Wider Networks Are Not Really That Wide · AAAI 2021
Machine learning › Learning theory
over-parameterization
0.512021
Frivolous Units: Wider Networks Are Not Really That Wide · AAAI 2021
Machine learning › Efficient and distributed learning › model compression
pruning
0.512021
Frivolous Units: Wider Networks Are Not Really That Wide · AAAI 2021
Computer vision › Image recognition and object detection
saliency prediction
0.522016
Predicting When Saliency Maps are Accurate and Eye Fixations Consistent · CVPR 2016
SALICON: Reducing the Semantic Gap in Saliency Prediction by Adapting Deep Neural Networks · ICCV 2015
Computer vision › Segmentation and scene understanding
semantic segmentation
0.542013
Active MAP Inference in CRFs for Efficient Semantic Segmentation · ICCV 2013
Harmony Potentials - Fusing Global and Local Scale for Semantic Image Segmentation · Int. J. Comput. Vis. 2012
Harmony potentials for joint classification and segmentation · CVPR 2010
Machine learning › Trustworthy machine learning
robustness
0.412019
Minimal Images in Deep Neural Networks: Fragile Object Recognition in Natural Images · ICLR (Poster) 2019
Image and video processing
image segmentation
0.422015
SEEDS: Superpixels Extracted Via Energy-Driven Sampling · Int. J. Comput. Vis. 2015
SEEDS: Superpixels Extracted via Energy-Driven Sampling · ECCV (7) 2012
Image and video processing › image segmentation
superpixel segmentation
0.422015
SEEDS: Superpixels Extracted Via Energy-Driven Sampling · Int. J. Comput. Vis. 2015
SEEDS: Superpixels Extracted via Energy-Driven Sampling · ECCV (7) 2012
Multimedia analysis and retrieval › video summarization
interactive video summarization
0.312017
Active Video Summarization: Customized Summaries via On-line Interaction with the User · AAAI 2017
Multimedia analysis and retrieval
video summarization
0.312017
Active Video Summarization: Customized Summaries via On-line Interaction with the User · AAAI 2017
Machine learning › Trustworthy machine learning
interpretability
0.212016
Predicting When Saliency Maps are Accurate and Eye Fixations Consistent · CVPR 2016
Computer vision › 3D vision › local feature descriptor
binary descriptor
0.212013
Sparse Quantization for Patch Description · CVPR 2013
Computer vision › Segmentation and scene understanding › semantic segmentation
CRF-based segmentation
0.212013
Active MAP Inference in CRFs for Efficient Semantic Segmentation · ICCV 2013
Computer vision › 3D vision › feature matching › local feature matching
keypoint matching
0.212013
Sparse Quantization for Patch Description · CVPR 2013
Computer vision › Image recognition and object detection › object detection
objectness estimation
0.212013
Online Video SEEDS for Temporal Window Objectness · ICCV 2013
Natural language and speech › Language models and text generation
compositional generalization
0.112021
How Modular should Neural Module Networks Be for Systematic Generalization? · NeurIPS 2021
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding
0.112012
Nested Sparse Quantization for Efficient Feature Coding · ECCV (2) 2012
Computer vision › Segmentation and scene understanding
image segmentation
0.112011
Are spatial and global constraints really necessary for segmentation? · ICCV 2011
Computer vision › Image recognition and object detection › object detection › multi-view object detection
multi-camera object detection
0.112011
Conditional Random Fields for multi-camera object detection · ICCV 2011
Computer vision › Image recognition and object detection › object detection
multi-class object detection
0.112011
Conditional Random Fields for multi-camera object detection · ICCV 2011
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
conditional random field
0.112010
Harmony potentials for joint classification and segmentation · CVPR 2010
Computer vision › Segmentation and scene understanding › image segmentation › multi-task segmentation
joint segmentation and classification
0.112010
Harmony potentials for joint classification and segmentation · CVPR 2010
Usability and user experience research › user modeling
preference elicitation
0.112017
Active Video Summarization: Customized Summaries via On-line Interaction with the User · AAAI 2017

Methods — techniques the papers use, named apart from their topics

neural module networks · 1.3transformer · 0.8modular composition · 0.8question selection · 0.6active learning · 0.6unit-level analysis · 0.5modularity tuning · 0.5linear combination redundancy · 0.5conditional random field · 0.4minimal image analysis · 0.4energy-driven sampling · 0.4sparse quantization · 0.3harmony potential · 0.3
YearPublicationVenuePosition
2024 Transformer Module Networks for Systematic Generalization in Visual Question Answering
abstract
Transformers achieve great performance on Visual Question Answering (VQA). However, their systematic generalization capabilities, i.e., handling novel combinations of known concepts, is unclear. We reveal that Neural Module Networks (NMNs), i.e., question-specific compositions of modules that tackle a sub-task, achieve better or similar systematic generalization performance than the conventional Transformers, even though NMNs' modules are CNN-based. In order to address this shortcoming of Transformers with respect to NMNs, in this paper we investigate whether and how modularity can bring benefits to Transformers. Namely, we introduce Transformer Module Network (TMN), a novel NMN based on compositions of Transformer modules. TMNs achieve state-of-the-art systematic generalization performance in three VQA datasets, improving more than 30% over standard Transformers for novel compositions of sub-tasks. We show that not only the module composition but also the module specialization for each sub-task are the key of such performance gain.
Moyuru Yamada, Vanessa D'Amario, Kentaro Takemoto, Xavier Boix, Tomotake Sasaki
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Robustness to Transformations Across Categories: Is Robustness Driven by Invariant Neural Representations?
abstract
Deep convolutional neural networks (DCNNs) have demonstrated impressive robustness to recognize objects under transformations (e.g., blur or noise) when these transformations are included in the training set. A hypothesis to explain such robustness is that DCNNs develop invariant neural representations that remain unaltered when the image is transformed. However, to what extent this hypothesis holds true is an outstanding question, as robustness to transformations could be achieved with properties different from invariance; for example, parts of the network could be specialized to recognize either transformed or nontransformed images. This article investigates the conditions under which invariant neural representations emerge by leveraging that they facilitate robustness to transformations beyond the training distribution. Concretely, we analyze a training paradigm in which only some object categories are seen transformed during training and evaluate whether the DCNN is robust to transformations across categories not seen transformed. Our results with state-of-the-art DCNNs indicate that invariant neural representations do not always drive robustness to transformations, as networks show robustness for categories seen transformed during training even in the absence of invariant neural representations. Invariance emerges only as the number of transformed categories in the training set is increased. This phenomenon is much more prominent with local transformations such as blurring and high-pass filtering than geometric transformations such as rotation and thinning, which entail changes in the spatial arrangement of the object. Our results contribute to a better understanding of invariant neural representations in deep learning and the conditions under which it spontaneously emerges.
Hojin Jang, Syed Suleman Abbas Zaidi, Xavier Boix, Neeraj Prasad, Sharon Gilad-Gutnick, Shlomit Ben-Ami, Pawan Sinha
Neural Comput.3
2022 Three approaches to facilitate invariant neurons and generalization to out-of-distribution orientations and illuminations
abstract
The training data distribution is often biased towards objects in certain orientations and illumination conditions. While humans have a remarkable capability of recognizing objects in out-of-distribution (OoD) orientations and illuminations, Deep Neural Networks (DNNs) severely suffer in this case, even when large amounts of training examples are available. Neurons that are invariant to orientations and illuminations have been proposed as a neural mechanism that could facilitate OoD generalization, but it is unclear how to encourage the emergence of such invariant neurons. In this paper, we investigate three different approaches that lead to the emergence of invariant neurons and substantially improve DNNs in recognizing objects in OoD orientations and illuminations. Namely, these approaches are (i) training much longer after convergence of the in-distribution (InD) validation accuracy, i.e., late-stopping, (ii) tuning the momentum parameter of the batch normalization layers, and (iii) enforcing invariance of the neural activity in an intermediate layer to orientation and illumination conditions. Each of these approaches substantially improves the DNN's OoD accuracy (more than 20% in some cases). We report results in four datasets: two datasets are modified from the MNIST and iLab datasets, and the other two are novel (one of 3D rendered cars and another of objects taken from various controlled orientations and illumination conditions). These datasets allow to study the effects of different amounts of bias and are challenging as DNNs perform poorly in OoD conditions. Finally, we demonstrate that even though the three approaches focus on different aspects of DNNs, they all tend to lead to the same underlying neural mechanism to enable OoD accuracy gains - individual neurons in the intermediate layers become invariant to OoD orientations and illuminations. We anticipate this study to be a basis for further improvement of deep neural networks' OoD generalization performance, which is highly demanded to achieve safe and fair AI applications.
Akira Sakai, Taro Sunagawa, Spandan Madan, Kanata Suzuki, Takashi Katoh, Hiromichi Kobashi, Hanspeter Pfister, Pawan Sinha, Xavier Boix, Tomotake Sasaki
Neural Networks9
2021 Frivolous Units: Wider Networks Are Not Really That Wide
abstract
A remarkable characteristic of overparameterized deep neural networks (DNNs) is that their accuracy does not degrade when the network width is increased. Recent evidence suggests that developing compressible representations allows the complexity of large networks to be adjusted for the learning task at hand. However, these representations are poorly understood. A promising strand of research inspired from biology involves studying representations at the unit level as it offers a more granular interpretation of the neural mechanisms. In order to better understand what facilitates increases in width without decreases in accuracy, we ask: Are there mechanisms at the unit level by which networks control their effective complexity? If so, how do these depend on the architecture, dataset, and hyperparameters? We identify two distinct types of “frivolous” units that proliferate when the network’s width increases: prunable units which can be dropped out of the network without significant change to the output and redundant units whose activities can be expressed as a linear combination of others. These units imply complexity constraints as the function the network computes could be expressed without them. We also identify how the development of these units can be influenced by architecture and a number of training factors. Together, these results help to explain why the accuracy of DNNs does not degrade when width is increased and highlight the importance of frivolous units toward understanding implicit regularization in DNNs.
Stephen Casper, Xavier Boix, Vanessa D'Amario, Martin Schrimpf, Kasper Vinken, Gabriel Kreiman
AAAI2
2021 How Modular should Neural Module Networks Be for Systematic Generalization?
abstract
Neural Module Networks (NMNs) aim at Visual Question Answering (VQA) via composition of modules that tackle a sub-task. NMNs are a promising strategy to achieve systematic generalization, i.e., overcoming biasing factors in the training distribution. However, the aspects of NMNs that facilitate systematic generalization are not fully understood. In this paper, we demonstrate that the degree of modularity of the NMN have large influence on systematic generalization. In a series of experiments on three VQA datasets (VQA-MNIST, SQOOP, and CLEVR-CoGenT), our results reveal that tuning the degree of modularity, especially at the image encoder stage, reaches substantially higher systematic generalization. These findings lead to new NMN architectures that outperform previous ones in terms of systematic generalization.
Vanessa D'Amario, Tomotake Sasaki, Xavier Boix
NeurIPS3
2021 Do Neural Networks for Segmentation Understand Insideness?
abstract
The insideness problem is an aspect of image segmentation that consists of determining which pixels are inside and outside a region. Deep neural networks (DNNs) excel in segmentation benchmarks, but it is unclear if they have the ability to solve the insideness problem as it requires evaluating long-range spatial dependencies. In this letter, we analyze the insideness problem in isolation, without texture or semantic cues, such that other aspects of segmentation do not interfere in the analysis. We demonstrate that DNNs for segmentation with few units have sufficient complexity to solve the insideness for any curve. Yet such DNNs have severe problems with learning general solutions. Only recurrent networks trained with small images learn solutions that generalize well to almost any curve. Recurrent networks can decompose the evaluation of long-range dependencies into a sequence of local operations, and learning with small images alleviates the common difficulties of training recurrent networks with a large number of unrolling steps.
Kimberly Villalobos, Vilim Stih, Amineh Ahmadinejad, Shobhita Sundaram, Jamell Dozier, Andrew Francl, Frederico Azevedo, Tomotake Sasaki, Xavier Boix
Neural Comput.9
2019 Minimal Images in Deep Neural Networks: Fragile Object Recognition in Natural Images
Sanjana Srivastava, Guy Ben-Yosef, Xavier Boix
ICLR (Poster)3
2017 Active Video Summarization: Customized Summaries via On-line Interaction with the User
abstract
To facilitate the browsing of long videos, automatic video summarization provides an excerpt that represents its content. In the case of egocentric and consumer videos, due to their personal nature, adapting the summary to specific user's preferences is desirable. Current approaches to customizable video summarization obtain the user's preferences prior to the summarization process. As a result, the user needs to manually modify the summary to further meet the preferences. In this paper, we introduce Active Video Summarization (AVS), an interactive approach to gather the user's preferences while creating the summary. AVS asks questions about the summary to update it on-line until the user is satisfied. To minimize the interaction, the best segment to inquire next is inferred from the previous feedback. We evaluate AVS in the commonly used UTEgo dataset. We also introduce a new dataset for customized video summarization (CSumm) recorded with a Google Glass. The results show that AVS achieves an excellent compromise between usability and quality. In 41% of the videos, AVS is considered the best over all tested baselines, including summaries manually generated. Also, when looking for specific events in the video, AVS provides an average level of satisfaction higher than those of all other baselines after only six questions to the user.
Ana Garcia del Molino, Xavier Boix, Joo-Hwee Lim, Ah-Hwee Tan
AAAI2
2016 Predicting When Saliency Maps are Accurate and Eye Fixations Consistent
abstract
Many computational models of visual attention use image features and machine learning techniques to predict eye fixation locations as saliency maps. Recently, the success of Deep Convolutional Neural Networks (DCNNs) for object recognition has opened a new avenue for computational models of visual attention due to the tight link between visual attention and object recognition. In this paper, we show that using features from DCNNs for object recognition we can make predictions that enrich the information provided by saliency models. Namely, we can estimate the reliability of a saliency model from the raw image, which serves as a meta-saliency measure that may be used to select the best saliency algorithm for an image. Analogously, the consistency of the eye fixations among subjects, i.e. the agreement between the eye fixation locations of different subjects, can also be predicted and used by a designer to assess whether subjects reach a consensus about salient image locations.
Anna Volokitin, Michael Gygli, Xavier Boix
CVPR3
2016 Learning to Predict Sequences of Human Visual Fixations
abstract
Most state-of-the-art visual attention models estimate the probability distribution of fixating the eyes in a location of the image, the so-called saliency maps. Yet, these models do not predict the temporal sequence of eye fixations, which may be valuable for better predicting the human eye fixations, as well as for understanding the role of the different cues during visual exploration. In this paper, we present a method for predicting the sequence of human eye fixations, which is learned from the recorded human eye-tracking data. We use least-squares policy iteration (LSPI) to learn a visual exploration policy that mimics the recorded eye-fixation examples. The model uses a different set of parameters for the different stages of visual exploration that capture the importance of the cues during the scanpath. In a series of experiments, we demonstrate the effectiveness of using LSPI for combining multiple cues at different stages of the scanpath. The learned parameters suggest that the low-level and high-level cues (semantics) are similarly important at the first eye fixation of the scanpath, and the contribution of high-level cues keeps increasing during the visual exploration. Results show that our approach obtains the state-of-the-art performances on two challenging data sets: 1) OSIE data set and 2) MIT data set.
Ming Jiang 0019, Xavier Boix, Gemma Roig, Luc Van Gool, Qi Zhao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2015 Saliency Prediction with Active Semantic Segmentation
abstract
Ming Jiang1 [email protected] Xavier Boix1,3 [email protected] Juan Xu1 [email protected] Gemma Roig2,3 [email protected] Luc Van Gool3 [email protected] Qi Zhao1 [email protected] 1 Department of Electrical and Computer Engineering National University of Singapore Singapore 2 CBMM, LCSL Massachusetts Institute of Technology Istituto Italiano di Tecnologia Cambridge, MA 3 Computer Vision Laboratory ETH Zurich Switzerland
Ming Jiang 0019, Xavier Boix, Gemma Roig, Luc Van Gool, Qi Zhao 0001
BMVC2
2015 SALICON: Reducing the Semantic Gap in Saliency Prediction by Adapting Deep Neural Networks
abstract
Saliency in Context (SALICON) is an ongoing effort that aims at understanding and predicting visual attention. Conventional saliency models typically rely on low-level image statistics to predict human fixations. While these models perform significantly better than chance, there is still a large gap between model prediction and human behavior. This gap is largely due to the limited capability of models in predicting eye fixations with strong semantic content, the so-called semantic gap. This paper presents a focused study to narrow the semantic gap with an architecture based on Deep Neural Network (DNN). It leverages the representational power of high-level semantics encoded in DNNs pretrained for object recognition. Two key components are fine-tuning the DNNs fully convolutionally with an objective function based on the saliency evaluation metrics, and integrating information at different image scales. We compare our method with 14 saliency models on 6 public eye tracking benchmark datasets. Results demonstrate that our DNNs can automatically learn features particularly for saliency prediction that surpass by a big margin the state-of-the-art. In addition, our model ranks top to date under all seven metrics on the MIT300 challenge set.
Chengyao Shen, Xavier Boix, Qi Zhao 0001
ICCV3
2015 SEEDS: Superpixels Extracted Via Energy-Driven Sampling
Michael Van den Bergh, Xavier Boix, Gemma Roig, Luc Van Gool
Int. J. Comput. Vis.2
2014 Reconstruction of Inextensible Surfaces on a Budget via Bootstrapping
abstract
Many methods for 3D reconstruction of deformable surfaces from a monocular view rely on inextensibility constraints. An interesting application with commercial potential lies in augmented reality in portable and wearable devices. Such applications add an additional challenge to the 3D reconstruction, since in portable platforms the availability of resources is limited and not always guaranteed. Towards this goal, we introduce a method to deliver the best possible 3D reconstruction of the deformable surface at any time. Since computational resources may vary, it is decided on-the-fly when to stop the reconstruction algorithm. We use an efficient optimization method to quickly deliver the reconstructed surface. We introduce bootstrapping to improve the robustness of the efficient 3D reconstruction algorithm by merging multiple versions of the reconstructed surface. Also, these multiple 3D surfaces can be used to estimate the confidence of the reconstruction. In a series of experiments, in both synthetic and real data, we show that our method is effective for timely reconstruction of 3D surfaces.
Alex Locher, Lennart Elsen, Xavier Boix, Luc Van Gool
3DV3
2014 Self-Adaptable Templates for Feature Coding
Xavier Boix, Gemma Roig, Salomon Diether, Luc Van Gool
NIPS1
2013 Sparse Quantization for Patch Description
abstract
The representation of local image patches is crucial for the good performance and efficiency of many vision tasks. Patch descriptors have been designed to generalize towards diverse variations, depending on the application, as well as the desired compromise between accuracy and efficiency. We present a novel formulation of patch description, that serves such issues well. Sparse quantization lies at its heart. This allows for efficient encodings, leading to powerful, novel binary descriptors, yet also to the generalization of existing descriptors like SIFT or BRIEF. We demonstrate the capabilities of our formulation for both key point matching and image classification. Our binary descriptors achieve state-of-the-art results for two key point matching benchmarks, namely those by Brown and Mikolajczyk. For image classification, we propose new descriptors, that perform similar to SIFT on Caltech101 and PASCAL VOC07.
Xavier Boix, Michael Gygli, Gemma Roig, Luc Van Gool
CVPR1
2013 Online Video SEEDS for Temporal Window Objectness
abstract
Super pixel and objectness algorithms are broadly used as a pre-processing step to generate support regions and to speed-up further computations. Recently, many algorithms have been extended to video in order to exploit the temporal consistency between frames. However, most methods are computationally too expensive for real-time applications. We introduce an online, real-time video super pixel algorithm based on the recently proposed SEEDS super pixels. A new capability is incorporated which delivers multiple diverse samples (hypotheses) of super pixels in the same image or video sequence. The multiple samples are shown to provide a strong cue to efficiently measure the objectness of image windows, and we introduce the novel concept of objectness in temporal windows. Experiments show that the video super pixels achieve comparable performance to state-of-the-art offline methods while running at 30 fps on a single 2.8 GHz i7 CPU. State-of-the-art performance on objectness is also demonstrated, yet orders of magnitude faster and extended to temporal windows in video.
Michael Van den Bergh, Gemma Roig, Xavier Boix, Santiago Manen, Luc Van Gool
ICCV3
2013 Active MAP Inference in CRFs for Efficient Semantic Segmentation
abstract
Most MAP inference algorithms for CRFs optimize an energy function knowing all the potentials. In this paper, we focus on CRFs where the computational cost of instantiating the potentials is orders of magnitude higher than MAP inference. This is often the case in semantic image segmentation, where most potentials are instantiated by slow classifiers fed with costly features. We introduce Active MAP inference 1) to on-the-fly select a subset of potentials to be instantiated in the energy function, leaving the rest of the parameters of the potentials unknown, and 2) to estimate the MAP labeling from such incomplete energy function. Results for semantic segmentation benchmarks, namely PASCAL VOC 2010 and MSRC-21, show that Active MAP inference achieves similar levels of accuracy but with major efficiency gains.
Gemma Roig, Xavier Boix, Roderick de Nijs, Sebastian Ramos, Kolja Kühnlenz, Luc Van Gool
ICCV2
2012 SEEDS: Superpixels Extracted via Energy-Driven Sampling
Michael Van den Bergh, Xavier Boix, Gemma Roig, Benjamin de Capitani, Luc Van Gool
ECCV (7)2
2012 Nested Sparse Quantization for Efficient Feature Coding
Xavier Boix, Gemma Roig, Christian Leistner, Luc Van Gool
ECCV (2)1
2012 On-line semantic perception using uncertainty
abstract
Visual perception capabilities are still highly unreliable in unconstrained settings, and solutions might not be accurate in all regions of an image. Awareness of the uncertainty of perception is a fundamental requirement for proper high level decision making in a robotic system. Yet, the uncertainty measure is often sacrificed to account for dependencies between object/region classifiers. This is the case of Conditional Random Fields (CRFs), the success of which stems from their ability to infer the most likely world configuration, but they do not directly allow to estimate the uncertainty of the solution. In this paper, we consider the setting of assigning semantic labels to the pixels of an image sequence. Instead of using a CRF, we employ a Perturb-and-MAP Random Field, a recently introduced probabilistic model that allows performing fast approximate sampling from its probability density function. This allows to effectively compute the uncertainty of the solution, indicating the reliability of the most likely labeling in each region of the image. We report results on the CamVid dataset, a standard benchmark for semantic labeling of urban image sequences. In our experiments, we show the benefits of exploiting the uncertainty by putting more computational effort on the regions of the image that are less reliable, and use more efficient techniques for other regions, showing little decrease of performance.
Roderick de Nijs, Sebastian Ramos, Gemma Roig, Xavier Boix, Luc Van Gool, Kolja Kühnlenz
IROS4
2012 Harmony Potentials - Fusing Global and Local Scale for Semantic Image Segmentation
Xavier Boix, Josep M. Gonfaus, Joost van de Weijer 0001, Andrew D. Bagdanov, Joan Serrat 0002, Jordi Gonzàlez 0001
Int. J. Comput. Vis.1
2011 Hierarchical CRF with product label spaces for parts-based models
abstract
Non-rigid object detection is a challenging open research problem in computer vision. It is a critical part in many applications such as image search, surveillance, human-computer interaction or image auto-annotation. Most successful approaches to non-rigid object detection make use of part-based models. In particular, Conditional Random Fields (CRF) have been successfully embedded into a discriminative parts-based model framework due to its effectiveness for learning and inference (usually based on a tree structure). However, CRF-based approaches do not incorporate global constraints and only model pairwise interactions. This is especially important when modeling object classes that may have complex parts interactions (e.g. facial features or body articulations), because neglecting them yields an oversimplified model with suboptimal performance. To overcome this limitation, this paper proposes a novel hierarchical CRF (HCRF). The main contribution is to build a hierarchy of part combinations by extending the label set to a hierarchy of product label spaces. In order to keep the inference computation tractable, we propose an effective method to reduce the new label set. We test our method on two applications: facial feature detection on the Multi-PIE database and human pose estimation on the Buffy dataset.
Gemma Roig, Xavier Boix, Fernando De la Torre, Joan Serrat 0002, Carles Vilella
FG2
2011 Are spatial and global constraints really necessary for segmentation?
abstract
Many state-of-the-art segmentation algorithms rely on Markov or Conditional Random Field models designed to enforce spatial and global consistency constraints. This is often accomplished by introducing additional latent variables to the model, which can greatly increase its complexity. As a result, estimating the model parameters or computing the best maximum a posteriori (MAP) assignment becomes a computationally expensive task. In a series of experiments on the PASCAL and the MSRC datasets, we were unable to find evidence of a significant performance increase attributed to the introduction of such constraints. On the contrary, we found that similar levels of performance can be achieved using a much simpler design that essentially ignores these constraints. This more simple approach makes use of the same local and global features to leverage evidence from the image, but instead directly biases the preferences of individual pixels. While our investigation does not prove that spatial and consistency constraints are not useful in principle, it points to the conclusion that they should be validated in a larger context.
Aurélien Lucchi, Yunpeng Li 0002, Xavier Boix, Kevin Smith 0001, Pascal Fua
ICCV3
2011 Conditional Random Fields for multi-camera object detection
abstract
We formulate a model for multi-class object detection in a multi-camera environment. From our knowledge, this is the first time that this problem is addressed taken into account different object classes simultaneously. Given several images of the scene taken from different angles, our system estimates the ground plane location of the objects from the output of several object detectors applied at each viewpoint. We cast the problem as an energy minimization modeled with a Conditional Random Field (CRF). Instead of predicting the presence of an object at each image location independently, we simultaneously predict the labeling of the entire scene. Our CRF is able to take into account occlusions between objects and contextual constraints among them. We propose an effective iterative strategy that renders tractable the underlying optimization problem, and learn the parameters of the model with the max-margin paradigm. We evaluate the performance of our model on several challenging multi-camera pedestrian detection datasets namely PETS 2009 [5] and EPFL terrace sequence [9]. We also introduce a new dataset in which multiple classes of objects appear simultaneously in the scene. It is here where we show that our method effectively handles occlusions in the multi-class case.
Gemma Roig, Xavier Boix, Horesh Ben Shitrit, Pascal Fua
ICCV2
2010 Harmony potentials for joint classification and segmentation
abstract
Hierarchical conditional random fields have been successfully applied to object segmentation. One reason is their ability to incorporate contextual information at different scales. However, these models do not allow multiple labels to be assigned to a single node. At higher scales in the image, this yields an oversimplified model, since multiple classes can be reasonable expected to appear within one region. This simplified model especially limits the impact that observations at larger scales may have on the CRF model. Neglecting the information at larger scales is undesirable since class-label estimates based on these scales are more reliable than at smaller, noisier scales. To address this problem, we propose a new potential, called harmony potential, which can encode any possible combination of class labels. We propose an effective sampling strategy that renders tractable the underlying optimization problem. Results show that our approach obtains state-of-the-art results on two challenging datasets: Pascal VOC 2009 and MSRC-21.
Josep M. Gonfaus, Xavier Boix, Joost van de Weijer 0001, Andrew D. Bagdanov, Joan Serrat 0002, Jordi Gonzàlez 0001
CVPR2