VLDB 2026 Research / reviewers in the wild / expert
Nuno Vasconcelos
dblp:78/4806
· DBLP profile ↗
214ranked-venue papers
34as first author
48since 2021 · last 2025
0000-0002-9024-4302ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 177 · 18 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 143 · 27 first-author · 37 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EditAR: Unified Conditional Generation with Autoregressive ModelsabstractRecent progress in controllable image generation and editing is largely driven by diffusion-based methods. Although diffusion models perform exceptionally well in specific tasks with tailored designs, establishing a unified model is still challenging. In contrast, autoregressive models inherently feature a unified tokenized representation, which simplifies the creation of a single foundational model for various tasks. In this work, we propose EditAR, a single unified autoregressive framework for a variety of conditional image generation tasks, e.g., image editing, depth-to-image, edge-to-image, segmentation-to-image. The model takes both images and instructions as inputs, and predicts the edited images tokens in a vanilla next-token paradigm. To enhance the text-to-image alignment, we further propose to distill the knowledge from foundation models into the autoregressive modeling process. We evaluate its effectiveness across diverse tasks on established benchmarks, showing competitive performance to various state-of-the-art task-specific methods. Project page: https://jitengmu.github.io/EditAR/ Jiteng Mu, Nuno Vasconcelos, Xiaolong Wang 0004 |
CVPR | 2 |
| 2025 | Diffusion-based Data Augmentation for Object Counting ProblemsabstractCrowd counting, an important problem in computer vision, is commonly solved with deep learning approaches, such as convolutional networks and transformers. However, Deep networks often overfit when the available labeled crowd data is scarce. To overcome this, we have designed a pipeline that utilizes a diffusion model to generate extensive training data. We pioneer using diffusion models to generate images from high-density head location dot maps (a binary dot map that specifies the location of human heads) and are the first to use these diverse synthetic data to augment the crowd counting models. Our proposed smoothed density map input for ControlNet significantly improves ControlNet’s performance in generating crowds in the correct locations. Also, our proposed counting loss and guidance sampling for the diffusion model effectively minimize the discrepancies between the location dot map and the crowd images generated. Moreover, our versatile framework can be easily adapted to all kinds of counting problems. Extensive experiments demonstrate that our framework improves the counting performance on the ShanghaiTech, NWPU-Crowd, UCF-QNRF, and TRANCOS datasets. Yuelei Li, Jia Wan 0001, Nuno Vasconcelos |
ICASSP | 4 |
| 2025 | Guiding Diffusion Models With Adaptive Negative Sampling Without External Resources
Alakh Desai, Nuno Vasconcelos |
ICCV | 2 |
| 2025 | IntroStyle: Training-Free Introspective Style Attribution Using Diffusion FeaturesabstractText-to-image (T2I) models have recently gained widespread adoption. This has spurred concerns about safeguarding intellectual property rights and an increasing demand for mechanisms that prevent the generation of specific artistic styles. Existing methods for style extraction typically necessitate the collection of custom datasets and the training of specialized models. This, however, is resource-intensive, time-consuming, and often impractical for real-time applications. We present a novel, training-free framework to solve the style attribution problem, using the features produced by a diffusion model alone, without any external modules or retraining. This is denoted as Introspective Style attribution (IntroStyle) and is shown to have superior performance to state-of-the-art models for style attribution. We also introduce a synthetic dataset of Artistic Style Split (ArtSplit) to isolate artistic style and evaluate fine-grained style attribution performance. Our experimental results on WikiArt and DomainNet datasets show that \ours is robust to the dynamic nature of artistic styles, outperforming existing methods by a wide margin. Jiteng Mu, Nuno Vasconcelos |
ICCV | 3 |
| 2025 | HQGS: High-Quality Novel View Synthesis with Gaussian Splatting in Degraded Scenesabstract3D Gaussian Splatting (3DGS) has shown promising results for Novel View Synthesis. However, while it is quite effective when based on high-quality images, its performance declines as image quality degrades, due to lack of resolution, motion blur, noise, compression artifacts, or other factors common in real-world data collection. While some solutions have been proposed for specific types of degradation, general techniques are still missing. To address the problem, we propose a robust HQGS that significantly enhances the 3DGS under various degradation scenarios. We first analyze that 3DGS lacks sufficient attention in some detailed regions in low-quality scenes, leading to the absence of Gaussian primitives in those areas and resulting in loss of detail in the rendered images. To address this issue, we focus on leveraging edge structural information to provide additional guidance for 3DGS, enhancing its robustness. First, we introduce an edge-semantic fusion guidance module that combines rich texture information from high-frequency edge-aware maps with semantic information from images. The fused features serve as prior guidance to capture detailed distribution across different regions, bringing more attention to areas with detailed edge information and allowing for a higher concentration of Gaussian primitives to be assigned to such areas. Additionally, we present a structural cosine similarity loss to complement pixel-level constraints, further improving the quality of the rendered images. Extensive experiments demonstrate that our method offers better robustness and achieves the best results across various degraded scenes. Source code and trained models are publicly available at: \url{https://github.com/linxin0/HQGS}. Shi Luo, Xiaojun Shan, Chao Ren 0002, Lu Qi 0001, Ming-Hsuan Yang 0001, Nuno Vasconcelos |
ICLR | 8 |
| 2025 | Core Knowledge Deficits in Multi-Modal Language ModelsabstractWhile Multi-modal Large Language Models (MLLMs) demonstrate impressive abilities over high-level perception and reasoning, their robustness in the wild remains limited, often falling short on tasks that are intuitive and effortless for humans. We examine the hypothesis that these deficiencies stem from the absence of core knowledge—rudimentary cognitive abilities innate to humans from early childhood.
To explore the core knowledge representation in MLLMs, we introduce CoreCognition, a large-scale benchmark encompassing 12 core knowledge concepts grounded in developmental cognitive science. We evaluate 230 models with 11 different prompts, leading to a total of 2,530 data points for analysis. Our experiments uncover four key findings, collectively demonstrating core knowledge deficits in MLLMs: they consistently underperform and show reduced, or even absent, scalability on low-level abilities relative to high-level ones.
Finally, we propose Concept Hacking, a novel controlled evaluation method, that reveals MLLMs fail to progress toward genuine core knowledge understanding, but instead rely on shortcut learning as they scale. Project page at https://williamium3000.github.io/core-knowledge/. Yijiang Li, Qingying Gao, Tianwei Zhao, Bingyang Wang, Haiyun Lyu, Robert D. Hawkins, Nuno Vasconcelos, Tal Golan, Dezhi Luo, Hokin Deng |
ICML | 8 |
| 2025 | EgoPrivacy: What Your First-Person Camera Says About You?abstractWhile the rapid proliferation of wearable cameras has raised significant concerns about egocentric video privacy, prior work has largely overlooked the unique privacy threats posed to the camera wearer. This work investigates the core question: How much privacy information about the camera wearer can be inferred from their first-person view videos? We introduce EgoPrivacy, the first large-scale benchmark for the comprehensive evaluation of privacy risks in egocentric vision. EgoPrivacy covers three types of privacy (demographic, individual, and situational), defining seven tasks that aim to recover private information ranging from fine-grained (e.g., wearer's identity) to coarse-grained (e.g., age group). To further emphasize the privacy threats inherent to egocentric vision, we propose Retrieval-Augmented Attack, a novel attack strategy that leverages ego-to-exo retrieval from an external pool of exocentric videos to boost the effectiveness of demographic privacy attacks. An extensive comparison of the different attacks possible under all threat models is presented, showing that private information of the wearer is highly susceptible to leakage. For instance, our findings indicate that foundation models can effectively compromise wearer privacy even in zero-shot settings by recovering attributes such as identity, scene, gender, and race with 70–80% accuracy. Our code and data are available at https://github.com/williamium3000/ego-privacy. Yijiang Li, Genpei Zhang, Yi Li 0051, Xiaojun Shan, Dashan Gao 0001, Jiancheng Lyu, Ning Bi, Nuno Vasconcelos |
ICML | 10 |
| 2025 | Ego-VPA: Egocentric Video Understanding with Parameter-Efficient AdaptationabstractVideo understanding typically requires fine-tuning the large backbone when adapting to new domains. In this paper, we leverage the egocentric video foundation models (EgoVFMs) based on video-language pre-training and propose a parameter-efficient adaptation for egocentric video tasks, namely Ego-VPA. It employs a local sparse approximation for each video frame/text feature using the basis prompts, and the selected basis prompts are used to synthesize video/text prompts. Since the basis prompts are shared across frames and modalities, it models context fusion and cross-modal transfer in an efficient fashion. Experiments show that EgoVPA excels in lightweight adaptation (with only 0.84% learnable parameters), largely improving over baselines and reaching the performance of full fine-tuning. Tz-Ying Wu, Kyle Min 0001, Subarna Tripathi, Nuno Vasconcelos |
WACV | 4 |
| 2024 | Towards Calibrated Multi-Label Deep Neural NetworksabstractThe problem of calibrating deep neural networks (DNNs) for multi-label learning is considered. It is well-known that DNNs trained by cross-entropy for single-label, or one-hot, classification are poorly calibrated. Many calibration techniques have been proposed to address the problem. However, little attention has been paid to the calibration of multi-label DNNs. In this literature, the focus has been on improving labeling accuracy in the face of severe dataset unbalance. This is addressed by the introduction of asymmetric losses, which have became very popular. However, these losses do not induce well calibrated classifiers. In this work, we first provide a theoretical explanation for this poor calibration performance, by showing that these loses losses lack the strictly proper property, a necessary condition for accurate probability estimation. To overcome this problem, we propose a new Strictly Proper Asymmetric (SPA) loss. This is complemented by a Label Pair Regularizer (LPR) that increases the number of calibration constraints introduced per training example. The effectiveness of both contributions is validated by extensive experiments on various multi-label datasets. The resulting training method is shown to significantly decrease the calibration error while maintaining state-of-the-art accuracy. Nuno Vasconcelos |
CVPR | 2 |
| 2024 | Long-Tailed Anomaly Detection with Learnable Class NamesabstractAnomaly detection (AD) aims to identify defective images and localize their defects (if any). Ideally, AD models should be able to detect defects over many image classes; without relying on hard-coded class names that can be uninformative or inconsistent across datasets; learn without anomaly supervision; and be robust to the long-tailed distributions of real-world applications. To address these challenges, we formulate the problem of long-tailed AD by introducing several datasets with different levels of class imbalance and metrics for performance evaluation. We then propose a novel method, LTAD, to detect defects from multiple and long-tailed classes, without relying on dataset class names. LTAD combines AD by reconstruction and semantic AD modules. AD by reconstruction is implemented with a transformer-based reconstruction module. Semantic AD is implemented with a binary classifier, which relies on learned pseudo class names and a pretrained foundation model. These modules are learned over two phases. Phase 1 learns the pseudo-class names and a variational autoencoder (VAE) for feature synthesis that augments the training data to combat long-tails. Phase 2 then learns the parameters of the reconstruction and classification modules of LTAD. Extensive experiments using the proposed long-tailed datasets show that LTAD substantially outperforms the state-of the-art methods for most forms of dataset imbalance. The long-tailed dataset split is available at https://zenodo.org/records/10854201. Chih-Hui Ho, Kuan-Chuan Peng, Nuno Vasconcelos |
CVPR | 3 |
| 2024 | ProTeCt: Prompt Tuning for Taxonomic Open Set ClassificationabstractVisual-language foundation models, like CLIP, learn generalized representations that enable zero-shot open-set clas-sification. Few-shot adaptation methods, based on prompt tuning, have been shown to further improve performance on downstream datasets. However, these methods do not fare well in the taxonomic open set (TOS) setting, where the classifier is asked to make prediction from label set across different levels of semantic granularity. Frequently, they infer incorrect labels at coarser taxonomic class levels, even when the inference at the leaf level (original class labels) is correct. To address this problem, we propose a prompt tuning technique that calibrates the hierarchical consistency of model predictions. A set of metrics of hierarchical consistency, the Hierarchical Consistent Accuracy (HCA) and the Mean Treecut Accuracy (MTA), are first proposed to evaluate TOS model performance. A new Prompt Tuning for Hierarchical Consistency (ProTeCt) technique is then proposed to calibrate classification across label set granularities. Results show that ProTeCt can be combined with existing prompt tuning methods to significantly improve TOS classification without degrading the leaf level classification performance. The code is available at https://github.com/gina9726/ProTeCt. Tz-Ying Wu, Chih-Hui Ho, Nuno Vasconcelos |
CVPR | 3 |
| 2024 | Learning a Dynamic Privacy-Preserving Camera Robust to Inversion Attacks
Jia Wan 0001, Nicholas Antipa, Nuno Vasconcelos |
ECCV (70) | 5 |
| 2024 | Improving Image Synthesis with Diffusion-Negative Sampling
Alakh Desai, Nuno Vasconcelos |
ECCV (53) | 2 |
| 2024 | Editable Image Elements for Controllable Synthesis
Jiteng Mu, Michaël Gharbi, Richard Zhang 0001, Eli Shechtman, Nuno Vasconcelos, Xiaolong Wang 0004, Taesung Park |
ECCV (2) | 5 |
| 2024 | Adapting Diffusion Models for Improved Prompt Compliance and Controllable Image SynthesisabstractRecent advances in generative modeling with diffusion processes (DPs) enabled breakthroughs in image synthesis. Despite impressive image quality, these models have various prompt compliance problems, including low recall in generating multiple objects, difficulty in generating text in images, and meeting constraints like object locations and pose. For fine-grained editing and manipulation, they also require fine-grained semantic or instance maps that are tedious to produce manually. While prompt compliance can be enhanced by addition of loss functions at inference, this is time consuming and does not scale to complex scenes.
To overcome these limitations, this work introduces a new family of $\textit{Factor Graph Diffusion Models}$ (FG-DMs) that models the joint distribution of images and conditioning variables, such as semantic, sketch, depth or normal maps via a factor graph decomposition. This joint structure has several advantages, including support for efficient sampling based prompt compliance schemes, which produce images of high object recall, semi-automated fine-grained editing, explainability at intermediate levels, ability to produce labeled datasets for the training of downstream models such as segmentation or depth, training with missing data, and continual learning where new conditioning variables can be added with minimal or no modifications to the existing structure. We propose an implementation of FG-DMs by adapting a pre-trained Stable Diffusion (SD) model to implement all FG-DM factors, using only COCO dataset, and show that it is effective in generating images with 15\% higher recall than SD while retaining its generalization ability. We introduce an attention distillation loss that encourages consistency among the attention maps of all factors, improving the fidelity of the generated conditions and image. We also show that training FG-DMs from scratch on MM-CelebA-HQ, Cityscapes, ADE20K, and COCO produce images of high quality (FID) and diversity (LPIPS). Deepak Sridhar, Abhishek Peri, Rohith Rachala, Nuno Vasconcelos |
NeurIPS | 4 |
| 2023 | Dense Network Expansion for Class Incremental LearningabstractThe problem of class incremental learning (CIL) is considered. State-of-the-art approaches use a dynamic architecture based on network expansion (NE), in which a task expert is added per task. While effective from a computational standpoint, these methods lead to models that grow quickly with the number of tasks. A new NE method, dense network expansion (DNE), is proposed to achieve a better trade-off between accuracy and model complexity. This is accomplished by the introduction of dense connections between the intermediate layers of the task expert networks, that enable the transfer of knowledge from old to new tasks via feature sharing and reusing. This sharing is implemented with a cross-task attention mechanism, based on a new task attention block (TAB), that fuses information across tasks. Unlike traditional attention mechanisms, TAB operates at the level of the feature mixing and is decoupled with spatial attentions. This is shown more effective than a joint spatial-and-task attention for CIL. The proposed DNE approach can strictly maintain the feature space of old classes while growing the network and feature scale at a much slower rate than previous methods. In result, it outperforms the previous SOTA methods by a margin of 4% in terms of accuracy, with similar or even smaller model scale. Yunsheng Li, Jiancheng Lyu, Dashan Gao 0001, Nuno Vasconcelos |
CVPR | 5 |
| 2023 | SViTT: Temporal Learning of Sparse Video-Text TransformersabstractDo video-text transformers learn to model temporal relationships across frames? Despite their immense capacity and the abundance of multimodal training data, recent work has revealed the strong tendency of video-text models towards frame-based spatial representations, while temporal reasoning remains largely unsolved. In this work, we identify several key challenges in temporal learning of video-text transformers: the spatiotemporal trade-off from limited network size; the curse of dimensionality for multi-frame modeling; and the diminishing returns of semantic information by extending clip length. Guided by these findings, we propose SVi TT, a sparse video-text architecture that performs multi-frame reasoning with significantly lower cost than naïve transformers with dense attention. Analogous to graph-based networks, SViTT employs two forms of sparsity: edge sparsity that limits the query-key communications between tokens in self-attention, and node sparsity that discards uninformative visual tokens. Trained with a curriculum which increases model sparsity with the clip length, SVi TT outperforms dense transformer baselines on multiple video-text retrieval and question answering benchmarks, with a fraction of computational cost. Project page: http://svcl.ucsd.edu/projects/svitt. Yi Li 0051, Kyle Min 0001, Subarna Tripathi, Nuno Vasconcelos |
CVPR | 4 |
| 2023 | Towards Professional Level Crowd Annotation of Expert Domain DataabstractImage recognition on expert domains is usually fine-grained and requires expert labeling, which is costly. This limits dataset sizes and the accuracy of learning systems. To address this challenge, we consider annotating expert data with crowdsourcing. This is denoted as PrOfeSsional lEvel cRowd (POSER) annotation. A new approach, based on semi-supervised learning (SSL) and denoted as SSL with human filtering (SSL-HF) is proposed. It is a human-in-the-loop SSL method, where crowd-source workers act as filters of pseudo-labels, replacing the unreliable confidence thresholding used by state-of-the-art SSL methods. To enable annotation by non-experts, classes are specified implicitly, via positive and negative sets of examples and augmented with deliberative explanations, which highlight regions of class ambiguity. In this way, SSL-HF leverages the strong low-shot learning and confidence estimation ability of humans to create an intuitive but effective labeling experience. Experiments show that SSL-HF significantly outperforms various alternative approaches in several benchmarks. Nuno Vasconcelos |
CVPR | 2 |
| 2023 | Toward Unsupervised Realistic Visual Question AnsweringabstractThe problem of realistic VQA (RVQA), where a model has to reject unanswerable questions (UQs) and answer answerable ones (AQs), is studied. We first point out 2 drawbacks in current RVQA research, where (1) datasets contain too many unchallenging UQs and (2) a large number of annotated UQs are required for training. To resolve the first drawback, we propose a new testing dataset, RGQA, which combines AQs from an existing VQA dataset with around 29K human-annotated UQs. These UQs consist of both fine-grained and coarse-grained image-question pairs generated with 2 approaches: CLIP-based and Perturbation-based. To address the second drawback, we introduce an unsupervised training approach. This combines pseudo UQs obtained by randomly pairing images and questions, with an RoI Mixup procedure to generate more fine-grained pseudo UQs, and model ensembling to regularize model confidence. Experiments show that using pseudo UQs significantly outperforms RVQA baselines. RoI Mixup and model ensembling further increase the gain. Finally, human evaluation reveals a performance gap between humans and models, showing that more RVQA research is needed. Code and dataset is released on https://github.com/chihhuiho/RGQA. Yuwei Zhang 0001, Chih-Hui Ho, Nuno Vasconcelos |
ICCV | 3 |
| 2023 | ActorsNeRF: Animatable Few-shot Human Rendering with Generalizable NeRFsabstractWhile NeRF-based human representations have shown impressive novel view synthesis results, most methods still rely on a large number of images / views for training. In this work, we propose a novel animatable NeRF called ActorsNeRF. It is first pre-trained on diverse human subjects, and then adapted with few-shot monocular video frames for a new actor with unseen poses. Building on previous generalizable NeRFs with parameter sharing using a ConvNet encoder, ActorsNeRF further adopts two human priors to capture the large human appearance, shape, and pose variations. Specifically, in the encoded feature space, we will first align different human subjects in a category-level canonical space, and then align the same human from different frames in an instance-level canonical space for rendering. We quantitatively and qualitatively demonstrate that ActorsNeRF significantly outperforms the existing state-of-the-art on few-shot generalization to new people and poses on multiple datasets. Project page: https://jitengmu.github.io/ActorsNeRF/. Jiteng Mu, Shen Sang, Nuno Vasconcelos, Xiaolong Wang 0004 |
ICCV | 3 |
| 2023 | A Generalized Explanation Framework for Visualization of Deep Learning Model PredictionsabstractAttribution-based explanations are popular in computer vision but of limited use for fine-grained classification problems typical of expert domains, where classes differ by subtle details. In these domains, users also seek understanding of "why" a class was chosen and "why not" an alternative class. A new GenerAlized expLanatiOn fRamEwork (GALORE) is proposed to satisfy all these requirements, by unifying attributive explanations with explanations of two other types. The first is a new class of explanations, denoted deliberative, proposed to address the "why" question, by exposing the network insecurities about a prediction. The second is the class of counterfactual explanations, which have been shown to address the "why not" question but are now more efficiently computed. GALORE unifies these explanations by defining them as combinations of attribution maps with respect to various classifier predictions and a confidence score. An evaluation protocol that leverages object recognition (CUB200) and scene classification (ADE20 K) datasets combining part and attribute annotations is also proposed. Experiments show that confidence scores can improve explanation accuracy, deliberative explanations provide insight into the network deliberation process, the latter correlates with that performed by humans, and counterfactual explanations enhance the performance of human students in machine teaching experiments. Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Calibrating Deep Neural Networks by Pairwise ConstraintsabstractIt is well known that deep neural networks (DNNs) pro-duce poorly calibrated estimates of class-posterior prob-abilities. We hypothesize that this is due to the limited calibration supervision provided by the cross-entropy loss, which places all emphasis on the probability of the true class and mostly ignores the remaining. We consider how each example can supervise all classes and show that the calibration of a C-way classification problem is equivalent to the calibration of C(C - 1) /2 pairwise binary classifi-cation problems that can be derived from it. This suggests the hypothesis that DNN calibration can be improved by providing calibration supervision to all such binary prob-lems. An implementation of this calibration by pairwise constraints (CPC) is then proposed, based on two types of binary calibration constraints. This is finally shown to be implementable with a very minimal increase in the complex-ity of cross-entropy training. Empirical evaluations of the proposed CPC method across multiple datasets and DNN architectures demonstrate state-of-the-art calibration per-formance. Nuno Vasconcelos |
CVPR | 2 |
| 2022 | VALHALLA: Visual Hallucination for Machine TranslationabstractDesigning better machine translation systems by considering auxiliary inputs such as images has attracted much attention in recent years. While existing methods show promising performance over the conventional text-only translation systems, they typically require paired text and image as input during inference, which limits their applicability to real-world scenarios. In this paper, we introduce a visual hallucination framework, called VALHALLA, which requires only source sentences at inference time and instead uses hallucinated visual representations for multi-modal machine translation. In particular, given a source sentence an autoregressive hallucination transformer is used to predict a discrete visual representation from the input text, and the combined text and hallucinated representations are utilized to obtain the target translation. We train the hallucination transformer jointly with the translation transformer using standard backpropagation with crossentropy losses while being guided by an additional loss that encourages consistency between predictions using either groundtruth or hallucinated visual representations. Extensive experiments on three standard translation datasets with a diverse set of language pairs demonstrate the effectiveness of our approach over both text-only baselines and state-of-the-art methods. Project page: http://www.svcl.ucsd.jects/valhalla.edu/pro. Yi Li 0051, Rameswar Panda, Chun-Fu Chen 0001, Rogério Feris, David D. Cox, Nuno Vasconcelos |
CVPR | 7 |
| 2022 | Improving Video Model Transfer with Dynamic Representation LearningabstractTemporal modeling is an essential element in video understanding. While deep convolution-based architectures have been successful at solving large-scale video recognition datasets, recent work has pointed out that they are biased towards modeling short-range relations, often failing to capture long-term temporal structures in the videos, leading to poor transfer and generalization to new datasets. In this work, the problem of dynamic representation learning (DRL) is studied. We propose dynamic score, a measure of video dynamic modeling that describes the additional amount of information learned by a video network that cannot be captured by pure spatial student through knowledge distillation. DRL is then formulated as an adversarial learning problem between the video and spatial models, with the objective of maximizing the dynamic score of learned spatiotemporal classifier. The quality of learned video representations are evaluated on a diverse set of transfer learning problems concerning many-shot and few-shot action classification. Experimental results show that models learned with DRL outperform baselines in dynamic modeling, demonstrating higher transferability and generalization capacity to novel domains and tasks. Yi Li 0051, Nuno Vasconcelos |
CVPR | 2 |
| 2022 | CoordGAN: Self-Supervised Dense Correspondences Emerge from GANsabstractRecent advances show that Generative Adversarial Networks (GANs) can synthesize images with smooth variations along semantically meaningful latent directions, such as pose, expression, layout, etc. While this indicates that GANs implicitly learn pixel-level correspondences across images, few studies explored how to extract them explicitly. In this work, we introduce Coordinate GAN (CoordGAN), a structure-texture disentangled GAN that learns a dense correspondence map for each generated image. We represent the correspondence maps of different images as warped coordinate frames transformed from a canonical coordinate frame, i.e., the correspondence map, which describes the structure (e.g., the shape of a face), is controlled via a transformation. Hence, finding correspondences boils down to locating the same coordinate in different correspondence maps. In CoordGAN, we sample a transformation to represent the structure of a synthesized instance, while an independent texture branch is responsible for rendering appearance details orthogonal to the structure. Our approach can also extract dense correspondence maps for real images by adding an encoder on top of the generator. We quantitatively demonstrate the quality of the learned dense correspondences through segmentation mask transfer on multiple datasets. We also show that the proposed generator achieves better structure and texture disentanglement compared to existing approaches. Project page: https://jitengmu.github.io/CoordGAN/ Jiteng Mu, Shalini De Mello, Zhiding Yu, Nuno Vasconcelos, Xiaolong Wang 0004, Jan Kautz, Sifei Liu |
CVPR | 4 |
| 2022 | Omni-DETR: Omni-Supervised Object Detection with TransformersabstractWe consider the problem of omni-supervised object detection, which can use unlabeled, fully labeled and weakly labeled annotations, such as image tags, counts, points, etc., for object detection. This is enabled by a unified architecture, Omni-DETR, based on the recent progress on student-teacher framework and end-to-end transformer based object detection. Under this unified architecture, different types of weak labels can be leveraged to generate accurate pseudo labels, by a bipartite matching based filtering mechanism, for the model to learn. In the experiments, Omni-DETR has achieved state-of-the-art results on multiple datasets and settings. And we have found that weak annotations can help to improve detection performance and a mixture of them can achieve a better trade-off between annotation cost and accuracy than the standard complete annotation. These findings could encourage larger object detection datasets with mixture annotations. The code is available at https://github.com/amazon-research/omni-detr. Zhaowei Cai, Hao Yang 0043, Gurumurthy Swaminathan, Nuno Vasconcelos, Bernt Schiele, Stefano Soatto |
CVPR | 5 |
| 2022 | Class-Incremental Learning with Strong Pre-trained ModelsabstractClass-incremental learning (CIL) has been widely stud-ied under the setting of starting from a small number of classes (base classes). Instead, we explore an understud-ied real-world setting of CIL that starts with a strong model pre-trained on a large number of base classes. We hypoth-esize that a strong base model can provide a good repre-sentation for novel classes and incremental learning can be done with small adaptations. We propose a 2-stage training scheme, i) feature augmentation - cloning part of the backbone and fine-tuning it on the novel data, and ii) fusion - combining the base and novel classifiers into a unified classifier. Experiments show that the proposed method sig-nificantly outperforms state-of-the-art CIL methods on the large-scale ImageNet dataset (e.g. + 10% overall accuracy than the best). We also propose and analyze understudied practical CIL scenarios, such as base-novel overlap with distribution shift. Our proposed method is robust and gen-eralizes to all analyzed CIL settings. Tz-Ying Wu, Gurumurthy Swaminathan, Zhizhong Li 0001, Avinash Ravichandran, Nuno Vasconcelos, Rahul Bhotika, Stefano Soatto |
CVPR | 5 |
| 2022 | Should All Proposals Be Treated Equally in Object Detection?
Yunsheng Li, Yinpeng Chen, Xiyang Dai, Dongdong Chen 0001, Mengchen Liu, Pei Yu, Lu Yuan 0001, Zicheng Liu 0001, Nuno Vasconcelos |
ECCV (25) | 10 |
| 2022 | Breadcrumbs: Adversarial Class-Balanced Sampling for Long-Tailed Recognition
Bo Liu 0043, Hao Kang, Gang Hua 0001, Nuno Vasconcelos |
ECCV (24) | 5 |
| 2022 | Meta-learning over time for destination prediction tasksabstractA need to understand and predict vehicles' behavior underlies both public and private goals in the transportation domain, including urban planning and management, ride-sharing services, and intelligent transportation systems. Individuals' preferences and intended destinations vary throughout the day, week, and year: for example, bars are most popular in the evenings, and beaches are most popular in the summer. Despite this principle, we note that recent studies on a popular benchmark dataset from Porto, Portugal have found, at best, only marginal improvements in predictive performance from incorporating temporal information. We propose an approach based on hypernetworks, a variant of meta-learning ("learning to learn") in which a neural network learns to change its own weights in response to an input. In our case, the weights responsible for destination prediction vary with the metadata, in particular the time, of the input trajectory. The time-conditioned weights notably improve the model's error relative to ablation studies and comparable prior work, and we confirm our hypothesis that knowledge of time should improve prediction of a vehicle's intended destination. Mark Tenzer, Zeeshan Rasheed 0002, Khurram Shafique, Nuno Vasconcelos |
SIGSPATIAL/GIS | 4 |
| 2022 | Single-Stage Visual Relationship Learning using Conditional QueriesabstractResearch in scene graph generation (SGG) usually considers two-stage models, that is, detecting a set of entities, followed by combining them and labeling all possible relationships. While showing promising results, the pipeline structure induces large parameter and computation overhead, and typically hinders end-to-end optimizations. To address this, recent research attempts to train single-stage models that are more computationally efficient. With the advent of DETR, a set-based detection model, one-stage models attempt to predict a set of subject-predicate-object triplets directly in a single shot. However, SGG is inherently a multi-task learning problem that requires modeling entity and predicate distributions simultaneously. In this paper, we propose Transformers with conditional queries for SGG, namely, TraCQ with a new formulation for SGG that avoids the multi-task learning problem and the combinatorial entity pair distribution. We employ a DETR-based encoder-decoder design and leverage conditional queries to significantly reduce the entity label space as well, which leads to 20% fewer parameters compared to state-of-the-art one-stage models. Experimental results show that TraCQ not only outperforms existing single-stage scene graph generation methods, it also beats state-of-the-art two-stage methods on the Visual Genome dataset, yet is capable of end-to-end training and faster inference. Alakh Desai, Tz-Ying Wu, Subarna Tripathi, Nuno Vasconcelos |
NeurIPS | 4 |
| 2022 | DISCO: Adversarial Defense with Local Implicit FunctionsabstractThe problem of adversarial defenses for image classification, where the goal is to robustify a classifier against adversarial examples, is considered. Inspired by the hypothesis that these examples lie beyond the natural image manifold, a novel aDversarIal defenSe with local impliCit functiOns (DISCO) is proposed to remove adversarial perturbations by localized manifold projections. DISCO consumes an adversarial image and a query pixel location and outputs a clean RGB value at the location. It is implemented with an encoder and a local implicit module, where the former produces per-pixel deep features and the latter uses the features in the neighborhood of query pixel for predicting the clean RGB value. Extensive experiments demonstrate that both DISCO and its cascade version outperform prior defenses, regardless of whether the defense is known to the attacker. DISCO is also shown to be data and parameter efficient and to mount defenses that transfers across datasets, classifiers and attacks. Chih-Hui Ho, Nuno Vasconcelos |
NeurIPS | 2 |
| 2021 | Robust Audio-Visual Instance DiscriminationabstractWe present a self-supervised learning method to learn audio and video representations. Prior work uses the natural correspondence between audio and video to define a standard cross-modal instance discrimination task, where a model is trained to match representations from the two modalities. However, the standard approach introduces two sources of training noise. First, audio-visual correspondences often produce faulty positives since the audio and video signals can be uninformative of each other. To limit the detrimental impact of faulty positives, we optimize a weighted contrastive learning loss, which down-weighs their contribution to the overall loss. Second, since self-supervised contrastive learning relies on random sampling of negative instances, instances that are semantically similar to the base instance can be used as faulty negatives. To alleviate the impact of faulty negatives, we propose to optimize an instance discrimination loss with a soft target distribution that estimates relationships between instances. We validate our contributions through extensive experiments on action recognition tasks and show that they address the problems of audio-visual instance discrimination and improve transfer learning performance. Pedro Morgado 0001, Ishan Misra, Nuno Vasconcelos |
CVPR | 3 |
| 2021 | Audio-Visual Instance Discrimination with Cross-Modal AgreementabstractWe present a self-supervised learning approach to learn audio-visual representations from video and audio. Our method uses contrastive learning for cross-modal discrimination of video from audio and vice-versa. We show that optimizing for cross-modal discrimination, rather than within-modal discrimination, is important to learn good representations from video and audio. With this simple but powerful insight, our method achieves highly competitive performance when finetuned on action recognition tasks. Furthermore, while recent work in contrastive learning defines positive and negative samples as individual instances, we generalize this definition by exploring cross-modal agreement. We group together multiple instances as positives by measuring their similarity in both the video and audio feature spaces. Cross-modal agreement creates better positive and negative sets, which allows us to calibrate visual similarities by seeking within-modal discrimination of positive instances, and achieve significant gains on downstream tasks. Pedro Morgado 0001, Nuno Vasconcelos, Ishan Misra |
CVPR | 2 |
| 2021 | Learning Deep Classifiers Consistent With Fine-Grained Novelty DetectionabstractThe problem of novelty detection in fine-grained visual classification (FGVC) is considered. An integrated understanding of the probabilistic and distance-based approaches to novelty detection is developed within the frame-work of convolutional neural networks (CNNs). It is shown that softmax CNN classifiers are inconsistent with novelty detection, because their learned class-conditional distributions and associated distance metrics are unidentifiable. A new regularization constraint, the class-conditional Gaussianity loss, is then proposed to eliminate this unidentifiability, and enforce Gaussian class-conditional distributions. This enables training Novelty Detection Consistent Classifiers (NDCCs) that are jointly optimal for classification and novelty detection. Empirical evaluations show that NDCCs achieve significant improvements over the state-of-the-art on both small- and large-scale FGVC datasets. Nuno Vasconcelos |
CVPR | 2 |
| 2021 | Dynamic Transfer for Multi-Source Domain AdaptationabstractRecent works of multi-source domain adaptation focus on learning a domain-agnostic model, of which the parameters are static. However, such a static model is difficult to handle conflicts across multiple domains, and suffers from a performance degradation in both source domains and target domain. In this paper, we present dynamic transfer to address domain conflicts, where the model parameters are adapted to samples. The key insight is that adapting model across domains is achieved via adapting model across samples. Thus, it breaks down source domain barriers and turns multi-source domains into a single-source domain. This also simplifies the alignment between source and target domains, as it only requires the target domain to be aligned with any part of the union of source domains. Furthermore, we find dynamic transfer can be simply modeled by aggregating residual matrices and a static convolution matrix. Experimental results show that, without using domain labels, our dynamic transfer outperforms the state-of-the-art method by more than 3% on the large multi-source domain adaptation datasets – DomainNet. Source code is at https://github.com/liyunsheng13/DRT. Yunsheng Li, Lu Yuan 0001, Yinpeng Chen, Nuno Vasconcelos |
CVPR | 5 |
| 2021 | IMAGINE: Image Synthesis by Image-Guided Model InversionabstractWe introduce an inversion based method, denoted as IMAge-Guided model INvErsion (IMAGINE), to generate high-quality and diverse images from only a single training sample. We leverage the knowledge of image semantics from a pre-trained classifier to achieve plausible generations via matching multi-level feature representations in the classifier, associated with adversarial training with an external discriminator. IMAGINE enables the synthesis procedure to simultaneously 1) enforce semantic specificity constraints during the synthesis, 2) produce realistic images without generator training, and 3) give users intuitive control over the generation process. With extensive experimental results, we demonstrate qualitatively and quantitatively that IMAGINE performs favorably against state-of-the-art GAN-based and inversion-based methods, across three different image domains (i.e., objects, scenes, and textures). Yijun Li 0001, Krishna Kumar Singh, Jingwan Lu, Nuno Vasconcelos |
CVPR | 5 |
| 2021 | Rethinking and Improving the Robustness of Image Style TransferabstractExtensive research in neural style transfer methods has shown that the correlation between features extracted by a pre-trained VGG network has a remarkable ability to capture the visual style of an image. Surprisingly, however, this stylization quality is not robust and often degrades significantly when applied to features from more advanced and lightweight networks, such as those in the ResNet family. By performing extensive experiments with different network architectures, we find that residual connections, which represent the main architectural difference between VGG and ResNet, produce feature maps of small entropy, which are not suitable for style transfer. To improve the robustness of the ResNet architecture, we then propose a simple yet effective solution based on a softmax transformation of the feature activations that enhances their entropy. Experimental results demonstrate that this small magic can greatly improve the quality of stylization results, even for networks with random weights. This suggests that the architecture used for feature extraction is more important than the use of learned weights for the task of style transfer. Yijun Li 0001, Nuno Vasconcelos |
CVPR | 3 |
| 2021 | Gradient-Based Algorithms for Machine TeachingabstractThe problem of machine teaching is considered. A new formulation is proposed under the assumption of an optimal student, where optimality is defined in the usual machine learning sense of empirical risk minimization. This is a sensible assumption for machine learning students and for human students in crowdsourcing platforms, who tend to perform at least as well as machine learning systems. It is shown that, if allowed unbounded effort, the optimal student always learns the optimal predictor for a classification task. Hence, the role of the optimal teacher is to select the teaching set that minimizes student effort. This is formulated as a problem of functional optimization where, at each teaching iteration, the teacher seeks to align the steepest descent directions of the risk of (1) the teaching set and (2) entire example population. The optimal teacher, denoted MaxGrad, is then shown to maximize the gradient of the risk on the set of new examples selected per iteration. MaxGrad teaching algorithms are finally provided for both binary and multiclass tasks, and shown to have some similarities with boosting algorithms. Experimental evaluations demonstrate the effectiveness of MaxGrad, which outperforms previous algorithms on the classification task, for both machine learning and human students from MTurk, by a substantial margin. Kabir Nagrecha, Nuno Vasconcelos |
CVPR | 3 |
| 2021 | BEV-Net: Assessing Social Distancing Compliance by Joint People Localization and Geometric ReasoningabstractSocial distancing, an essential public health measure to limit the spread of contagious diseases, has gained significant attention since the outbreak of the COVID-19 pandemic. In this work, the problem of visual social distancing compliance assessment in busy public areas, with wide field-of-view cameras, is considered. A dataset of crowd scenes with people annotations under a bird’s eye view (BEV) and ground truth for metric distances is introduced, and several measures for the evaluation of social distance detection systems are proposed. A multi-branch network, BEV-Net, is proposed to localize individuals in world coordinates and identify high-risk regions where social distancing is violated. BEV-Net combines detection of head and feet locations, camera pose estimation, a differentiable homography module to map image into BEV coordinates, and geometric reasoning to produce a BEV map of the people locations in the scene. Experiments on complex crowded scenes demonstrate the power of the approach and show superior performance over baselines derived from methods in the literature. Applications of interest for public health decision makers are finally discussed. Datasets, code and pretrained models are publicly available at GitHub1. Zhirui Dai, Yuepeng Jiang, Yi Li 0051, Bo Liu 0043, Antoni B. Chan, Nuno Vasconcelos |
ICCV | 6 |
| 2021 | Learning of Visual Relations: The Devil is in the TailsabstractSignificant effort has been recently devoted to modeling visual relations. This has mostly addressed the design of architectures, typically by adding parameters and increasing model complexity. However, visual relation learning is a long-tailed problem, due to the combinatorial nature of joint reasoning about groups of objects. Increasing model complexity is, in general, illsuited for long-tailed problems due to their tendency to overfit. In this paper, we explore an alternative hypothesis, denoted the Devil is in the Tails. Under this hypothesis, better performance is achieved by keeping the model simple but improving its ability to cope with long-tailed distributions. To test this hypothesis, we devise a new approach for training visual relationships models, which is inspired by state-of-the-art long-tailed recognition literature. This is based on an iterative decoupled training scheme, denoted Decoupled Training for Devil in the Tails (DT2). DT2 employs a novel sampling approach, Alternating Class-Balanced Sampling (ACBS), to capture the interplay between the long-tailed entity and predicate distributions of visual relations. Results show that, with an extremely simple architecture, DT2-ACBS significantly out-performs much more complex state-of-the-art methods on scene graph generation tasks. This suggests that the development of sophisticated models must be considered in tandem with the long-tailed nature of the problem. Alakh Desai, Tz-Ying Wu, Subarna Tripathi, Nuno Vasconcelos |
ICCV | 4 |
| 2021 | MicroNet: Improving Image Recognition with Extremely Low FLOPsabstractThis paper aims at addressing the problem of substantial performance degradation at extremely low computational cost (e.g. 5M FLOPs on ImageNet classification). We found that two factors, sparse connectivity and dynamic activation function, are effective to improve the accuracy. The former avoids the significant reduction of network width, while the latter mitigates the detriment of reduction in network depth. Technically, we propose micro-factorized convolution, which factorizes a convolution matrix into low rank matrices, to integrate sparse connectivity into convolution. We also present a new dynamic activation function, named Dynamic Shift Max, to improve the non-linearity via maxing out multiple dynamic fusions between an input feature map and its circular channel shift. Building upon these two new operators, we arrive at a family of networks, named MicroNet, that achieves significant performance gains over the state of the art in the low FLOP regime. For instance, under the constraint of 12M FLOPs, MicroNet achieves 59.4% top-1 accuracy on ImageNet classification, outperforming MobileNetV3 by 9.6%. Source code is at https://github.com/liyunsheng13/micronet. Yunsheng Li, Yinpeng Chen, Xiyang Dai, Dongdong Chen 0001, Mengchen Liu, Lu Yuan 0001, Zicheng Liu 0001, Lei Zhang 0001, Nuno Vasconcelos |
ICCV | 9 |
| 2021 | GistNet: a Geometric Structure Transfer Network for Long-Tailed RecognitionabstractThe problem of long-tailed recognition, where the number of examples per class is highly unbalanced, is considered. It is hypothesized that the well known tendency of standard classifier training to overfit to popular classes can be exploited for effective transfer learning. Rather than eliminating this overfitting, e.g. by adopting popular class-balanced sampling methods, the learning algorithm should instead leverage this overfitting to transfer geometric information from popular to low-shot classes. A new classifier architecture, GistNet, is proposed to support this goal, using constellations of classifier parameters to encode the class geometry. A new learning algorithm is then proposed for GeometrIc Structure Transfer (GIST), with resort to a combination of loss functions that combine class-balanced and random sampling to guarantee that, while overfitting to the popular classes is restricted to geometric parameters, it is leveraged to transfer class geometry from popular to few-shot classes. This enables better generalization for few-shot classes without the need for the manual specification of class weights, or even the explicit grouping of classes into different types. Experiments on two popular long-tailed recognition datasets show that GistNet outperforms existing solutions to this problem. Bo Liu 0043, Hao Kang, Gang Hua 0001, Nuno Vasconcelos |
ICCV | 5 |
| 2021 | A-SDF: Learning Disentangled Signed Distance Functions for Articulated Shape RepresentationabstractRecent work has made significant progress on using implicit functions, as a continuous representation for 3D rigid object shape reconstruction. However, much less effort has been devoted to modeling general articulated objects. Compared to rigid objects, articulated objects have higher degrees of freedom, which makes it hard to generalize to unseen shapes. To deal with the large shape variance, we introduce Articulated Signed Distance Functions (A-SDF) to represent articulated shapes with a disentangled latent space, where we have separate codes for encoding shape and articulation. With this disentangled continuous representation, we demonstrate that we can control the articulation input and animate unseen instances with unseen joint angles. Furthermore, we propose a Test-Time Adaptation inference algorithm to adjust our model during inference. We demonstrate our model generalize well to out-of-distribution and unseen data, e.g., partial point clouds and real-world depth images. Project page with code: https://jitengmu.github.io/A-SDF/. Jiteng Mu, Weichao Qiu, Adam Kortylewski, Alan L. Yuille, Nuno Vasconcelos, Xiaolong Wang 0004 |
ICCV | 5 |
| 2021 | A Machine Teaching Framework for Scalable RecognitionabstractWe consider the scalable recognition problem in the fine-grained expert domain where large-scale data collection is easy whereas annotation is difficult. Existing solutions are typically based on semi-supervised or self-supervised learning. We propose an alternative new framework, MEMORABLE, based on machine teaching and online crowd-sourcing platforms. A small amount of data is first labeled by experts and then used to teach online annotators for the classes of interest, who finally label the entire dataset. Preliminary studies show that the accuracy of classifiers trained on the final dataset is a function of the accuracy of the student annotators. A new machine teaching algorithm, CMaxGrad, is then proposed to enhance this accuracy by introducing explanations in a state-of-the-art machine teaching algorithm. For this, CMaxGrad leverages counterfactual explanations, which take into account student predictions, thereby proving feedback that is student-specific, explicitly addresses the causes of student confusion, and adapts to the level of competence of the student. Experiments show that both MEMORABLE and CMaxGrad outperform existing solutions to their respective problems. Nuno Vasconcelos |
ICCV | 2 |
| 2021 | Revisiting Dynamic Convolution via Matrix Decomposition
Yunsheng Li, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen 0001, Lu Yuan 0001, Zicheng Liu 0001, Nuno Vasconcelos |
ICLR | 10 |
| 2021 | Deep Hashing with Hash-Consistent Large Margin Proxy Embeddings
Pedro Morgado 0001, Yunsheng Li, José Costa Pereira, Mohammad J. Saberian, Nuno Vasconcelos |
Int. J. Comput. Vis. | 5 |
| 2021 | Cascade R-CNN: High Quality Object Detection and Instance SegmentationabstractIn object detection, the intersection over union (IoU) threshold is frequently used to define positives/negatives. The threshold used to train a detector defines its quality. While the commonly used threshold of 0.5 leads to noisy (low-quality) detections, detection performance frequently degrades for larger thresholds. This paradox of high-quality detection has two causes: 1) overfitting, due to vanishing positive samples for large thresholds, and 2) inference-time quality mismatch between detector and test hypotheses. A multi-stage object detection architecture, the Cascade R-CNN, composed of a sequence of detectors trained with increasing IoU thresholds, is proposed to address these problems. The detectors are trained sequentially, using the output of a detector as training set for the next. This resampling progressively improves hypotheses quality, guaranteeing a positive training set of equivalent size for all detectors and minimizing overfitting. The same cascade is applied at inference, to eliminate quality mismatches between hypotheses and detectors. An implementation of the Cascade R-CNN without bells or whistles achieves state-of-the-art performance on the COCO dataset, and significantly improves high-quality detection on generic and specific object datasets, including VOC, KITTI, CityPerson, and WiderFace. Finally, the Cascade R-CNN is generalized to instance segmentation, with nontrivial improvements over the Mask R-CNN. Zhaowei Cai, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Rethinking Differentiable Search for Mixed-Precision Neural NetworksabstractLow-precision networks, with weights and activations quantized to low bit-width, are widely used to accelerate inference on edge devices. However, current solutions are uniform, using identical bit-width for all filters. This fails to account for the different sensitivities of different filters and is suboptimal. Mixed-precision networks address this problem, by tuning the bit-width to individual filter requirements. In this work, the problem of optimal mixed-precision network search (MPS) is considered. To circumvent its difficulties of discrete search space and combinatorial optimization, a new differentiable search architecture is proposed, with several novel contributions to advance the efficiency by leveraging the unique properties of the MPS problem. The resulting Efficient differentiable MIxed-Precision network Search (EdMIPS) method is effective at finding the optimal bit allocation for multiple popular networks, and can search a large model, e.g. Inception-V3, directly on ImageNet without proxy task in a reasonable amount of time. The learned mixed-precision networks significantly outperform their uniform counterparts. Zhaowei Cai, Nuno Vasconcelos |
CVPR | 2 |
| 2020 | Exploit Clues From Views: Self-Supervised and Regularized Learning for Multiview Object RecognitionabstractMultiview recognition has been well studied in the literature and achieves decent performance in object recognition and retrieval task. However, most previous works rely on supervised learning and some impractical underlying assumptions, such as the availability of all views in training and inference time. In this work, the problem of multiview self-supervised learning (MV-SSL) is investigated, where only image to object association is given. Given this setup, a novel surrogate task for self-supervised learning is proposed by pursuing "object invariant" representation. This is solved by randomly selecting an image feature of an object as object prototype, accompanied with multiview consistency regularization, which results in view invariant stochastic prototype embedding (VISPE). Experiments shows that the recognition and retrieval results using VISPE outperform that of other self-supervised learning methods on seen and unseen data. VISPE can also be applied to semi-supervised scenario and demonstrates robust performance with limited data available. Code is available at https://github.com/chihhuiho/VISPE. Chih-Hui Ho, Bo Liu 0043, Tz-Ying Wu, Nuno Vasconcelos |
CVPR | 4 |
| 2020 | Background Data Resampling for Outlier-Aware ClassificationabstractThe problem of learning an image classifier that allows detection of out-of-distribution (OOD) examples, with the help of auxiliary background datasets, is studied. While training with background has been shown to improve OOD detection performance, the optimal choice of such dataset remains an open question, and challenges of data imbalance and computational complexity make it a potentially inefficient or even impractical solution. Targeted at balancing between efficiency and detection quality, a dataset resampling approach is proposed for obtaining a compact yet representative set of background data points. The resampling algorithm takes inspiration from prior work on hard negative mining, performing an iterative adversarial weighting on the background examples and using the learned weights to obtain the subset of desired size. Experiments on different datasets, model architectures and training strategies validate the universal effectiveness and efficiency of adversarially resampled background data. Code is available at https://github.com/JerryYLi/bg-resample-ood. Yi Li 0051, Nuno Vasconcelos |
CVPR | 2 |
| 2020 | Few-Shot Open-Set Recognition Using Meta-LearningabstractThe problem of open-set recognition is considered. While previous approaches only consider this problem in the context of large-scale classifier training, we seek a unified solution for this and the low-shot classification setting. It is argued that the classic softmax classifier is a poor solution for open-set recognition, since it tends to overfit on the training classes. Randomization is then proposed as a solution to this problem. This suggests the use of meta-learning techniques, commonly used for few-shot classification, for the solution of open-set recognition. A new oPen sEt mEta LEaRning (PEELER) algorithm is then introduced. This combines the random selection of a set of novel classes per episode, a loss that maximizes the posterior entropy for examples of those classes, and a new metric learning formulation based on the Mahalanobis distance. Experimental results show that PEELER achieves state of the art open set recognition performance for both few-shot and large-scale recognition. On CIFAR and miniImageNet, it achieves substantial gains in seen/unseen class detection AUROC for a given seen-class classification accuracy. Bo Liu 0043, Hao Kang, Gang Hua 0001, Nuno Vasconcelos |
CVPR | 5 |
| 2020 | SCOUT: Self-Aware Discriminant Counterfactual ExplanationsabstractThe problem of counterfactual visual explanations is considered. A new family of discriminant explanations is introduced. These produce heatmaps that attribute high scores to image regions informative of a classifier prediction but not of a counter class. They connect attributive explanations, which are based on a single heat map, to counterfactual explanations, which account for both predicted class and counter class. The latter are shown to be computable by combination of two discriminant explanations, with reversed class pairs. It is argued that self-awareness, namely the ability to produce classification confidence scores, is important for the computation of discriminant explanations, which seek to identify regions where it is easy to discriminate between prediction and counter class. This suggests the computation of discriminant explanations by the combination of three attribution maps. The resulting counterfactual explanations are optimization free and thus much faster than previous methods. To address the difficulty of their evaluation, a proxy task and set of quantitative metrics are also proposed. Experiments under this protocol show that the proposed counterfactual explanations outperform the state of the art while achieving speeds much faster, for popular networks. In a human-learning machine teaching experiment, they are also shown to improve mean student accuracy from chance level to 95%. Nuno Vasconcelos |
CVPR | 2 |
| 2020 | Explainable Object-Induced Action Decision for Autonomous VehiclesabstractA new paradigm is proposed for autonomous driving. The new paradigm lies between the end-to-end and pipelined approaches, and is inspired by how humans solve the problem. While it relies on scene understanding, the latter only considers objects that could originate hazard. These are denoted as action inducing, since changes in their state should trigger vehicle actions. They also define a set of explanations for these actions, which should be produced jointly with the latter. An extension of the BDD100K dataset, annotated for a set of 4 actions and 21 explanations, is proposed. A new multi-task formulation of the problem, which optimizes the accuracy of both action commands and explanations, is then introduced. A CNN architecture is finally proposed to solve this problem, by combining reasoning about action inducing objects and global scene context. Experimental results show that the requirement of explanations improves the recognition of action-inducing objects, which in turn leads to better action predictions. Xiaoyin Yang, Lihang Gong, Hsuan-Chu Lin, Tz-Ying Wu, Yunsheng Li, Nuno Vasconcelos |
CVPR | 7 |
| 2020 | SPOT: Selective Point Cloud Voting for Better Proposal in Point Cloud Object Detection
Hongyuan Du, Linjun Li, Bo Liu 0043, Nuno Vasconcelos |
ECCV (11) | 4 |
| 2020 | Solving Long-Tailed Recognition with Deep Realistic Taxonomic Classifier
Tz-Ying Wu, Pedro Morgado 0001, Chih-Hui Ho, Nuno Vasconcelos |
ECCV (8) | 5 |
| 2020 | Contrastive Learning with Adversarial ExamplesabstractContrastive learning (CL) is a popular technique for self-supervised learning (SSL) of visual representations. It uses pairs of augmentations of unlabeled training examples to define a classification task for pretext learning of a deep embedding. Despite extensive works in augmentation procedures, prior works do not address the selection of challenging negative pairs, as images within a sampled batch are treated independently. This paper addresses the problem, by introducing a new family of adversarial examples for constrastive learning and using these examples to define a new adversarial training algorithm for SSL, denoted as CLAE. When compared to standard CL, the use of adversarial examples creates more challenging positive pairs and adversarial training produces harder negative pairs by accounting for all images in a batch during the optimization. CLAE is compatible with many CL methods in the literature. Experiments show that it improves the performance of several existing CL baselines on multiple datasets. Chih-Hui Ho, Nuno Vasconcelos |
NeurIPS | 2 |
| 2020 | Learning Representations from Audio-Visual Spatial AlignmentabstractWe introduce a novel self-supervised pretext task for learning representations from audio-visual content. Prior work on audio-visual representation learning leverages correspondences at the video level. Approaches based on audio-visual correspondence (AVC) predict whether audio and video clips originate from the same or different video instances. Audio-visual temporal synchronization (AVTS) further discriminates negative pairs originated from the same video instance but at different moments in time. While these approaches learn high-quality representations for downstream tasks such as action recognition, they completely disregard the spatial cues of audio and visual signals naturally occurring in the real world. To learn from these spatial cues, we tasked a network to perform contrastive audio-visual spatial alignment of 360\degree video and spatial audio. The ability to perform spatial alignment is enhanced by reasoning over the full spatial content of the 360\degree video using a transformer architecture to combine representations from multiple viewpoints. The advantages of the proposed pretext task are demonstrated on a variety of audio and visual downstream tasks, including audio-visual correspondence, spatial alignment, action recognition and video semantic segmentation. Dataset and code are available at https://github.com/pedro-morgado/AVSpatialAlignment. Pedro Morgado 0001, Yi Li 0051, Nuno Vasconcelos |
NeurIPS | 3 |
| 2020 | Learning Complexity-Aware Cascades for Pedestrian DetectionabstractThe problem of pedestrian detection is considered. The design of complexity-aware cascaded pedestrian detectors, combining features of very different complexities, is investigated. A new cascade design procedure is introduced, by formulating cascade learning as the Lagrangian optimization of a risk that accounts for both accuracy and complexity. A boosting algorithm, denoted as complexity aware cascade training (CompACT), is then derived to solve this optimization. CompACT cascades are shown to seek an optimal trade-off between accuracy and complexity by pushing features of higher complexity to the later cascade stages, where only a few difficult candidate patches remain to be classified. This enables the use of features of vastly different complexities in a single detector. In result, the feature pool can be expanded to features previously impractical for cascade design, such as the responses of a deep convolutional neural network (CNN). This is demonstrated through the design of pedestrian detectors with a pool of features whose complexities span orders of magnitude. The resulting cascade generalizes the combination of a CNN with an object proposal mechanism: rather than a pre-processing stage, CompACT cascades seamlessly integrate CNNs in their stages. This enables accurate detection at fairly fast speeds. Zhaowei Cai, Mohammad J. Saberian, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Semantic Fisher Scores for Task Transfer: Using Objects to Classify ScenesabstractThe transfer of a neural network (CNN) trained to recognize objects to the task of scene classification is considered. A Bag-of-Semantics (BoS) representation is first induced, by feeding scene image patches to the object CNN, and representing the scene image by the ensuing bag of posterior class probability vectors (semantic posteriors). The encoding of the BoS with a Fisher vector (FV) is then studied. A link is established between the FV of any probabilistic model and the Q-function of the expectation-maximization (EM) algorithm used to estimate its parameters by maximum likelihood. This enables 1) immediate derivation of FVs for any model for which an EM algorithm exists, and 2) leveraging efficient implementations from the EM literature for the computation of FVs. It is then shown that standard FVs, such as those derived from Gaussian or even Dirichlet mixtures, are unsuccessful for the transfer of semantic posteriors, due to the highly non-linear nature of the probability simplex. The analysis of these FVs shows that significant benefits can ensue by 1) designing FVs in the natural parameter space of the multinomial distribution, and 2) adopting sophisticated probabilistic models of semantic feature covariance. The combination of these two insights leads to the encoding of the BoS in the natural parameter space of the multinomial, using a vector of Fisher scores derived from a mixture of factor analyzers (MFA). A network implementation of the MFA Fisher Score (MFA-FS), denoted as the MFAFSNet, is finally proposed to enable end-to-end training. Experiments with various object CNNs and datasets show that the approach has state-of-the-art transfer performance. Somewhat surprisingly, the scene classification results are superior to those of a CNN explicitly trained for scene classification, using a large scene dataset (Places). This suggests that holistic analysis is insufficient for scene classification. The modeling of local object semantics appears to be at least equally important. The two approaches are also shown to be strongly complementary, leading to very large scene classification gains when combined, and outperforming all previous scene classification approaches by a sizable margin. Mandar Dixit, Yunsheng Li, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Super Diffusion for Salient Object DetectionabstractOne major branch of saliency object detection methods are diffusion-based which construct a graph model on a given image and diffuse seed saliency values to the whole graph by a diffusion matrix. While their performance is sensitive to specific feature spaces and scales used for the diffusion matrix definition, little work has been published to systematically promote the robustness and accuracy of salient object detection under the generic mechanism of diffusion. In this work, we firstly present a novel view of the working mechanism of the diffusion process based on mathematical analysis, which reveals that the diffusion process is actually computing the similarity of nodes with respect to the seeds based on diffusion maps. Following this analysis, we propose super diffusion, a novel inclusive learning-based framework for salient object detection, which makes the optimum and robust performance by integrating a large pool of feature spaces, scales and even features originally computed for non-diffusion-based salient object detection. A closed-form solution of the optimal parameters for the integration is determined through supervised learning. At the local level, we propose to promote each individual diffusion before the integration. Our mathematical analysis reveals the close relationship between saliency diffusion and spectral clustering. Based on this, we propose to re-synthesize each individual diffusion matrix from the most discriminative eigenvectors and the constant eigenvector (for saliency normalization). The proposed framework is implemented and experimented on prevalently used benchmark datasets, consistently leading to state-of-the-art performance. Peng Jiang 0002, Zhiyi Pan 0001, Changhe Tu, Nuno Vasconcelos, Baoquan Chen, Jingliang Peng |
IEEE Trans. Image Process. | 4 |
| 2019 | Towards Universal Object Detection by Domain AttentionabstractDespite increasing efforts on universal representations for visual recognition, few have addressed object detection. In this paper, we develop an effective and efficient universal object detection system that is capable of working on various image domains, from human faces and traffic signs to medical CT images. Unlike multi-domain models, this universal model does not require prior knowledge of the domain of interest. This is achieved by the introduction of a new family of adaptation layers, based on the principles of squeeze and excitation, and a new domain-attention mechanism. In the proposed universal detector, all parameters and computations are shared across domains, and a single network processes all domains all the time. Experiments, on a newly established universal object detection benchmark of 11 diverse datasets, show that the proposed detector outperforms a bank of individual detectors, a multi-domain detector, and a baseline universal detector, with a 1.3x parameter increase over a single-domain baseline detector. The code and benchmark are available at http://www.svcl.ucsd.edu/projects/universal-detection/. Xudong Wang 0007, Zhaowei Cai, Dashan Gao 0001, Nuno Vasconcelos |
CVPR | 4 |
| 2019 | Catastrophic Child's Play: Easy to Perform, Hard to Defend Adversarial AttacksabstractThe problem of adversarial CNN attacks is considered, with an emphasis on attacks that are trivial to perform but difficult to defend. A framework for the study of such attacks is proposed, using real world object manipulations. Unlike most works in the past, this framework supports the design of attacks based on both small and large image perturbations, implemented by camera shake and pose variation. A setup is proposed for the collection of such perturbations and determination of their perceptibility. It is argued that perceptibility depends on context, and a distinction is made between imperceptible and semantically imperceptible perturbations. While the former survives image comparisons, the latter are perceptible but have no impact on human object recognition. A procedure is proposed to determine the perceptibility of perturbations using Turk experiments, and a dataset of both perturbation classes which enables replicable studies of object manipulation attacks, is assembled. Experiments using defenses based on many datasets, CNN models, and algorithms from the literature elucidate the difficulty of defending these attacks -- in fact, none of the existing defenses is found effective against them. Better results are achieved with real world data augmentation, but even this is not foolproof. These results confirm the hypothesis that current CNNs are vulnerable to attacks implementable even by a child, and that such attacks may prove difficult to defend. Chih-Hui Ho, Brandon Leung, Erik Sandström, Yen Chang, Nuno Vasconcelos |
CVPR | 5 |
| 2019 | PIEs: Pose Invariant EmbeddingsabstractThe role of pose invariance in image recognition and retrieval is studied. A taxonomic classification of embeddings, according to their level of invariance, is introduced and used to clarify connections between existing embeddings, identify missing approaches, and propose invariant generalizations. This leads to a new family of pose invariant embeddings (PIEs), derived from existing approaches by a combination of two models, which follow from the interpretation of CNNs as estimators of class posterior probabilities: a view-to-object model and an object-to-class model. The new pose-invariant models are shown to have interesting properties, both theoretically and through experiments, where they outperform existing multiview approaches. Most notably, they achieve good performance for both 1) classification and retrieval, and 2) single and multiview inference. These are important properties for the design of real vision systems, where universal embeddings are preferable to task specific ones, and multiple images are usually not available at inference time. Finally, a new multiview dataset of real objects, imaged in the wild against complex backgrounds, is introduced. We believe that this is a much needed complement to the synthetic datasets in wide use and will contribute to the advancement of multiview recognition and retrieval. Chih-Hui Ho, Pedro Morgado 0001, Amir Persekian, Nuno Vasconcelos |
CVPR | 4 |
| 2019 | Efficient Multi-Domain Learning by Covariance NormalizationabstractThe problem of multi-domain learning of deep networks is considered. An adaptive layer is induced per target domain and a novel procedure, denoted covariance normalization (CovNorm), proposed to reduce its parameters. CovNorm is a data driven method of fairly simple implementation, requiring two principal component analyzes (PCA) and fine-tuning of a mini-adaptation layer. Nevertheless, it is shown, both theoretically and experimentally, to have several advantages over previous approaches, such as batch normalization or geometric matrix approximations. Furthermore, CovNorm can be deployed both when target datasets are available sequentially or simultaneously. Experiments show that, in both cases, it has performance comparable to a fully fine-tuned network, using as few as 0.13% of the corresponding parameters per target domain. Yunsheng Li, Nuno Vasconcelos |
CVPR | 2 |
| 2019 | REPAIR: Removing Representation Bias by Dataset ResamplingabstractModern machine learning datasets can have biases for certain representations that are leveraged by algorithms to achieve high performance without learning to solve the underlying task. This problem is referred to as “representation bias”. The question of how to reduce the representation biases of a dataset is investigated and a new dataset REPresentAtion bIas Removal (REPAIR) procedure is proposed. This formulates bias minimization as an optimization problem, seeking a weight distribution that penalizes examples easy for a classifier built on a given feature representation. Bias reduction is then equated to maximizing the ratio between the classification loss on the reweighted dataset and the uncertainty of the ground-truth class labels. This is a minimax problem that REPAIR solves by alternatingly updating classifier parameters and dataset resampling weights, using stochastic gradient descent. An experimental set-up is also introduced to measure the bias of any dataset for a given representation, and the impact of this bias on the performance of recognition models. Experiments with synthetic and action recognition data show that dataset REPAIR can significantly reduce representation bias, and lead to improved generalization of models trained on REPAIRed datasets. The tools used for characterizing representation bias, and the proposed dataset REPAIR algorithm, are available at https://github.com/JerryYLi/Dataset-REPAIR/. Yi Li 0051, Nuno Vasconcelos |
CVPR | 2 |
| 2019 | Bidirectional Learning for Domain Adaptation of Semantic SegmentationabstractDomain adaptation for semantic image segmentation is very necessary since manually labeling large datasets with pixel-level labels is expensive and time consuming. Existing domain adaptation techniques either work on limited datasets, or yield not so good performance compared with supervised learning. In this paper, we propose a novel bidirectional learning framework for domain adaptation of segmentation. Using the bidirectional learning, the image translation model and the segmentation adaptation model can be learned alternatively and promote to each other.Furthermore, we propose a self-supervised learning algorithm to learn a better segmentation adaptation model and in return improve the image translation model. Experiments show that our method superior to the state-of-the-art methods in domain adaptation of segmentation with a big margin. The source code is available at https://github.com/liyunsheng13/BDL. Yunsheng Li, Lu Yuan 0001, Nuno Vasconcelos |
CVPR | 3 |
| 2019 | NetTailor: Tuning the Architecture, Not Just the WeightsabstractReal-world applications of object recognition often require the solution of multiple tasks in a single platform. Under the standard paradigm of network fine-tuning, an entirely new CNN is learned per task, and the final network size is independent of task complexity. This is wasteful, since simple tasks require smaller networks than more complex tasks, and limits the number of tasks that can be solved simultaneously. To address these problems, we propose a transfer learning procedure, denoted NetTailor, in which layers of a pre-trained CNN are used as universal blocks that can be combined with small task-specific layers to generate new networks. Besides minimizing classification error, the new network is trained to mimic the internal activations of a strong unconstrained CNN, and minimize its complexity by the combination of 1) a soft-attention mechanism over blocks and 2) complexity regularization constraints. In this way, NetTailor can adapt the network architecture, not just its weights, to the target task. Experiments show that networks adapted to simple tasks, such as character or traffic sign recognition, become significantly smaller than those adapted to hard tasks, such as fine-grained recognition. More importantly, due to the modular nature of the procedure, this reduction in network complexity is achieved without compromise of either parameter sharing across tasks, or classification accuracy. Pedro Morgado 0001, Nuno Vasconcelos |
CVPR | 2 |
| 2019 | Volumetric Attention for 3D Medical Image Segmentation and Detection
Xudong Wang 0007, Shizhong Han, Yunqiang Chen, Dashan Gao 0001, Nuno Vasconcelos |
MICCAI (6) | 5 |
| 2019 | Deliberative Explanations: visualizing network insecuritiesabstractA new approach to explainable AI, denoted {\it deliberative explanations,\/} is proposed. Deliberative explanations are a visualization technique that aims to go beyond the simple visualization of the image regions (or, more generally, input variables) responsible for a network prediction. Instead, they aim to expose the deliberations carried by the network to arrive at that prediction, by uncovering the insecurities of the network about the latter. The explanation consists of a list of insecurities, each composed of 1) an image region (more generally, a set of input variables), and 2) an ambiguity formed by the pair of classes responsible for the network uncertainty about the region. Since insecurity detection requires quantifying the difficulty of network predictions, deliberative explanations combine ideas from the literatures on visual explanations and assessment of classification difficulty. More specifically, the proposed implementation combines attributions with respect to both class predictions and a difficulty score. An evaluation protocol that leverages object recognition (CUB200) and scene classification (ADE20K) datasets that combine part and attribute annotations is also introduced to evaluate the accuracy of deliberative explanations. Finally, an experimental evaluation shows that the most accurate explanations are achieved by combining non self-referential difficulty scores and second-order attributions. The resulting insecurities are shown to correlate with regions of attributes that are shared by different classes. Since these regions are also ambiguous for humans, deliberative explanations are intuitive, suggesting that the deliberative process of modern networks correlates with human reasoning. Nuno Vasconcelos |
NeurIPS | 2 |
| 2019 | Cost-sensitive support vector machines
Arya Iranmehr, Hamed Masnadi-Shirazi, Nuno Vasconcelos |
Neurocomputing | 3 |
| 2019 | Multiclass Boosting: Margins, Codewords, Losses, and AlgorithmsabstractThe problem of multiclass boosting is considered. A new formulation is presented, combining multi-dimensional predictors, multi-dimensional real-valued codewords, and proper multiclass margin loss functions. This leads to a number of contributions, such as maximum capacity codeword sets, a family of proper and margin enforcing losses, denoted as $\gamma-\phi$ losses, and two new multiclass boosting algorithms. These are descent procedures on the functional space spanned by a set of weak learners. The first, CD-MCBoost, is a coordinate descent procedure that updates one predictor component at a time. The second, GD-MCBoost, a gradient descent procedure that updates all components jointly. Both MCBoost algorithms are defined with respect to a $\gamma-\phi$ loss and can reduce to classical boosting procedures (such as AdaBoost and LogitBoost) for binary problems. Beyond the algorithms themselves, the proposed formulation enables a unified treatment of many previous multiclass boosting algorithms. This is used to show that the latter implement different combinations of optimization strategy, codewords, weak learners, and loss function, highlighting some of their deficiencies. It is shown that no previous method matches the support of MCBoost for real codewords of maximum capacity, a proper margin-enforcing loss function, and any family of multidimensional predictors and weak learners. Experimental results confirm the superiority of MCBoost, showing that the two proposed MCBoost algorithms outperform comparable prior methods on a number of datasets.\\ \\ \textbf{Keywords}: Boosting, Multiclass Boosting, Multiclass Classification, Margin Maximization, Loss Function. Mohammad J. Saberian, Nuno Vasconcelos |
J. Mach. Learn. Res. | 2 |
| 2018 | Cascade R-CNN: Delving Into High Quality Object DetectionabstractIn object detection, an intersection over union (IoU) threshold is required to define positives and negatives. An object detector, trained with low IoU threshold, e.g. 0.5, usually produces noisy detections. However, detection performance tends to degrade with increasing the IoU thresholds. Two main factors are responsible for this: 1) overfitting during training, due to exponentially vanishing positive samples, and 2) inference-time mismatch between the IoUs for which the detector is optimal and those of the input hypotheses. A multi-stage object detection architecture, the Cascade R-CNN, is proposed to address these problems. It consists of a sequence of detectors trained with increasing IoU thresholds, to be sequentially more selective against close false positives. The detectors are trained stage by stage, leveraging the observation that the output of a detector is a good distribution for training the next higher quality detector. The resampling of progressively improved hypotheses guarantees that all detectors have a positive set of examples of equivalent size, reducing the overfitting problem. The same cascade procedure is applied at inference, enabling a closer match between the hypotheses and the detector quality of each stage. A simple implementation of the Cascade R-CNN is shown to surpass all single-model object detectors on the challenging COCO dataset. Experiments also show that the Cascade R-CNN is widely applicable across detector architectures, achieving consistent gains independently of the baseline detector strength. The code is available at https://github.com/zhaoweicai/cascade-rcnn. Zhaowei Cai, Nuno Vasconcelos |
CVPR | 2 |
| 2018 | Feature Space Transfer for Data AugmentationabstractThe problem of data augmentation in feature space is considered. A new architecture, denoted the FeATure TransfEr Network (FATTEN), is proposed for the modeling of feature trajectories induced by variations of object pose. This architecture exploits a parametrization of the pose manifold in terms of pose and appearance. This leads to a deep encoder/decoder network architecture, where the encoder factors into an appearance and a pose predictor. Unlike previous attempts at trajectory transfer, FATTEN can be efficiently trained end-to-end, with no need to train separate feature transfer functions. This is realized by supplying the decoder with information about a target pose and the use of a multi-task loss that penalizes category- and pose-mismatches. In result, FATTEN discourages discontinuous or non-smooth trajectories that fail to capture the structure of the pose manifold, and generalizes well on object recognition tasks involving large pose variation. Experimental results on the artificial ModelNet database show that it can successfully learn to map source features to target features of a desired pose, while preserving class identity. Most notably, by using feature space transfer for data augmentation (w.r.t. pose and depth) on SUN-RGBD objects, we demonstrate considerable performance improvements on one/few-shot object recognition in a transfer learning setup, compared to current state-of-the-art methods. Bo Liu 0043, Xudong Wang 0007, Mandar Dixit, Roland Kwitt, Nuno Vasconcelos |
CVPR | 5 |
| 2018 | RESOUND: Towards Action Recognition Without Representation Bias
Yingwei Li 0001, Yi Li 0051, Nuno Vasconcelos |
ECCV (6) | 3 |
| 2018 | Towards Realistic Predictors
Nuno Vasconcelos |
ECCV (13) | 2 |
| 2018 | Self-Supervised Generation of Spatial Audio for 360° VideoabstractWe introduce an approach to convert mono audio recorded by a 360° video camera into spatial audio, a representation of the distribution of sound over the full viewing sphere. Spatial audio is an important component of immersive 360° video viewing, but spatial audio microphones are still rare in current 360° video production. Our system consists of end-to-end trainable neural networks that separate individual sound sources and localize them on the viewing sphere, conditioned on multi-modal analysis from the audio and 360° video frames. We introduce several datasets, including one filmed ourselves, and one collected in-the-wild from YouTube, consisting of 360° videos uploaded with spatial audio. During training, ground truth spatial audio serves as self-supervision and a mixed down mono track forms the input to our network. Using our approach we show that it is possible to infer the spatial localization of sounds based only on a synchronized 360° video and the mono audio track. Pedro Morgado 0001, Nuno Vasconcelos, Timothy R. Langlois, Oliver Wang |
NeurIPS | 2 |
| 2017 | Deep Learning with Low Precision by Half-Wave Gaussian QuantizationabstractThe problem of quantizing the activations of a deep neural network is considered. An examination of the popular binary quantization approach shows that this consists of approximating a classical non-linearity, the hyperbolic tangent, by two functions: a piecewise constant sign function, which is used in feedforward network computations, and a piecewise linear hard tanh function, used in the backpropagation step during network learning. The problem of approximating the widely used ReLU non-linearity is then considered. An half-wave Gaussian quantizer (HWGQ) is proposed for forward approximation and shown to have efficient implementation, by exploiting the statistics of of network activations and batch normalization operations. To overcome the problem of gradient mismatch, due to the use of different forward and backward approximations, several piece-wise backward approximators are then investigated. The implementation of the resulting quantized network, denoted as HWGQ-Net, is shown to achieve much closer performance to full precision networks, such as AlexNet, ResNet, GoogLeNet and VGG-Net, than previously available low-precision networks, with 1-bit binary weights and 2-bit quantized activations. Zhaowei Cai, Xiaodong He 0001, Jian Sun 0001, Nuno Vasconcelos |
CVPR | 4 |
| 2017 | AGA: Attribute-Guided AugmentationabstractWe consider the problem of data augmentation, i.e., generating artificial samples to extend a given corpus of training data. Specifically, we propose attributed-guided augmentation (AGA) which learns a mapping that allows to synthesize data such that an attribute of a synthesized sample is at a desired value or strength. This is particularly interesting in situations where little data with no attribute annotation is available for learning, but we have access to a large external corpus of heavily annotated samples. While prior works primarily augment in the space of images, we propose to perform augmentation in feature space instead. We implement our approach as a deep encoder-decoder architecture that learns the synthesis function in an end-to-end manner. We demonstrate the utility of our approach on the problems of (1) one-shot object recognition in a transfer-learning setting where we have no prior knowledge of the new classes, as well as (2) object-based one-shot scene recognition. As external data, we leverage 3D depth and pose information from the SUN RGB-D dataset. Our experiments show that attribute-guided augmentation of high-level CNN features considerably improves one-shot recognition performance on both problems. Mandar Dixit, Roland Kwitt, Marc Niethammer, Nuno Vasconcelos |
CVPR | 4 |
| 2017 | Semantically Consistent Regularization for Zero-Shot RecognitionabstractThe role of semantics in zero-shot learning is considered. The effectiveness of previous approaches is analyzed according to the form of supervision provided. While some learn semantics independently, others only supervise the semantic subspace explained by training classes. Thus, the former is able to constrain the whole space but lacks the ability to model semantic correlations. The latter addresses this issue but leaves part of the semantic space unsupervised. This complementarity is exploited in a new convolutional neural network (CNN) framework, which proposes the use of semantics as constraints for recognition. Although a CNN trained for classification has no transfer ability, this can be encouraged by learning an hidden semantic layer together with a semantic code for classification. Two forms of semantic constraints are then introduced. The first is a loss-based regularizer that introduces a generalization constraint on each semantic predictor. The second is a codeword regularizer that favors semantic-to-class mappings consistent with prior semantic knowledge while allowing these to be learned from data. Significant improvements over the state-of-the-art are achieved on several datasets. Pedro Morgado 0001, Nuno Vasconcelos |
CVPR | 2 |
| 2017 | Deep Scene Image Classification with the MFAFVNetabstractThe problem of transferring a deep convolutional network trained for object recognition to the task of scene image classification is considered. An embedded implementation of the recently proposed mixture of factor analyzers Fisher vector (MFA-FV) is proposed. This enables the design of a network architecture, the MFAFVNet, that can be trained in an end to end manner. The new architecture involves the design of a MFA-FV layer that implements a statistically correct version of the MFA-FV, through a combination of network computations and regularization. When compared to previous neural implementations of Fisher vectors, the MFAFVNet relies on a more powerful statistical model and a more accurate implementation. When compared to previous non-embedded models, the MFAFVNet relies on a state of the art model, which is now embedded into a CNN. This enables end to end training. Experiments show that the MFAFVNet has state of the art performance on scene classification. Yunsheng Li, Mandar Dixit, Nuno Vasconcelos |
ICCV | 3 |
| 2017 | Complex Activity Recognition Via Attribute Dynamics
Weixin Li 0002, Nuno Vasconcelos |
Int. J. Comput. Vis. | 2 |
| 2016 | Boosted Convolutional Neural Networks
Mohammad Moghimi, Serge J. Belongie, Mohammad J. Saberian, Jian Yang 0003, Nuno Vasconcelos, Li-Jia Li 0001 |
BMVC | 5 |
| 2016 | VLAD3: Encoding Dynamics of Deep Features for Action RecognitionabstractPrevious approaches to action recognition with deep features tend to process video frames only within a small temporal region, and do not model long-range dynamic information explicitly. However, such information is important for the accurate recognition of actions, especially for the discrimination of complex activities that share sub-actions, and when dealing with untrimmed videos. Here, we propose a representation, VLAD for Deep Dynamics (VLAD3), that accounts for different levels of video dynamics. It captures short-term dynamics with deep convolutional neural network features, relying on linear dynamic systems (LDS) to model medium-range dynamics. To account for long-range inhomogeneous dynamics, a VLAD descriptor is derived for the LDS and pooled over the whole video, to arrive at the final VLAD3representation. An extensive evaluation was performed on Olympic Sports, UCF101 and THUMOS15, where the use of the VLAD3representation leads to state-of-the-art results. Yingwei Li 0001, Weixin Li 0002, Vijay Mahadevan, Nuno Vasconcelos |
CVPR | 4 |
| 2016 | A Unified Multi-scale Deep Convolutional Neural Network for Fast Object Detection
Zhaowei Cai, Quanfu Fan, Rogério Feris, Nuno Vasconcelos |
ECCV (4) | 4 |
| 2016 | Semantic Clustering for Robust Fine-Grained Scene Recognition
Marian George, Mandar Dixit, Gábor Zogg, Nuno Vasconcelos |
ECCV (1) | 4 |
| 2016 | Peak-Piloted Deep Network for Facial Expression Recognition
Xiangyun Zhao, Xiaodan Liang, Luoqi Liu, Teng Li 0001, Yugang Han, Nuno Vasconcelos, Shuicheng Yan |
ECCV (2) | 6 |
| 2016 | Object based Scene Representations using Fisher Scores of Local Subspace ProjectionsabstractSeveral works have shown that deep CNN classifiers can be easily transferred across datasets, e.g. the transfer of a CNN trained to recognize objects on ImageNET to an object detector on Pascal VOC. Less clear, however, is the ability of CNNs to transfer knowledge across tasks. A common example of such transfer is the problem of scene classification that should leverage localized object detections to recognize holistic visual concepts. While this problem is currently addressed with Fisher vector representations, these are now shown ineffective for the high-dimensional and highly non-linear features extracted by modern CNNs. It is argued that this is mostly due to the reliance on a model, the Gaussian mixture of diagonal covariances, which has a very limited ability to capture the second order statistics of CNN features. This problem is addressed by the adoption of a better model, the mixture of factor analyzers (MFA), which approximates the non-linear data manifold by a collection of local subspaces. The Fisher score with respect to the MFA (MFA-FS) is derived and proposed as an image representation for holistic image classifiers. Extensive experiments show that the MFA-FS has state of the art performance for object-to-scene transfer and this transfer actually outperforms the training of a scene CNN from a large scene dataset. The two representations are also shown to be complementary, in the sense that their combination outperforms each of the representations by itself. When combined, they produce a state of the art scene classifier. Mandar Dixit, Nuno Vasconcelos |
NIPS | 2 |
| 2016 | Large Margin Discriminant Dimensionality Reduction in Prediction SpaceabstractIn this paper we establish a duality between boosting and SVM, and use this to derive a novel discriminant dimensionality reduction algorithm. In particular, using the multiclass formulation of boosting and SVM we note that both use a combination of mapping and linear classification to maximize the multiclass margin. In SVM this is implemented using a pre-defined mapping (induced by the kernel) and optimizing the linear classifiers. In boosting the linear classifiers are pre-defined and the mapping (predictor) is learned through combination of weak learners. We argue that the intermediate mapping, e.g. boosting predictor, is preserving the discriminant aspects of the data and by controlling the dimension of this mapping it is possible to achieve discriminant low dimensional representations for the data. We use the aforementioned duality and propose a new method, Large Margin Discriminant Dimensionality Reduction (LADDER) that jointly learns the mapping and the linear classifiers in an efficient manner. This leads to a data-driven mapping which can embed data into any number of dimensions. Experimental results show that this embedding can significantly improve performance on tasks such as hashing and image/scene classification. Mohammad J. Saberian, José Costa Pereira, Nuno Vasconcelos |
NIPS | 3 |
| 2016 | Person-following UAVsabstractWe consider the design of vision-based control algorithms for unmanned aerial vehicles (UAVs), so as to enable a UAV to autonomously follow a person. A new vision-based control architecture is proposed with the goals of 1) robustly following the user and 2) implementing following behaviors programmed by manipulation of visual patterns. This is achieved within a detection/tracking paradigm, where the target is a programmable badge worn by the user. This badge contains a visual pattern with two components. The first is fixed and used to locate the user. The second is variable and implements a code used to program the UAV behavior. A biologically inspired tracking/recognition architecture, combining bottom-up and top-down saliency mechanisms, a novel image similarity measure, and an affine validation procedure, is proposed to detect the badge in the scene. The badge location is used by a control algorithm to adjust the UAV flight parameters so as to maintain the user in the center of the field of view. The detected badge is further analyzed to extract the visual code that commands the UAV behavior This is used to control the height and distance of the UAV relative to the user. Francisca Vasconcelos, Nuno Vasconcelos |
WACV | 2 |
| 2016 | Parametric Regression on the GrassmannianabstractWe address the problem of fitting parametric curves on the Grassmann manifold for the purpose of intrinsic parametric regression. We start from the energy minimization formulation of linear least-squares in Euclidean space and generalize this concept to general nonflat Riemannian manifolds, following an optimal-control point of view. We then specialize this idea to the Grassmann manifold and demonstrate that it yields a simple, extensible and easy-to-implement solution to the parametric regression problem. In fact, it allows us to extend the basic geodesic model to (1) a "time-warped" variant and (2) cubic splines. We demonstrate the utility of the proposed solution on different vision problems, such as shape regression as a function of age, traffic-speed estimation and crowd-counting from surveillance video clips. Most notably, these problems can be conveniently solved within the same framework without any specifically-tailored steps along the processing pipeline. Yi Hong 0006, Roland Kwitt, Nikhil Singh 0002, Nuno Vasconcelos, Marc Niethammer |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | Guest Editorial Special Section on Visual Saliency Computing and LearningabstractVision and multimedia communities have long attempted to enable computers to understand image or video content in a manner analogous to humans. Humans’ comprehension to an image or a video clip often depends on the objects that draw their attention. As a result, one fundamental and open problem is to automatically infer the attention attracting or interesting areas in an image or a video sequence. Recently, a large number of researchers explore visual saliency models to address this problem. The study on visual saliency models is originally motivated by simulating humans’ bottom-up visual attention and it is mainly based on the biological evidence that humans’ visual attention is automatically attracted by highly salient features in the visual scene, which are discriminative with respect to the surrounding environment. Junwei Han 0001, Ling Shao 0001, Nuno Vasconcelos, Jungong Han, Dong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Scene classification with semantic Fisher vectorsabstractWith the help of a convolutional neural network (CNN) trained to recognize objects, a scene image is represented as a bag of semantics (BoS). This involves classifying image patches using the network and considering the class posterior probability vectors as locally extracted semantic descriptors. The image BoS is summarized using a Fisher vector (FV) embedding that exploits the properties of the space of these descriptors. The resulting representation is referred to as a semantic Fisher vector. Two implementations of a semantic FV are investigated. First involves modeling the BoS with a Dirichlet Mixture and computing the Fisher gradients for this model. Due to the difficulty of mixture modeling on a non-Euclidean probability simplex, this approach is shown to be unsuccessful. A second implementation is derived using the interpretation of semantic descriptors as parameters of a multinomial distribution. Like the parameters of any exponential family, these can be projected into their natural parameter space. For a CNN, this is shown equivalent to using inputs of its soft-max layer as patch descriptors. A semantic FV is then computed as a Gaussian Mixture FV in the space of these natural parameters. This representation is shown to outperform other alternatives such as FVs of features from the intermediate CNN layers or a classifier obtained by adapting (fine-tuning) the CNN. The proposed FV represents an embedding for object classification probabilities. As an image representation, therefore, it is complementary to the features obtained from a scene classification CNN. A combination of the two representations is shown to achieve state-of-the-art results on MIT Indoor scenes and SUN datasets. Mandar Dixit, Dashan Gao 0001, Nikhil Rasiwasia, Nuno Vasconcelos |
CVPR | 5 |
| 2015 | How many bits does it take for a stimulus to be salient?abstractVisual saliency has been shown to depend on the unpredictability of the visual stimulus given its surround. Various previous works have advocated the equivalence between stimulus saliency and uncompressibility. We propose a direct measure of this quantity, namely the number of bits required by an optimal video compressor to encode a given video patch, and show that features derived from this measure are highly predictive of eye fixations. To account for global saliency effects, these are embedded in a Markov random field model. The resulting saliency measure is shown to achieve state-of-the-art accuracy for the prediction of fixations, at a very low computational cost. Since most modern cameras incorporate video encoders, this paves the way for in-camera saliency estimation, which could be useful in a variety of computer vision applications. Seyed Hossein Khatoonabadi, Nuno Vasconcelos, Ivan V. Bajic, Yufeng Shan |
CVPR | 2 |
| 2015 | Multiple instance learning for soft bags via top instancesabstractA generalized formulation of the multiple instance learning problem is considered. Under this formulation, both positive and negative bags are soft, in the sense that negative bags can also contain positive instances. This reflects a problem setting commonly found in practical applications, where labeling noise appears on both positive and negative training samples. A novel bag-level representation is introduced, using instances that are most likely to be positive (denoted top instances), and its ability to separate soft bags, depending on their relative composition in terms of positive and negative instances, is studied. This study inspires a new large-margin algorithm for soft-bag classification, based on a latent support vector machine that efficiently explores the combinatorial space of bag compositions. Empirical evaluation on three datasets is shown to confirm the main findings of the theoretical analysis and the effectiveness of the proposed soft-bag classifier. Weixin Li 0002, Nuno Vasconcelos |
CVPR | 2 |
| 2015 | Learning Complexity-Aware Cascades for Deep Pedestrian DetectionabstractThe design of complexity-aware cascaded detectors, combining features of very different complexities, is considered. A new cascade design procedure is introduced, by formulating cascade learning as the Lagrangian optimization of a risk that accounts for both accuracy and complexity. A boosting algorithm, denoted as complexity aware cascade training (CompACT), is then derived to solve this optimization. CompACT cascades are shown to seek an optimal trade-off between accuracy and complexity by pushing features of higher complexity to the later cascade stages, where only a few difficult candidate patches remain to be classified. This enables the use of features of vastly different complexities in a single detector. In result, the feature pool can be expanded to features previously impractical for cascade design, such as the responses of a deep convolutional neural network (CNN). This is demonstrated through the design of a pedestrian detector with a pool of features whose complexities span orders of magnitude. The resulting cascade generalizes the combination of a CNN with an object proposal mechanism: rather than a pre-processing stage, CompACT cascades seamlessly integrate CNNs in their stages. This enables state of the art performance on the Caltech and KITTI datasets, at fairly fast speeds. Zhaowei Cai, Mohammad J. Saberian, Nuno Vasconcelos |
ICCV | 3 |
| 2015 | Generic Promotion of Diffusion-Based Salient Object DetectionabstractIn this work, we propose a generic scheme to promote any diffusion-based salient object detection algorithm by original ways to re-synthesize the diffusion matrix and construct the seed vector. We first make a novel analysis of the working mechanism of the diffusion matrix, which reveals the close relationship between saliency diffusion and spectral clustering. Following this analysis, we propose to re-synthesize the diffusion matrix from the most discriminative eigenvectors after adaptive re-weighting. Further, we propose to generate the seed vector based on the readily available diffusion maps, avoiding extra computation for color-based seed search. As a particular instance, we use inverse normalized Laplacian matrix as the original diffusion matrix and promote the corresponding salient object detection algorithm, which leads to superior performance as experimentally demonstrated. Peng Jiang 0002, Nuno Vasconcelos, Jingliang Peng |
ICCV | 2 |
| 2015 | Bayesian Model Adaptation for Crowd CountsabstractThe problem of transfer learning is considered in the domain of crowd counting. A solution based on Bayesian model adaptation of Gaussian processes is proposed. This is shown to produce intuitive model updates, which are tractable, and lead to an adapted model (predictive distribution) that accounts for all information in both training and adaptation data. The new adaptation procedure achieves significant gains over previous approaches, based on multi-task learning, while requiring much less computation to deploy. This makes it particularly suited for the problem of expanding the capacity of crowd counting camera networks. A large video dataset for the evaluation of adaptation approaches to crowd counting is also introduced. This contains a number of adaptation tasks, involving information transfer across video collected by 1) a single camera under different scene conditions (different times of the day) and 2) video collected from different cameras. Evaluation of the proposed model adaptation procedure in this dataset shows good performance in realistic operating conditions. Bo Liu 0043, Nuno Vasconcelos |
ICCV | 2 |
| 2015 | A view of margin losses as regularizers of probability estimates
Hamed Masnadi-Shirazi, Nuno Vasconcelos |
J. Mach. Learn. Res. | 2 |
| 2014 | Learning Optimal Seeds for Diffusion-Based Salient Object DetectionabstractIn diffusion-based saliency detection, an image is partitioned into superpixels and mapped to a graph, with superpixels as nodes and edge strengths proportional to superpixel similarity. Saliency information is then propagated over the graph using a diffusion process, whose equilibrium state yields the object saliency map. The optimal solution is the product of a propagation matrix and a saliency seed vector that contains a prior saliency assessment. This is obtained from either a bottom-up saliency detector or some heuristics. In this work, we propose a method to learn optimal seeds for object saliency. Two types of features are computed per superpixel: the bottom-up saliency of the superpixel region and a set of mid-level vision features informative of how likely the superpixel is to belong to an object. The combination of features that best discriminates between object and background saliency is then learned, using a large-margin formulation of the discriminant saliency principle. The propagation of the resulting saliency seeds, using a diffusion process, is finally shown to outperform the state of the art on a number of salient object detection datasets. Vijay Mahadevan, Nuno Vasconcelos |
CVPR | 3 |
| 2014 | Learning Receptive Fields for Pooling from Tensors of Feature ResponseabstractA new method for learning pooling receptive fields for recognition is presented. The method exploits the statistics of the 3D tensor of SIFT responses to an image. It is argued that the eigentensors of this tensor contain the information necessary for learning class-specific pooling recep- tive fields. It is shown that this information can be extracted by a simple PCA analysis of a specific tensor flattening. A novel algorithm is then proposed for fitting box-like receptive fields to the eigenimages extracted from a collection of images. The resulting receptive fields can be combined with any of the recently popular coding strategies for image classification. This combination is experimentally shown to improve classification accuracy for both vector quantization and Fisher vector (FV) encodings. It is then shown that the combination of the FV encoding with the proposed receptive fields has state-of-the-art performance for both object recognition and scene classification. Finally, when compared with previous attempts at learning receptive fields for pooling, the method is simpler and achieves better results. Nuno Vasconcelos |
CVPR | 2 |
| 2014 | Geodesic Regression on the Grassmannian
Yi Hong 0006, Roland Kwitt, Nikhil Singh 0002, Bradley C. Davis, Nuno Vasconcelos, Marc Niethammer |
ECCV (2) | 5 |
| 2014 | Guess-Averse Loss Functions For Cost-Sensitive Multiclass BoostingabstractCost-sensitive multiclass classification has recently acquired significance in several applications, through the introduction of multiclass datasets with well-defined misclassification costs. The design of classification algorithms for this setting is considered. It is argued that the unreliable performance of current algorithms is due to the inability of the underlying loss functions to enforce a certain fundamental underlying property. This property, denoted guess-aversion, is that the loss should encourage correct classifications over the arbitrary guessing that ensues when all classes are equally scored by the classifier. While guess-aversion holds trivially for binary classification, this is not true in the multiclass setting. A new family of cost-sensitive guess-averse loss functions is derived, and used to design new cost-sensitive multiclass boosting algorithms, denoted GEL- and GLL-MCBoost. Extensive experiments demonstrate (1) the general importance of guess-aversion and (2) that the GLL loss function outperforms other loss functions for multiclass boosting. Oscar Beijbom, Mohammad J. Saberian, David J. Kriegman, Nuno Vasconcelos |
ICML | 4 |
| 2014 | Multi-Resolution Cascades for Multiclass Object Detection
Mohammad J. Saberian, Nuno Vasconcelos |
NIPS | 2 |
| 2014 | Cross-modal domain adaptation for text-based regularization of image semantics in image retrieval systems
José Costa Pereira, Nuno Vasconcelos |
Comput. Vis. Image Underst. | 2 |
| 2014 | Boosting algorithms for detector cascade learning
Mohammad J. Saberian, Nuno Vasconcelos |
J. Mach. Learn. Res. | 2 |
| 2014 | Anomaly Detection and Localization in Crowded ScenesabstractThe detection and localization of anomalous behaviors in crowded scenes is considered, and a joint detector of temporal and spatial anomalies is proposed. The proposed detector is based on a video representation that accounts for both appearance and dynamics, using a set of mixture of dynamic textures models. These models are used to implement 1) a center-surround discriminant saliency detector that produces spatial saliency scores, and 2) a model of normal behavior that is learned from training data and produces temporal saliency scores. Spatial and temporal anomaly maps are then defined at multiple spatial scales, by considering the scores of these operators at progressively larger regions of support. The multiscale scores act as potentials of a conditional random field that guarantees global consistency of the anomaly judgments. A data set of densely crowded pedestrian walkways is introduced and used to evaluate the proposed anomaly detector. Experiments on this and other data sets show that the latter achieves state-of-the-art anomaly detection results. Weixin Li 0002, Vijay Mahadevan, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | On the Role of Correlation and Abstraction in Cross-Modal Multimedia RetrievalabstractThe problem of cross-modal retrieval from multimedia repositories is considered. This problem addresses the design of retrieval systems that support queries across content modalities, for example, using an image to search for texts. A mathematical formulation is proposed, equating the design of cross-modal retrieval systems to that of isomorphic feature spaces for different content modalities. Two hypotheses are then investigated regarding the fundamental attributes of these spaces. The first is that low-level cross-modal correlations should be accounted for. The second is that the space should enable semantic abstraction. Three new solutions to the cross-modal retrieval problem are then derived from these hypotheses: correlation matching (CM), an unsupervised method which models cross-modal correlations, semantic matching (SM), a supervised technique that relies on semantic representation, and semantic correlation matching (SCM), which combines both. An extensive evaluation of retrieval performance is conducted to test the validity of the hypotheses. All approaches are shown successful for text retrieval in response to image queries and vice versa. It is concluded that both hypotheses hold, in a complementary form, although evidence in favor of the abstraction hypothesis is stronger than that for correlation. José Costa Pereira, Emanuele Coviello, Gabriel Doyle, Nikhil Rasiwasia, Gert R. G. Lanckriet, Roger Levy, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2014 | Robust Deformable and Occluded Object Tracking With Dynamic GraphabstractWhile some efforts have been paid to handle deformation and occlusion in visual tracking, they are still great challenges. In this paper, a dynamic graph-based tracker (DGT) is proposed to address these two challenges in a unified framework. In the dynamic target graph, nodes are the target local parts encoding appearance information, and edges are the interactions between nodes encoding inner geometric structure information. This graph representation provides much more information for tracking in the presence of deformation and occlusion. The target tracking is then formulated as tracking this dynamic undirected graph, which is also a matching problem between the target graph and the candidate graph. The local parts within the candidate graph are separated from the background with Markov random field, and spectral clustering is used to solve the graph matching. The final target state is determined through a weighted voting procedure according to the reliability of part correspondence, and refined with recourse to a foreground/background segmentation. An effective online updating mechanism is proposed to update the model, allowing DGT to robustly adapt to variations of target structure. Experimental results show improved performance over several state-of-the-art trackers, in various challenging scenarios. Zhaowei Cai, Longyin Wen, Zhen Lei 0001, Nuno Vasconcelos, Stan Z. Li |
IEEE Trans. Image Process. | 4 |
| 2013 | Recognizing Activities via Bag of Words for Attribute DynamicsabstractIn this work, we propose a novel video representation for activity recognition that models video dynamics with attributes of activities. A video sequence is decomposed into short-term segments, which are characterized by the dynamics of their attributes. These segments are modeled by a dictionary of attribute dynamics templates, which are implemented by a recently introduced generative model, the binary dynamic system~(BDS). We propose methods for learning a dictionary of BDSs from a training corpus, and for quantizing attribute sequences extracted from videos into these BDS code words. This procedure produces a representation of the video as a histogram of BDS code words, which is denoted the bag-of-words for attribute dynamics (BoWAD). An extensive experimental evaluation reveals that this representation outperforms other state-of-the-art approaches in temporal structure modeling for complex activity recognition. Weixin Li 0002, Harpreet Sawhney, Nuno Vasconcelos |
CVPR | 4 |
| 2013 | Class-Specific Simplex-Latent Dirichlet Allocation for Image ClassificationabstractAn extension of the latent Dirichlet allocation (LDA), denoted class-specific-simplex LDA (css-LDA), is proposed for image classification. An analysis of the supervised LDA models currently used for this task shows that the impact of class information on the topics discovered by these models is very weak in general. This implies that the discovered topics are driven by general image regularities, rather than the semantic regularities of interest for classification. To address this, we introduce a model that induces supervision in topic discovery, while retaining the original flexibility of LDA to account for unanticipated structures of interest. The proposed css-LDA is an LDA model with class supervision at the level of image features. In css-LDA topics are discovered per class, i.e. a single set of topics shared across classes is replaced by multiple class-specific topic sets. This model can be used for generative classification using the Bayes decision rule or even extended to discriminative classification with support vector machines (SVMs). A css-LDA model can endow an image with a vector of class and topic specific count statistics that are similar to the Bag-of-words (BoW) histogram. SVM-based discriminants can be learned for classes in the space of these histograms. The effectiveness of css-LDA model in both generative and discriminative classification frameworks is demonstrated through an extensive experimental evaluation, involving multiple benchmark datasets, where it is shown to outperform all existing LDA based image classification approaches. Mandar Dixit, Nikhil Rasiwasia, Nuno Vasconcelos |
ICCV | 3 |
| 2013 | Dynamic Pooling for Complex Event RecognitionabstractThe problem of adaptively selecting pooling regions for the classification of complex video events is considered. Complex events are defined as events composed of several characteristic behaviors, whose temporal configuration can change from sequence to sequence. A dynamic pooling operator is defined so as to enable a unified solution to the problems of event specific video segmentation, temporal structure modeling, and event detection. Video is decomposed into segments, and the segments most informative for detecting a given event are identified, so as to dynamically determine the pooling operator most suited for each sequence. This dynamic pooling is implemented by treating the locations of characteristic segments as hidden information, which is inferred, on a sequence-by-sequence basis, via a large-margin classification rule with latent variables. Although the feasible set of segment selections is combinatorial, it is shown that a globally optimal solution to the inference problem can be obtained efficiently, through the solution of a series of linear programs. Besides the coarse-level location of segments, a finer model of video structure is implemented by jointly pooling features of segment-tuples. Experimental evaluation demonstrates that the resulting event detector has state-of-the-art performance on challenging video datasets. Weixin Li 0002, Ajay Divakaran, Nuno Vasconcelos |
ICCV | 4 |
| 2013 | Localizing target structures in ultrasound video - A phantom study
Roland Kwitt, Nuno Vasconcelos, Sharif Razzaque, Stephen R. Aylward |
Medical Image Anal. | 2 |
| 2013 | Biologically Inspired Object Tracking Using Center-Surround Saliency MechanismsabstractA biologically inspired discriminant object tracker is proposed. It is argued that discriminant tracking is a consequence of top-down tuning of the saliency mechanisms that guide the deployment of visual attention. The principle of discriminant saliency is then used to derive a tracker that implements a combination of center-surround saliency, a spatial spotlight of attention, and feature-based attention. In this framework, the tracking problem is formulated as one of continuous target-background classification, implemented in two stages. The first, or learning stage, combines a focus of attention (FoA) mechanism, and bottom-up saliency to identify a maximally discriminant set of features for target detection. The second, or detection stage, uses a feature-based attention mechanism and a target-tuned top-down discriminant saliency detector to detect the target. Overall, the tracker iterates between learning discriminant features from the target location in a video frame and detecting the location of the target in the next. The statistics of natural images are exploited to derive an implementation which is conceptually simple and computationally efficient. The saliency formulation is also shown to establish a unified framework for classifier design, target detection, automatic tracker initialization, and scale adaptation. Experimental results show that the proposed discriminant saliency tracker outperforms a number of state-of-the-art trackers in the literature. Vijay Mahadevan, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Latent Dirichlet Allocation Models for Image ClassificationabstractTwo new extensions of latent Dirichlet allocation (LDA), denoted topic-supervised LDA (ts-LDA) and class-specific-simplex LDA (css-LDA), are proposed for image classification. An analysis of the supervised LDA models currently used for this task shows that the impact of class information on the topics discovered by these models is very weak in general. This implies that the discovered topics are driven by general image regularities, rather than the semantic regularities of interest for classification. To address this, ts-LDA models are introduced which replace the automated topic discovery of LDA with specified topics, identical to the classes of interest for classification. While this results in improvements in classification accuracy over existing LDA models, it compromises the ability of LDA to discover unanticipated structure of interest. This limitation is addressed by the introduction of css-LDA, an LDA model with class supervision at the level of image features. In css-LDA topics are discovered per class, i.e., a single set of topics shared across classes is replaced by multiple class-specific topic sets. The css-LDA model is shown to combine the labeling strength of topic-supervision with the flexibility of topic-discovery. Its effectiveness is demonstrated through an extensive experimental evaluation, involving multiple benchmark datasets, where it is shown to outperform existing LDA-based image classification approaches. Nikhil Rasiwasia, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | On the regularization of image semantics by modal expansionabstractRecent research efforts in semantic representations and context modeling are based on the principle of task expansion: that vision problems such as object recognition, scene classification, or retrieval (RCR) cannot be solved in isolation. The extended principle of modality expansion (that RCR problems cannot be solved from visual information alone) is investigated in this work. A semantic image labeling system is augmented with text. Pairs of images and text are mapped to a semantic space, and the text features used to regularize their image counterparts. This is done with a new cross-modal regularizer, which learns the mapping of the image features that maximizes their average similarity to those derived from text. The proposed regularizer is class-sensitive, combining a set of class-specific denoising transformations and nearest neighbor interpolation of text-based class assignments. Regularization of a state-of-the-art approach to image retrieval is then shown to produce substantial gains in retrieval accuracy, outperforming recent image retrieval approaches. José Costa Pereira, Nuno Vasconcelos |
CVPR | 2 |
| 2012 | Boosting algorithms for simultaneous feature extraction and selectionabstractThe problem of simultaneous feature extraction and selection, for classifier design, is considered. A new framework is proposed, based on boosting algorithms that can either 1) select existing features or 2) assemble a combination of these features. This framework is simple and mathematically sound, derived from the statistical view of boosting and Taylor series approximations in functional space. Unlike classical boosting, which is limited to linear feature combinations, the new algorithms support more sophisticated combinations of weak learners, such as “sums of products” or “products of sums”. This is shown to enable the design of fairly complex predictor structures with few weak learners in a fully automated manner, leading to faster and more accurate classifiers, based on more informative features. Extensive experiments on synthetic data, UCI datasets, object detection and scene recognition show that these predictors consistently lead to more accurate classifiers than classical boosting algorithms. Mohammad J. Saberian, Nuno Vasconcelos |
CVPR | 2 |
| 2012 | Scene Recognition on the Semantic Manifold
Roland Kwitt, Nuno Vasconcelos, Nikhil Rasiwasia |
ECCV (4) | 2 |
| 2012 | Recognition in Ultrasound Videos: Where Am I?
Roland Kwitt, Nuno Vasconcelos, Sharif Razzaque, Stephen R. Aylward |
MICCAI (3) | 2 |
| 2012 | Recognizing Activities by Attribute DynamicsabstractIn this work, we consider the problem of modeling the dynamic structure of human activities in the attributes space. A video sequence is first represented in a semantic feature space, where each feature encodes the probability of occurrence of an activity attribute at a given time. A generative model, denoted the binary dynamic system (BDS), is proposed to learn both the distribution and dynamics of different activities in this space. The BDS is a non-linear dynamic system, which extends both the binary principal component analysis (PCA) and classical linear dynamic systems (LDS), by combining binary observation variables with a hidden Gauss-Markov state process. In this way, it integrates the representation power of semantic modeling with the ability of dynamic systems to capture the temporal structure of time-varying processes. An algorithm for learning BDS parameters, inspired by a popular LDS learning method from dynamic textures, is proposed. A similarity measure between BDSs, which generalizes the Binet-Cauchy kernel for LDS, is then introduced and used to design activity classifiers. The proposed method is shown to outperform similar classifiers derived from the kernel dynamic system (KDS) and state-of-the-art approaches for dynamics-based or attribute-based action recognition. Weixin Li 0002, Nuno Vasconcelos |
NIPS | 2 |
| 2012 | On the connections between saliency and trackingabstractA model connecting visual tracking and saliency has recently been proposed. This model is based on the saliency hypothesis for tracking which postulates that tracking is achieved by the top-down tuning, based on target features, of discriminant center-surround saliency mechanisms over time. In this work, we identify three main predictions that must hold if the hypothesis were true: 1) tracking reliability should be larger for salient than for non-salient targets, 2) tracking reliability should have a dependence on the defining variables of saliency, namely feature contrast and distractor heterogeneity, and must replicate the dependence of saliency on these variables, and 3) saliency and tracking can be implemented with common low level neural mechanisms. We confirm that the first two predictions hold by reporting results from a set of human behavior studies on the connection between saliency and tracking. We also show that the third prediction holds by constructing a common neurophysiologically plausible architecture that can computationally solve both saliency and tracking. This architecture is fully compliant with the standard physiological models of V1 and MT, and with what is known about attentional control in area LIP, while explaining the results of the human behavior experiments. Vijay Mahadevan, Nuno Vasconcelos |
NIPS | 2 |
| 2012 | Endoscopic image analysis in semantic space
Roland Kwitt, Nuno Vasconcelos, Nikhil Rasiwasia, Andreas Uhl, Bradley C. Davis, Michael Häfner, Friedrich Wrba |
Medical Image Anal. | 2 |
| 2012 | Holistic Context Models for Visual RecognitionabstractA novel framework to context modeling based on the probability of co-occurrence of objects and scenes is proposed. The modeling is quite simple, and builds upon the availability of robust appearance classifiers. Images are represented by their posterior probabilities with respect to a set of contextual models, built upon the bag-of-features image representation, through two layers of probabilistic modeling. The first layer represents the image in a semantic space, where each dimension encodes an appearance-based posterior probability with respect to a concept. Due to the inherent ambiguity of classifying image patches, this representation suffers from a certain amount of contextual noise. The second layer enables robust inference in the presence of this noise by modeling the distribution of each concept in the semantic space. A thorough and systematic experimental evaluation of the proposed context modeling is presented. It is shown that it captures the contextual “gist” of natural images. Scene classification experiments show that contextual classifiers outperform their appearance-based counterparts, irrespective of the precise choice and accuracy of the latter. The effectiveness of the proposed approach to context modeling is further demonstrated through a comparison to existing approaches on scene classification and image retrieval, on benchmark data sets. In all cases, the proposed approach achieves superior results. Nikhil Rasiwasia, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Learning Optimal Embedded CascadesabstractThe problem of automatic and optimal design of embedded object detector cascades is considered. Two main challenges are identified: optimization of the cascade configuration and optimization of individual cascade stages, so as to achieve the best tradeoff between classification accuracy and speed, under a detection rate constraint. Two novel boosting algorithms are proposed to address these problems. The first, RCBoost, formulates boosting as a constrained optimization problem which is solved with a barrier penalty method. The constraint is the target detection rate, which is met at all iterations of the boosting process. This enables the design of embedded cascades of known configuration without extensive cross validation or heuristics. The second, ECBoost, searches over cascade configurations to achieve the optimal tradeoff between classification risk and speed. The two algorithms are combined into an overall boosting procedure, RCECBoost, which optimizes both the cascade configuration and its stages under a detection rate constraint, in a fully automated manner. Extensive experiments in face, car, pedestrian, and panda detection show that the resulting detectors achieve an accuracy versus speed tradeoff superior to those of previous methods. Mohammad J. Saberian, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Counting People With Low-Level Features and Bayesian RegressionabstractAn approach to the problem of estimating the size of inhomogeneous crowds, which are composed of pedestrians that travel in different directions, without using explicit object segmentation or tracking is proposed. Instead, the crowd is segmented into components of homogeneous motion, using the mixture of dynamic-texture motion model. A set of holistic low-level features is extracted from each segmented region, and a function that maps features into estimates of the number of people per segment is learned with Bayesian regression. Two Bayesian regression models are examined. The first is a combination of Gaussian process regression with a compound kernel, which accounts for both the global and local trends of the count mapping but is limited by the real-valued outputs that do not match the discrete counts. We address this limitation with a second model, which is based on a Bayesian treatment of Poisson regression that introduces a prior distribution on the linear weights of the model. Since exact inference is analytically intractable, a closed-form approximation is derived that is computationally efficient and kernelizable, enabling the representation of nonlinear functions. An approximate marginal likelihood is also derived for kernel hyperparameter learning. The two regression-based crowd counting methods are evaluated on a large pedestrian data set, containing very distinct camera views, pedestrian traffic, and outliers, such as bikes or skateboarders. Experimental results show that regression-based counts are accurate regardless of the crowd size, outperforming the count estimates produced by state-of-the-art pedestrian detectors. Results on 2 h of video demonstrate the efficiency and robustness of the regression-based crowd size estimation over long periods of time. Antoni B. Chan, Nuno Vasconcelos |
IEEE Trans. Image Process. | 2 |
| 2011 | Adapted Gaussian models for image classificationabstractA general formulation of “Bayesian Adaptation” for generative and discriminative classification in the topic model framework is proposed. A generic topic-independent Gaussian mixture model, known as the background GMM, is learned using all available training data and adapted to the individual topics. In the generative framework, a Gaussian variant of the spatial pyramid model is used with a Bayes classifier. For the discriminative case, a novel predictive histogram representation for an image is presented. This builds upon the adapted topic model structure, using the individual class dictionaries and Bayesian weighting. The resulting histogram representation is evaluated for classification using a Support Vector Machine (SVM). A comparative evaluation of the proposed image models with the standard ones in the image classification literature is provided on three benchmark datasets. Mandar Dixit, Nikhil Rasiwasia, Nuno Vasconcelos |
CVPR | 3 |
| 2011 | TaylorBoost: First and second-order boosting algorithms with explicit margin controlabstractA new family of boosting algorithms, denoted Taylor-Boost, is proposed. It supports any combination of loss function and first or second order optimization, and includes classical algorithms such as AdaBoost, Gradient-Boost, or LogitBoost as special cases. Its restriction to the set of canonical losses makes it possible to have boosting algorithms with explicit margin control. A new large family of losses with this property, based on the set of cumulative distributions of zero mean random variables, is then proposed. A novel loss function in this family, the Laplace loss, is finally derived. The combination of this loss and second order TaylorBoost produces a boosting algorithm with explicit margin control. Mohammad J. Saberian, Hamed Masnadi-Shirazi, Nuno Vasconcelos |
CVPR | 3 |
| 2011 | Learning Pit Pattern Concepts for Gastroenterological Training
Roland Kwitt, Nikhil Rasiwasia, Nuno Vasconcelos, Andreas Uhl, Michael Häfner, Friedrich Wrba |
MICCAI (3) | 3 |
| 2011 | Maximum Covariance Unfolding : Manifold Learning for Bimodal DataabstractWe propose maximum covariance unfolding (MCU), a manifold learning algorithm for simultaneous dimensionality reduction of data from different input modalities. Given high dimensional inputs from two different but naturally aligned sources, MCU computes a common low dimensional embedding that maximizes the cross-modal (inter-source) correlations while preserving the local (intra-source) distances. In this paper, we explore two applications of MCU. First we use MCU to analyze EEG-fMRI data, where an important goal is to visualize the fMRI voxels that are most strongly correlated with changes in EEG traces. To perform this visualization, we augment MCU with an additional step for metric learning in the high dimensional voxel space. Second, we use MCU to perform cross-modal retrieval of matched image and text samples from Wikipedia. To manage large applications of MCU, we develop a fast implementation based on ideas from spectral graph theory. These ideas transform the original problem for MCU, one of semidefinite programming, into a simpler problem in semidefinite quadratic linear programming. Vijay Mahadevan, Chi Wah Wong, José Costa Pereira, Tom Liu, Nuno Vasconcelos, Lawrence K. Saul |
NIPS | 5 |
| 2011 | Multiclass Boosting: Theory and AlgorithmsabstractThe problem of multiclass boosting is considered. A new framework,based on multi-dimensional codewords and predictors is introduced. The optimal set of codewords is derived, and a margin enforcing loss proposed. The resulting risk is minimized by gradient descent on a multidimensional functional space. Two algorithms are proposed: 1) CD-MCBoost, based on coordinate descent, updates one predictor component at a time, 2) GD-MCBoost, based on gradient descent, updates all components jointly. The algorithms differ in the weak learners that they support but are both shown to be 1) Bayes consistent, 2) margin enforcing, and 3) convergent to the global minimum of the risk. They also reduce to AdaBoost when there are only two classes. Experiments show that both methods outperform previous multiclass boosting approaches on a number of datasets. Mohammad J. Saberian, Nuno Vasconcelos |
NIPS | 2 |
| 2011 | Generalized Stauffer-Grimson background subtraction for dynamic scenes
Antoni B. Chan, Vijay Mahadevan, Nuno Vasconcelos |
Mach. Vis. Appl. | 3 |
| 2011 | Cost-Sensitive BoostingabstractA novel framework is proposed for the design of cost-sensitive boosting algorithms. The framework is based on the identification of two necessary conditions for optimal cost-sensitive learning that 1) expected losses must be minimized by optimal cost-sensitive decision rules and 2) empirical loss minimization must emphasize the neighborhood of the target cost-sensitive boundary. It is shown that these conditions enable the derivation of cost-sensitive losses that can be minimized by gradient descent, in the functional space of convex combinations of weak learners, to produce novel boosting algorithms. The proposed framework is applied to the derivation of cost-sensitive extensions of AdaBoost, RealBoost, and LogitBoost. Experimental evidence, with a synthetic problem, standard data sets, and the computer vision problems of face and car detection, is presented in support of the cost-sensitive optimality of the new algorithms. Their performance is also compared to those of various previous cost-sensitive boosting proposals, as well as the popular combination of large-margin classifiers and probability calibration. Cost-sensitive boosting is shown to consistently outperform all other methods. Hamed Masnadi-Shirazi, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Anomaly detection in crowded scenesabstractA novel framework for anomaly detection in crowded scenes is presented. Three properties are identified as important for the design of a localized video representation suitable for anomaly detection in such scenes: (1) joint modeling of appearance and dynamics of the scene, and the abilities to detect (2) temporal, and (3) spatial abnormalities. The model for normal crowd behavior is based on mixtures of dynamic textures and outliers under this model are labeled as anomalies. Temporal anomalies are equated to events of low-probability, while spatial anomalies are handled using discriminant saliency. An experimental evaluation is conducted with a new dataset of crowded scenes, composed of 100 video sequences and five well defined abnormality categories. The proposed representation is shown to outperform various state of the art anomaly detection techniques. Vijay Mahadevan, Weixin Li 0002, Viral Bhalodia, Nuno Vasconcelos |
CVPR | 4 |
| 2010 | On the design of robust classifiers for computer visionabstractThe design of robust classifiers, which can contend with the noisy and outlier ridden datasets typical of computer vision, is studied. It is argued that such robustness requires loss functions that penalize both large positive and negative margins. The probability elicitation view of classifier design is adopted, and a set of necessary conditions for the design of such losses is identified. These conditions are used to derive a novel robust Bayes-consistent loss, denoted Tangent loss, and an associated boosting algorithm, denoted TangentBoost. Experiments with data from the computer vision problems of scene classification, object tracking, and multiple instance learning show that TangentBoost consistently outperforms previous boosting algorithms. Hamed Masnadi-Shirazi, Vijay Mahadevan, Nuno Vasconcelos |
CVPR | 3 |
| 2010 | Motion vector refinement for FRUC using saliency and segmentationabstractMotion-Compensated Frame Interpolation (MCFI) is a technique used extensively for increasing the temporal frequency of a video sequence. In order to obtain a high quality interpolation, the motion field between frames must be well-estimated. However, many current techniques for determining the motion are prone to errors in occlusion regions, as well as regions with repetitive structure. An algorithm is proposed for improving both the objective and subjective quality of MCFI by refining the motion vector field. A Discriminant Saliency classifier is employed to determine regions of the motion field which are most important to a human observer. These regions are refined using a multi-stage motion vector refinement which promotes candidates based on their likelihood given a local neighborhood. For regions which fall below the saliency threshold, frame segmentation is used to locate regions of homogeneous color and texture via Normalized Cuts. Motion vectors are promoted such that each homogeneous region has a consistent motion. Experimental results demonstrate an improvement over previous methods in both objective and subjective picture quality. Natan Jacobson, Yen-Lin Lee, Vijay Mahadevan, Nuno Vasconcelos, Truong Q. Nguyen |
ICME | 4 |
| 2010 | Risk minimization, probability elicitation, and cost-sensitive SVMs
Hamed Masnadi-Shirazi, Nuno Vasconcelos |
ICML | 2 |
| 2010 | A new approach to cross-modal multimedia retrievalabstractThe problem of joint modeling the text and image components of multimedia documents is studied. The text component is represented as a sample from a hidden topic model, learned with latent Dirichlet allocation, and images are represented as bags of visual (SIFT) features. Two hypotheses are investigated: that 1) there is a benefit to explicitly modeling correlations between the two components, and 2) this modeling is more effective in feature spaces with higher levels of abstraction. Correlations between the two components are learned with canonical correlation analysis. Abstraction is achieved by representing text and images at a more general, semantic level. The two hypotheses are studied in the context of the task of cross-modal document retrieval. This includes retrieving the text that most closely matches a query image, or retrieving the images that most closely match a query text. It is shown that accounting for cross-modal correlations and semantic abstraction both improve retrieval accuracy. The cross-modal model is also shown to outperform state-of-the-art image retrieval systems on a unimodal retrieval task. Nikhil Rasiwasia, José Costa Pereira, Emanuele Coviello, Gabriel Doyle, Gert R. G. Lanckriet, Roger Levy, Nuno Vasconcelos |
ACM Multimedia | 7 |
| 2010 | Variable margin losses for classifier designabstractThe problem of controlling the margin of a classifier is studied. A detailed analytical study is presented on how properties of the classification risk, such as its optimal link and minimum risk functions, are related to the shape of the loss, and its margin enforcing properties. It is shown that for a class of risks, denoted canonical risks, asymptotic Bayes consistency is compatible with simple analytical relationships between these functions. These enable a precise characterization of the loss for a popular class of link functions. It is shown that, when the risk is in canonical form and the link is inverse sigmoidal, the margin properties of the loss are determined by a single parameter. Novel families of Bayes consistent loss functions, of variable margin, are derived. These families are then used to design boosting style algorithms with explicit control of the classification margin. The new algorithms generalize well established approaches, such as LogitBoost. Experimental results show that the proposed variable margin losses outperform the fixed margin counterparts used by existing algorithms. Finally, it is shown that best performance can be achieved by cross-validating the margin parameter. Hamed Masnadi-Shirazi, Nuno Vasconcelos |
NIPS | 2 |
| 2010 | A biologically plausible network for the computation of orientation dominanceabstractThe determination of dominant orientation at a given image location is formulated as a decision-theoretic question. This leads to a novel measure for the dominance of a given orientation $\theta$, which is similar to that used by SIFT. It is then shown that the new measure can be computed with a network that implements the sequence of operations of the standard neurophysiological model of V1. The measure can thus be seen as a biologically plausible version of SIFT, and is denoted as bioSIFT. The network units are shown to exhibit trademark properties of V1 neurons, such as cross-orientation suppression, sparseness and independence. The connection between SIFT and biological vision provides a justification for the success of SIFT-like features and reinforces the importance of contrast normalization in computer vision. We illustrate this by replacing the Gabor units of an HMAX network with the new bioSIFT units. This is shown to lead to significant gains for classification tasks, leading to state-of-the-art performance among biologically inspired network models and performance competitive with the best non-biological object recognition systems. Kritika Muralidharan, Nuno Vasconcelos |
NIPS | 2 |
| 2010 | Boosting Classifier CascadesabstractThe problem of optimal and automatic design of a detector cascade is considered. A novel mathematical model is introduced for a cascaded detector. This model is analytically tractable, leads to recursive computation, and accounts for both classification and complexity. A boosting algorithm, FCBoost, is proposed for fully automated cascade design. It exploits the new cascade model, minimizes a Lagrangian cost that accounts for both classification risk and complexity. It searches the space of cascade configurations to automatically determine the optimal number of stages and their predictors, and is compatible with bootstrapping of negative examples and cost sensitive learning. Experiments show that the resulting cascades have state-of-the-art performance in various computer vision problems. Mohammad J. Saberian, Nuno Vasconcelos |
NIPS | 2 |
| 2010 | Spatiotemporal Saliency in Dynamic ScenesabstractA spatiotemporal saliency algorithm based on a center-surround framework is proposed. The algorithm is inspired by biological mechanisms of motion-based perceptual grouping and extends a discriminant formulation of center-surround saliency previously proposed for static imagery. Under this formulation, the saliency of a location is equated to the power of a predefined set of features to discriminate between the visual stimuli in a center and a surround window, centered at that location. The features are spatiotemporal video patches and are modeled as dynamic textures, to achieve a principled joint characterization of the spatial and temporal components of saliency. The combination of discriminant center-surround saliency with the modeling power of dynamic textures yields a robust, versatile, and fully unsupervised spatiotemporal saliency algorithm, applicable to scenes with highly dynamic backgrounds and moving cameras. The related problem of background subtraction is treated as the complement of saliency detection, by classifying nonsalient (with respect to appearance and motion dynamics) points in the visual field as background. The algorithm is tested for background subtraction on challenging sequences, and shown to substantially outperform various state-of-the-art techniques. Quantitatively, its average error rate is almost half that of the closest competitor. Vijay Mahadevan, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | A Novel Approach to FRUC Using Discriminant Saliency and Frame SegmentationabstractMotion-compensated frame interpolation (MCFI) is a technique used extensively for increasing the temporal frequency of a video sequence. In order to obtain a high quality interpolation, the motion field between frames must be well-estimated. However, many current techniques for determining the motion are prone to errors in occlusion regions, as well as regions with repetitive structure. We propose an algorithm for improving both the objective and subjective quality of MCFI by refining the motion vector field. We first utilize a discriminant saliency classifier to determine which regions of the motion field are most important to a human observer. These regions are refined using a multistage motion vector refinement (MVR), which promotes motion vector candidates based upon their likelihood given a local neighborhood. For regions which fall below the saliency-threshold, a frame segmentation is used to locate regions of homogeneous color and texture via normalized cuts. Motion vectors are promoted such that each homogeneous region has a consistent motion. Experimental results demonstrate an improvement over previous frame rate up-conversion (FRUC) methods in both objective and subjective picture quality. Natan Jacobson, Yen-Lin Lee, Vijay Mahadevan, Nuno Vasconcelos, Truong Q. Nguyen |
IEEE Trans. Image Process. | 4 |
| 2009 | Variational layered dynamic texturesabstractThe layered dynamic texture (LDT) is a generative model, which represents video as a collection of stochastic layers of different appearance and dynamics. Each layer is modeled as a temporal texture sampled from a different linear dynamical system, with regions of the video assigned to a layer using a Markov random field. Model parameters are learned from training video using the EM algorithm. However, exact inference for the E-step is intractable. In this paper, we propose a variational approximation for the LDT that enables efficient learning of the model. We also propose a temporally-switching LDT (TS-LDT), which allows the layer shape to change over time, along with the associated EM algorithm and variational approximation. The ability of the LDT to segment video into layers of coherent appearance and dynamics is also extensively evaluated, on both synthetic and natural video. These experiments show that the model possesses an ability to group regions of globally homogeneous, but locally heterogeneous, stochastic dynamics currently unparalleled in the literature. Antoni B. Chan, Nuno Vasconcelos |
CVPR | 2 |
| 2009 | Saliency-based discriminant trackingabstractWe propose a biologically inspired framework for visual tracking based on discriminant center surround saliency. At each frame, discrimination of the target from the background is posed as a binary classification problem. From a pool of feature descriptors for the target and background, a subset that is most informative for classification between the two is selected using the principle of maximum marginal diversity. Using these features, the location of the target in the next frame is identified using top-down saliency, completing one iteration of the tracking algorithm. We also show that a simple extension of the framework to include motion features in a bottom-up saliency mode can robustly identify salient moving objects and automatically initialize the tracker. The connections of the proposed method to existing works on discriminant tracking are discussed. Experimental results comparing the proposed method to the state of the art in tracking are presented, showing improved performance. Vijay Mahadevan, Nuno Vasconcelos |
CVPR | 2 |
| 2009 | Holistic context modeling using semantic co-occurrencesabstractWe present a simple framework to model contextual relationships between visual concepts. The new framework combines ideas from previous object-centric methods (which model contextual relationships between objects in an image, such as their co-occurrence patterns) and scene-centric methods (which learn a holistic context model from the entire image, known as its “gist”). This is accomplished without demarcating individual concepts or regions in the image. First, using the output of a generic appearance based concept detection system, a semantic space is formulated, where each axis represents a semantic feature. Next, context models are learned for each of the concepts in the semantic space, using mixtures of Dirichlet distributions. Finally, an image is represented as a vector of posterior concept probabilities under these contextual concept models. It is shown that these posterior probabilities are remarkably noise-free, and an effective model of the contextual relationships between semantic concepts in natural images. This is further demonstrated through an experimental evaluation with respect to two vision tasks, viz. scene classification and image annotation, on benchmark datasets. The results show that, besides quite simple to compute, the proposed context models attain superior performance than state of the art systems in both tasks. Nikhil Rasiwasia, Nuno Vasconcelos |
CVPR | 2 |
| 2009 | Bayesian Poisson regression for crowd countingabstractPoisson regression models the noisy output of a counting function as a Poisson random variable, with a log-mean parameter that is a linear function of the input vector. In this work, we analyze Poisson regression in a Bayesian setting, by introducing a prior distribution on the weights of the linear function. Since exact inference is analytically unobtainable, we derive a closed-form approximation to the predictive distribution of the model. We show that the predictive distribution can be kernelized, enabling the representation of non-linear log-mean functions. We also derive an approximate marginal likelihood that can be optimized to learn the hyperparameters of the kernel. We then relate the proposed approximate Bayesian Poisson regression to Gaussian processes. Finally, we present experimental results using Bayesian Poisson regression for crowd counting from low-level features. Antoni B. Chan, Nuno Vasconcelos |
ICCV | 2 |
| 2009 | Minimum Bayes error features for visual recognition
Gustavo Carneiro 0001, Nuno Vasconcelos |
Image Vis. Comput. | 2 |
| 2009 | Decision-Theoretic Saliency: Computational Principles, Biological Plausibility, and Implications for Neurophysiology and PsychophysicsabstractA decision-theoretic formulation of visual saliency, first proposed for top-down processing (object recognition) (Gao & Vasconcelos, 2005a), is extended to the problem of bottom-up saliency. Under this formulation, optimality is defined in the minimum probability of error sense, under a constraint of computational parsimony. The saliency of the visual features at a given location of the visual field is defined as the power of those features to discriminate between the stimulus at the location and a null hypothesis. For bottom-up saliency, this is the set of visual features that surround the location under consideration. Discrimination is defined in an information-theoretic sense and the optimal saliency detector derived for a class of stimuli that complies with known statistical properties of natural images. It is shown that under the assumption that saliency is driven by linear filtering, the optimal detector consists of what is usually referred to as the standard architecture of V1: a cascade of linear filtering, divisive normalization, rectification, and spatial pooling. The optimal detector is also shown to replicate the fundamental properties of the psychophysics of saliency: stimulus pop-out, saliency asymmetries for stimulus presence versus absence, disregard of feature conjunctions, and Weber's law. Finally, it is shown that the optimal saliency architecture can be applied to the solution of generic inference problems. In particular, for the class of stimuli studied, it performs the three fundamental operations of statistical inference: assessment of probabilities, implementation of Bayes decision rule, and feature selection. Dashan Gao 0001, Nuno Vasconcelos |
Neural Comput. | 2 |
| 2009 | Layered Dynamic TexturesabstractA novel video representation, the layered dynamic texture (LDT), is proposed. The LDT is a generative model, which represents a video as a collection of stochastic layers of different appearance and dynamics. Each layer is modeled as a temporal texture sampled from a different linear dynamical system. The LDT model includes these systems, a collection of hidden layer assignment variables (which control the assignment of pixels to layers), and a Markov random field prior on these variables (which encourages smooth segmentations). An EM algorithm is derived for maximum-likelihood estimation of the model parameters from a training video. It is shown that exact inference is intractable, a problem which is addressed by the introduction of two approximate inference procedures: a Gibbs sampler and a computationally efficient variational approximation. The trade-off between the quality of the two approximations and their complexity is studied experimentally. The ability of the LDT to segment videos into layers of coherent appearance and dynamics is also evaluated, on both synthetic and natural videos. These experiments show that the model possesses an ability to group regions of globally homogeneous, but locally heterogeneous, stochastic dynamics currently unparalleled in the literature. Antoni B. Chan, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Discriminant Saliency, the Detection of Suspicious Coincidences, and Applications to Visual RecognitionabstractA discriminant formulation of top-down visual saliency, intrinsically connected to the recognition problem, is proposed. The new formulation is shown to be closely related to a number of classical principles for the organization of perceptual systems, including infomax, inference by detection of suspicious coincidences, classification with minimal uncertainty, and classification with minimum probability of error. The implementation of these principles with computational parsimony, by exploitation of the statistics of natural images, is investigated. It is shown that Barlow's principle of inference by the detection of suspicious coincidences enables computationally efficient saliency measures which are nearly optimal for classification. This principle is adopted for the solution of the two fundamental problems in discriminant saliency, feature selection and saliency detection. The resulting saliency detector is shown to have a number of interesting properties, and act effectively as a focus of attention mechanism for the selection of interest points according to their relevance for visual recognition. Experimental evidence shows that the selected points have good performance with respect to 1) the ability to localize objects embedded in significant amounts of clutter, 2) the ability to capture information relevant for image classification, and 3) the richness of the set of visual attributes that can be considered salient. Dashan Gao 0001, Sunhyoung Han, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | Natural Image Statistics and Low-Complexity Feature SelectionabstractLow-complexity feature selection is analyzed in the context of visual recognition. It is hypothesized that high-order dependences of bandpass features contain little information for discrimination of natural images. This hypothesis is characterized formally by the introduction of the concepts of conjunctive interference and decomposability order of a feature set. Necessary and sufficient conditions for the feasibility of low-complexity feature selection are then derived in terms of these concepts. It is shown that the intrinsic complexity of feature selection is determined by the decomposability order of the feature set and not its dimension. Feature selection algorithms are then derived for all levels of complexity and are shown to be approximated by existing information-theoretic methods, which they consistently outperform. The new algorithms are also used to objectively test the hypothesis of low decomposability order through comparison of classification performance. It is shown that, for image classification, the gain of modeling feature dependencies has strongly diminishing returns: best results are obtained under the assumption of decomposability order 1. This suggests a generic law for bandpass features extracted from natural images: that the effect, on the dependence of any two features, of observing any other feature is constant across image classes. Manuela Vasconcelos, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Privacy preserving crowd monitoring: Counting people without people models or trackingabstractWe present a privacy-preserving system for estimating the size of inhomogeneous crowds, composed of pedestrians that travel in different directions, without using explicit object segmentation or tracking. First, the crowd is segmented into components of homogeneous motion, using the mixture of dynamic textures motion model. Second, a set of simple holistic features is extracted from each segmented region, and the correspondence between features and the number of people per segment is learned with Gaussian Process regression. We validate both the crowd segmentation algorithm, and the crowd counting system, on a large pedestrian dataset (2000 frames of video, containing 49,885 total pedestrian instances). Finally, we present results of the system running on a full hour of video. Antoni B. Chan, Zhang-Sheng John Liang, Nuno Vasconcelos |
CVPR | 3 |
| 2008 | Background subtraction in highly dynamic scenesabstractA new algorithm is proposed for background subtraction in highly dynamic scenes. Background subtraction is equated to the dual problem of saliency detection: background points are those considered not salient by suitable comparison of object and background appearance and dynamics. Drawing inspiration from biological vision, saliency is defined locally, using center-surround computations that measure local feature contrast. A discriminant formulation is adopted, where the saliency of a location is the discriminant power of a set of features with respect to the binary classification problem which opposes center to surround. To account for both motion and appearance, and achieve robustness to highly dynamic backgrounds, these features are spatiotemporal patches, which are modeled as dynamic textures. The resulting background subtraction algorithm is fully unsupervised, requires no training stage to learn background parameters, and depends only on the relative disparity of motion between the center and surround regions. This makes it insensitive to camera motion. The algorithm is tested on challenging video sequences, and shown to outperform various state-of-the-art techniques for background subtraction. Vijay Mahadevan, Nuno Vasconcelos |
CVPR | 2 |
| 2008 | Scene classification with low-dimensional semantic spaces and weak supervisionabstractA novel approach to scene categorization is proposed. Similar to previous works of [11, 15, 3, 12], we introduce an intermediate space, based on a low dimensional semantic ldquothemerdquo image representation. However, instead of learning the themes in an unsupervised manner, they are learned with weak supervision, from casual image annotations. Each theme induces a probability density on the space of low-level features, and images are represented as vectors of posterior theme probabilities. This enables an image to be associated with multiple themes, even when there are no multiple associations in the training labels. An implementation is presented and compared to various existing algorithms, on benchmark datasets. It is shown that the proposed low dimensional representation correlates well with human scene understanding, and is able to learn theme co-occurrences without explicit training. It is also shown to outperform unsupervised latent-space methods, with much smaller training complexity, and to achieve performance close to the state of the art methods, which rely on much higher-dimensional image representations. Finally a study of the effect of dimensionality on the classification performance is presented, indicating that the dimensionality of theme space grows sub-linearly with the number of scene categories. Nikhil Rasiwasia, Nuno Vasconcelos |
CVPR | 2 |
| 2008 | Object-Based Regions of Interest for Image CompressionabstractA fully automated architecture for object-based region of interest (ROI) detection is proposed. ROI's are defined as regions containing user defined objects of interest, and an efficient algorithm is developed for the detection of such regions. The algorithm is based on the principle of discriminant saliency, which defines as salient the image regions of strongest response to a set of features that optimally discriminate the object class of interest from all the others. It consists of two stages, saliency detection and saliency validation. The first detects salient points, the second verifies the consistency of their geometric configuration with that of training examples. Both the saliency detector and the configuration model can be learned from cluttered images downloaded from the web. Learning and ROI detection are optimal in the minimum probability of error (MPE) sense, and computationally efficient. This enables interactive user training of ROI-based image coders, with minimal amounts of manual supervision. Experimental results are presented for images of complex scenes, containing both objects and background clutter, and demonstrate good object-based ROI image compression performance. Sunhyoung Han, Nuno Vasconcelos |
DCC | 2 |
| 2008 | Complex discriminant features for object classificationabstractA new algorithm for the design of complex features, to be used in the discriminant saliency approach to object classification, is presented. The algorithm consists of sequential rotations of an initial basis of simple features, so as to maximize the discriminant power of the feature set for image classification. Discrimination is measured in an information theoretic sense. The proposed algorithm has lower complexity than popular techniques for learning parts, and is evaluated on classification tasks from the PASCAL challenge. It is shown that complex features consistently outperform simple features. Sunhyoung Han, Nuno Vasconcelos |
ICIP | 2 |
| 2008 | A systematic study of the role of context on image classificationabstractWe present the results of a systematic study of the contextual gain hypothesis for image classification. This hypothesis relates the traditional strategy of direct visual classification (DVC), and an alternative strategy based on indirect contextual classification (ICC). DVC is composed of classifiers that operate directly on pixel or feature based image representations. ICC relies on DVC to label images with respect to a pre-defined set of contextual semantic features. Image classification is then performed by a classifier that operates on the semantic space of these classifier outputs. The contextual gain hypothesis states that, in this semantic space, it is possible to design classifiers with better accuracy than those achievable with DVC. A framework for the systematic comparison of the DVC and ICC strategies is introduced, and an extensive comparison of the performance of the two strategies is carried out. Its results strongly suggest that the contextual gain hypothesis holds. Nikhil Rasiwasia, Nuno Vasconcelos |
ICIP | 2 |
| 2008 | Tumor Targeting for Lung Cancer Radiotherapy Using Machine Learning TechniquesabstractAccurate lung tumor targeting in real time plays a fundamental role in image-guide radiotherapy of lung cancers. Precise tumor targeting is required for both respiratory gating and tracking. Gating is considered as the current state of the art for precise lung cancer radiotherapy, which irradiates the tumor when it moves into a predefined gating window. Tracking seems to be a next-generation technique, and it operates in a more aggressive fashion by following the tumor position with radiation beam in real time. Existing methods for gating and tracking often rely on observed motion patterns of external surrogates or implanted fiducial markers. However, external surrogates suffer from certain degrees of inaccuracy, and implanted fiducial markers are in limited uses due to the risk of pneumothorax. Therefore, direct tumor targeting techniques without implanting fiducial markers are desired. Previous studies in fluoroscopic markerless targeting are mainly based on template matching methods, which may fail when tumor boundary is unclear in fluoroscopic images. In this paper, we propose a novel framework of markerless gating and tracking based on machine learning algorithms. Specifically, gating is treated as a two-class classification problem, which is solved by principal component analysis (PCA) and artificial neural network (ANN). Further, we formulate the tracking problem as a regression task, which employs the correlation between the tumor position and nearby surrogate anatomic features in the image. Four regression methods were tested in this study: 1-degree and 2-degree linear regression, artificial neural network (ANN), and support vector machine (SVM). Finally, we demonstrate the superb performance of the proposed markerless gating and tracking algorithms on 10 fluoroscopic image sequences of 9 patients. For gating, the target coverage (the precision) ranges from 90% to 99%, with mean of 96.5%. For tracking, the mean localization error is about 2.1 pixels and the maximum error at 95% confidence level is about 4.6 pixels (pixel size is about 0.5 mm). Tong Lin 0002, Laura I. Cervino, Nuno Vasconcelos, Steve B. Jiang |
ICMLA | 4 |
| 2008 | On the Design of Loss Functions for Classification: theory, robustness to outliers, and SavageBoostabstractThe machine learning problem of classifier design is studied from the perspective of probability elicitation, in statistics. This shows that the standard approach of proceeding from the specification of a loss, to the minimization of conditional risk is overly restrictive. It is shown that a better alternative is to start from the specification of a functional form for the minimum conditional risk, and derive the loss function. This has various consequences of practical interest, such as showing that 1) the widely adopted practice of relying on convex loss functions is unnecessary, and 2) many new losses can be derived for classification problems. These points are illustrated by the derivation of a new loss which is not convex, but does not compromise the computational tractability of classifier design, and is robust to the contamination of data with outliers. A new boosting algorithm, SavageBoost, is derived for the minimization of this loss. Experimental results show that it is indeed less sensitive to outliers than conventional methods, such as Ada, Real, or LogitBoost, and converges in fewer iterations. Hamed Masnadi-Shirazi, Nuno Vasconcelos |
NIPS | 2 |
| 2008 | Modeling, Clustering, and Segmenting Video with Mixtures of Dynamic TexturesabstractA dynamic texture is a spatio-temporal generative model for video, which represents video sequences as observations from a linear dynamical system. This work studies the mixture of dynamic textures, a statistical model for an ensemble of video sequences that is sampled from a finite collection of visual processes, each of which is a dynamic texture. An expectationmaximization (EM) algorithm is derived for learning the parameters of the model, and the model is related to previous works in linear systems, machine learning, time-series clustering, control theory, and computer vision. Through experimentation, it is shown that the mixture of dynamic textures is a suitable representation for both the appearance and dynamics of a variety of visual processes that have traditionally been challenging for computer vision (e.g. fire, steam, water, vehicle and pedestrian traffic, etc.). When compared with state-of-the-art methods in motion segmentation, including both temporal texture methods and traditional representations (e.g. optical flow or other localized motion representations), the mixture of dynamic textures achieves superior performance in the problems of clustering and segmenting video of such processes. Antoni B. Chan, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Classifying Video with Kernel Dynamic TexturesabstractThe dynamic texture is a stochastic video model that treats the video as a sample from a linear dynamical system. The simple model has been shown to be surprisingly useful in domains such as video synthesis, video segmentation, and video classification. However, one major disadvantage of the dynamic texture is that it can only model video where the motion is smooth, i.e. video textures where the pixel values change smoothly. In this work, we propose an extension of the dynamic texture to address this issue. Instead of learning a linear observation function with PCA, we learn a non-linear observation function using kernel-PCA. The resulting kernel dynamic texture is capable of modeling a wider range of video motion, such as chaotic motion (e.g. turbulent water) or camera motion (e.g. panning). We derive the necessary steps to compute the Martin distance between kernel dynamic textures, and then validate the new model through classification experiments on video containing camera motion. Antoni B. Chan, Nuno Vasconcelos |
CVPR | 2 |
| 2007 | Discriminant Interest Points are StableabstractA study of the performance of recently introduced discriminant methods for interest point detection [6,14] is presented. It has been previously shown that the resulting interest points are more informative for object recognition than those produced by the detectors currently used in computer vision. Little is, however, known about the properties of discriminant points with respect to the metrics, such as repeatability, that have been traditionally used to evaluate interest point detection. A thorough experimental evaluation of the stability of discriminant points is presented, and this stability compared to those of four popular methods. In particular, we consider image correspondence under geometric and photometric transformations, and extend the experimental protocol proposed by Mikolajczyk et al. [13] for the evaluation of stability with respect to such transformations. The extended protocol is suitable for the evaluation of both bottom-up and top-down (learned) detectors. It is shown that the stability of discriminant interest points is comparable, and frequently superior, to those of interest points produced by various currently popular techniques. Dashan Gao 0001, Nuno Vasconcelos |
CVPR | 2 |
| 2007 | Bottom-up saliency is a discriminant processabstractA bottom-up visual saliency detector is proposed, following a decision-theoretic formulation of saliency, previously developed for top-down processing (object recognition) [5]. The saliency of a given location of the visual field is defined as the power of a Gabor-like feature set to discriminate between the visual appearance of 1) a neighborhood centered at that location (the center) and 2) a neighborhood that surrounds it (the surround). Discrimination is defined in an information-theoretic sense and the optimal saliency detector derived for a class of stimuli that complies with known statistical properties of natural images, so as to achieve a computationally efficient solution. The resulting saliency detector is shown to replicate the fundamental properties of the psychophysics of pre-attentive vision, including stimulus pop-out, inability to detect feature conjunctions, asymmetries with respect to feature presence vs. absence, and compliance with Weber's law. It is also shown that the detector produces better predictions of human eye fixations than two previously proposed bottom-up saliency detectors. Dashan Gao 0001, Nuno Vasconcelos |
ICCV | 2 |
| 2007 | High Detection-rate Cascades for Real-Time Object DetectionabstractA new strategy is proposed for the design of cascaded object detectors of high detection-rate. The problem of jointly minimizing the false-positive rate and classification complexity of a cascade, given a constraint on its detection rate, is considered. It is shown that it reduces to the problem of minimizing false-positive rate given detection- rate and is, therefore, an instance of the classic problem of cost-sensitive learning. A cost-sensitive extension of boosting, denoted by asymmetric boosting, is introduced. It maintains a high detection-rate across the boosting iterations, and allows the design of cascaded detectors of high overall detection-rate. Experimental evaluation shows that, when compared to previous cascade design algorithms, the cascades produced by asymmetric boosting achieve significantly higher detection-rates, at the cost of a marginal increase in computation. Hamed Masnadi-Shirazi, Nuno Vasconcelos |
ICCV | 2 |
| 2007 | Direct convex relaxations of sparse SVMabstractAlthough support vector machines (SVMs) for binary classification give rise to a decision rule that only relies on a subset of the training data points (support vectors), it will in general be based on all available features in the input space. We propose two direct, novel convex relaxations of a non-convex sparse SVM formulation that explicitly constrains the cardinality of the vector of feature weights. One relaxation results in a quadratically-constrained quadratic program (QCQP), while the second is based on a semidefinite programming (SDP) relaxation. The QCQP formulation can be interpreted as applying an adaptive soft-threshold on the SVM hyperplane, while the SDP formulation learns a weighted inner-product (i.e. a kernel) that results in a sparse hyperplane. Experimental results show an increase in sparsity while conserving the generalization performance compared to a standard as well as a linear programming SVM. Antoni B. Chan, Nuno Vasconcelos, Gert R. G. Lanckriet |
ICML | 2 |
| 2007 | Asymmetric boostingabstractA cost-sensitive extension of boosting, denoted as asymmetric boosting, is presented. Unlike previous proposals, the new algorithm is derived from sound decision-theoretic principles, which exploit the statistical interpretation of boosting to determine a principled extension of the boosting loss. Similarly to AdaBoost, the cost-sensitive extension minimizes this loss by gradient descent on the functional space of convex combinations of weak learners, and produces large margin detectors. It is shown that asymmetric boosting is fully compatible with AdaBoost, in the sense that it becomes the latter when errors are weighted equally. Experimental evidence is provided to demonstrate the claims of cost-sensitivity and large margin. The algorithm is also applied to the computer vision problem of face detection, where it is shown to outperform a number of previous heuristic proposals for cost-sensitive boosting (AdaCost, CSB0, CSB1, CSB2, asymmetric-AdaBoost, AdaC1, AdaC2 and AdaC3). Hamed Masnadi-Shirazi, Nuno Vasconcelos |
ICML | 2 |
| 2007 | The discriminant center-surround hypothesis for bottom-up saliencyabstractThe classical hypothesis, that bottom-up saliency is a center-surround process, is combined with a more recent hypothesis that all saliency decisions are optimal in a decision-theoretic sense. The combined hypothesis is denoted as discriminant center-surround saliency, and the corresponding optimal saliency architecture is derived. This architecture equates the saliency of each image location to the discriminant power of a set of features with respect to the classification problem that opposes stimuli at center and surround, at that location. It is shown that the resulting saliency detector makes accurate quantitative predictions for various aspects of the psychophysics of human saliency, including non-linear properties beyond the reach of previous saliency models. Furthermore, it is shown that discriminant center-surround saliency can be easily generalized to various stimulus modalities (such as color, orientation and motion), and provides optimal solutions for many other saliency problems of interest for computer vision. Optimal solutions, under this hypothesis, are derived for a number of the former (including static natural images, dense motion fields, and even dynamic textures), and applied to a number of the latter (the prediction of human eye fixations, motion-based saliency in the presence of ego-motion, and motion-based saliency in the presence of highly dynamic backgrounds). In result, discriminant saliency is shown to predict eye fixations better than previous models, and produce background subtraction algorithms that outperform the state-of-the-art in computer vision. Dashan Gao 0001, Vijay Mahadevan, Nuno Vasconcelos |
NIPS | 3 |
| 2007 | Supervised Learning of Semantic Classes for Image Annotation and RetrievalabstractA probabilistic formulation for semantic image annotation and retrieval is proposed. Annotation and retrieval are posed as classification problems where each class is defined as the group of database images labeled with a common semantic label. It is shown that, by establishing this one-to-one correspondence between semantic labels and semantic classes, a minimum probability of error annotation and retrieval are feasible with algorithms that are 1) conceptually simple, 2) computationally efficient, and 3) do not require prior semantic segmentation of training images. In particular, images are represented as bags of localized feature vectors, a mixture density estimated for each image, and the mixtures associated with all images annotated with a common semantic label pooled into a density estimate for the corresponding semantic class. This pooling is justified by a multiple instance learning argument and performed efficiently with a hierarchical extension of expectation-maximization. The benefits of the supervised formulation over the more complex, and currently popular, joint modeling of semantic label and visual feature distributions are illustrated through theoretical arguments and extensive experiments. The supervised formulation is shown to achieve higher accuracy than various previously published methods at a fraction of their computational cost. Finally, the proposed method is shown to be fairly robust to parameter tuning. Gustavo Carneiro 0001, Antoni B. Chan, Pedro J. Moreno 0001, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2007 | Bridging the Gap: Query by Semantic ExampleabstractA combination of query-by-visual-example (QBVE) and semantic retrieval (SR), denoted as query-by-semantic-example (QBSE), is proposed. Images are labeled with respect to a vocabulary of visual concepts, as is usual in SR. Each image is then represented by a vector, referred to as a semantic multinomial, of posterior concept probabilities. Retrieval is based on the query-by-example paradigm: the user provides a query image, for which 1) a semantic multinomial is computed and 2) matched to those in the database. QBSE is shown to have two main properties of interest, one mostly practical and the other philosophical. From a practical standpoint, because it inherits the generalization ability of SR inside the space of known visual concepts (referred to as the semantic space) but performs much better outside of it, QBSE produces retrieval systems that are more accurate than what was previously possible. Philosophically, because it allows a direct comparison of visual and semantic representations under a common query paradigm, QBSE enables the design of experiments that explicitly test the value of semantic representations for image retrieval. An implementation of QBSE under the minimum probability of error (MPE) retrieval framework, previously applied with success to both QBVE and SR, is proposed, and used to demonstrate the two properties. In particular, an extensive objective comparison of QBSE with QBVE is presented, showing that the former significantly outperforms the latter both inside and outside the semantic space. By carefully controlling the structure of the semantic space, it is also shown that this improvement can only be attributed to the semantic nature of the representation on which QBSE is based. Nikhil Rasiwasia, Pedro J. Moreno 0001, Nuno Vasconcelos |
IEEE Trans. Multim. | 3 |
| 2006 | Weakly Supervised Top-down Image SegmentationabstractThere has recently been significant interest in top-down image segmentation methods, which incorporate the recognition of visual concepts as an intermediate step of segmentation. This work addresses the problem of top-down segmentation with weak supervision. Under this framework, learning does not require a set of manually segmented examples for each concept of interest, but simply a weakly labeled training set. This is a training set where images are annotated with a set of keywords describing their contents, but visual concepts are not explicitly segmented and no correspondence is specified between keywords and image regions. We demonstrate, both analytically and empirically, that weakly supervised segmentation is feasible when certain conditions hold. We also propose a simple weakly supervised segmentation algorithm that extends state-of-theart bottom-up segmentation methods in the direction of perceptually meaningful segmentation1. Manuela Vasconcelos, Nuno Vasconcelos, Gustavo Carneiro 0001 |
CVPR (1) | 2 |
| 2006 | Single Image Superresolution Based on Support Vector RegressionabstractSupport vector machine (SVM) regression is considered for a statistical method of single frame superresolution in both the spatial and discrete cosine transform (DCT) domains. As opposed to current classification techniques, regression allows considerably more freedom in the determination of missing high-resolution information. In addition, since SVM regression approaches the superresolution problem as an estimation problem with a criterion of image correctness rather than visual acceptableness, its optimization results have better mean-squared error. With the addition of structure in the DCT coefficients, DCT domain image superresolution is further improved Karl S. Ni, Sanjeev Kumar 0003, Nuno Vasconcelos, Truong Q. Nguyen |
ICASSP (2) | 3 |
| 2006 | Image Compression using Object-Based Regions of InterestabstractA new architecture for region of interest (ROI) image coding is proposed. ROIs are defined as image regions containing objects of interest, and an efficient algorithm proposed for the detection of such regions. This algorithm is based on the principle of discriminant saliency, under which salient regions are the image regions of strongest response for a set of features that discriminate the object class of interest from all others. The resulting ROI masks are fully compatible with the JPEG2000 standard. Experimental results are presented for images of complex scenes, which contain both objects and background clutter, demonstrating significant gains for object-based ROI coding, in terms of both subjective image quality and SNR. The proposed ROI-based coder is also shown to be trainable with small, informally collected, image collections (e.g. by simple Web search). This suggests the possibility of user-trained image coders. Sunhyoung Han, Nuno Vasconcelos |
ICIP | 2 |
| 2005 | Formulating Semantic Image Annotation as a Supervised Learning ProblemabstractWe introduce a new method to automatically annotate and retrieve images using a vocabulary of image semantics. The novel contributions include a discriminant formulation of the problem, a multiple instance learning solution that enables the estimation of concept probability distributions without prior image segmentation, and a hierarchical description of the density of each image class that enables very efficient training. Compared to current methods of image annotation and retrieval, the one now proposed has significantly smaller time complexity and better recognition performance. Specifically, its recognition complexity is O(C/spl times/R), where C is the number of classes (or image annotations) and R is the number of image regions, while the best results in the literature have complexity O(T/spl times/R), where T is the number of training images. Since the number of classes grows substantially slower than that of training images, the proposed method scales better during training, and processes test images faster This is illustrated through comparisons in terms of complexity, time, and recognition performance with current state-of-the-art methods. Gustavo Carneiro 0001, Nuno Vasconcelos |
CVPR (2) | 2 |
| 2005 | Probabilistic Kernels for the Classification of Auto-Regressive Visual ProcessesabstractWe present a framework for the classification of visual processes that are best modeled with spatio-temporal autoregressive models. The new framework combines the modeling power of a family of models known as dynamic textures and the generalization guarantees, for classification, of the support vector machine classifier. This combination is achieved by the derivation of a new probabilistic kernel based on the Kullback-Leibler divergence (KL) between Gauss-Markov processes. In particular, we derive the KL-kernel for dynamic textures in both 1) the image space, which describes both the motion and appearance components of the spatio-temporal process, and 2) the hidden state space, which describes the temporal component alone. Together, the two kernels cover a large variety of video classification problems, including the cases where classes can differ in both appearance and motion and the cases where appearance is similar for all classes and only motion is discriminant. Experimental evaluation on two databases shows that the new classifier achieves superior performance over existing solutions. 1. Antoni B. Chan, Nuno Vasconcelos |
CVPR (1) | 2 |
| 2005 | Integrated Learning of Saliency, Complex Features, and Object Detectors from Cluttered ScenesabstractA novel procedure for object detection from cluttered scenes is proposed. It consists of an integrated solution to the problems of learning 1) a saliency detection module tuned to a class of objects of interest, 2) a set of complex features that achieves the optimal trade-off, in a minimum probability of error sense, between discrimination and generalization ability, and 3) a large-margin object detector. All stages of the new procedure have some degree of biological motivation and this is shown to enable a computationally efficient solution that is scalable to problems containing large numbers of object classes. Experimental evidence is given in support of the arguments that different levels of feature complexity are optimal for different object classes, and that optimal features range from parts to templates, depending on the variability of the object class. Dashan Gao 0001, Nuno Vasconcelos |
CVPR (2) | 2 |
| 2005 | Mixtures of Dynamic TexturesabstractA dynamic texture is a linear dynamical system used to model a single video as a sample from a spatio-temporal stochastic process. In this work, we introduce the mixture of dynamic textures, which models a collection of videos consisting of different visual processes as samples from a set of dynamic textures. We derive the EM algorithm for learning a mixture of dynamic textures, and relate the learning algorithm and the dynamic texture mixture model to previous works. Finally, we demonstrate the applicability of the proposed model to problems that have traditionally been challenging for computer vision. Antoni B. Chan, Nuno Vasconcelos |
ICCV | 2 |
| 2005 | Layered Dynamic TexturesabstractA dynamic texture is a video model that treats a video as a sample from a spatio-temporal stochastic process, specifically a linear dynamical sys- tem. One problem associated with the dynamic texture is that it cannot model video where there are multiple regions of distinct motion. In this work, we introduce the layered dynamic texture model, which addresses this problem. We also introduce a variant of the model, and present the EM algorithm for learning each of the models. Finally, we demonstrate the efficacy of the proposed model for the tasks of segmentation and syn- thesis of video. Antoni B. Chan, Nuno Vasconcelos |
NIPS | 2 |
| 2005 | A database centric view of semantic image annotation and retrievalabstractWe introduce a new model for semantic annotation and retrieval from image databases. The new model is based on a probabilistic formulation that poses annotation and retrieval as classification problems, and produces solutions that are optimal in the minimum probability of error sense. It is also database centric, by establishing a one-to-one mapping between semantic classes and the groups of database images that share the associated semantic labels. In this work we show that, under the database centric probabilistic model, optimal annotation and retrieval can be implemented with algorithms that are conceptually simple, computationally efficient, and do not require prior semantic segmentation of training images. Due to its simplicity, the annotation and retrieval architecture is also amenable to sophisticated parameter tuning, a property that is exploited to investigate the role of feature selection in the design of optimal annotation and retrieval systems. Finally, we demonstrate the benefits of simply establishing a one-to-one mapping between keywords and the states of the semantic classification problem over the more complex, and currently popular, joint modeling of keyword and visual feature distributions. The database centric probabilistic retrieval model is compared to existing semantic labeling and retrieval methods, and shown to achieve higher accuracy than the previously best published results, at a fraction of their computational cost. Gustavo Carneiro 0001, Nuno Vasconcelos |
SIGIR | 2 |
| 2005 | Content-Based Image and Video Retrieval
Nuno Vasconcelos |
Signal Process. | 1 |
| 2005 | A multiresolution manifold distance for invariant image similarityabstractAccounting for spatial image transformations is a requirement for multimedia problems such as video classification and retrieval, face/object recognition or the creation of image mosaics from video sequences. We analyze a transformation invariant metric recently proposed in the machine learning literature to measure the distance between image manifolds - the tangent distance (TD) - and show that it is closely related to alignment techniques from the motion analysis literature. Exposing these relationships results in benefits for the two domains. On one hand, it allows leveraging on the knowledge acquired in the alignment literature to build better classifiers. On the other, it provides a new interpretation of alignment techniques as one component of a decomposition that has interesting properties for the classification of video. In particular, we embed the TD into a multiresolution framework that makes it significantly less prone to local minima. The new metric - multiresolution tangent distance (MRTD) - can be easily combined with robust estimation procedures, and exhibits significantly higher invariance to image transformations than the TD and the Euclidean distance (ED). For classification, this translates into significant improvements in face recognition accuracy. For video characterization, it leads to a decomposition of image dissimilarity into "differences due to camera motion" plus "differences due to scene activity" that is useful for classification. Experimental results on a movie database indicate that the distance could be used as a basis for the extraction of semantic primitives such as action and romance. Nuno Vasconcelos, Andy Lippman |
IEEE Trans. Multim. | 1 |
| 2004 | Scalable Discriminant Feature Selection for Image Retrieval and Recognition
Nuno Vasconcelos, Manuela Vasconcelos |
CVPR (2) | 1 |
| 2004 | The Kullback-Leibler Kernel as a Framework for Discriminant and Localized Representations for Visual Recognition
Nuno Vasconcelos, Purdy Ho, Pedro J. Moreno 0001 |
ECCV (3) | 1 |
| 2004 | Discriminant Saliency for Visual Recognition from Cluttered ScenesabstractSaliency mechanisms play an important role when visual recognition must be performed in cluttered scenes. We propose a computational defi- nition of saliency that deviates from existing models by equating saliency to discrimination. In particular, the salient attributes of a given visual class are defined as the features that enable best discrimination between that class and all other classes of recognition interest. It is shown that this definition leads to saliency algorithms of low complexity, that are scalable to large recognition problems, and is compatible with existing models of early biological vision. Experimental results demonstrating success in the context of challenging recognition problems are also pre- sented. Dashan Gao 0001, Nuno Vasconcelos |
NIPS | 2 |
| 2004 | On the efficient evaluation of probabilistic similarity functions for image retrievalabstractProbabilistic approaches are a promising solution to the image retrieval problem that, when compared to standard retrieval methods, can lead to a significant gain in retrieval accuracy. However, this occurs at the cost of a significant increase in computational complexity. In fact, closed-form solutions for probabilistic retrieval are currently available only for simple probabilistic models such as the Gaussian or the histogram. We analyze the case of mixture densities and exploit the asymptotic equivalence between likelihood and Kullback-Leibler (KL) divergence to derive solutions for these models. In particular, 1) we show that the divergence can be computed exactly for vector quantizers (VQs) and 2) has an approximate solution for Gauss mixtures (GMs) that, in high-dimensional feature spaces, introduces no significant degradation of the resulting similarity judgments. In both cases, the new solutions have closed-form and computational complexity equivalent to that of standard retrieval approaches. Nuno Vasconcelos |
IEEE Trans. Inf. Theory | 1 |
| 2003 | Feature Selection by Maximum Marginal Diversity: optimality and implications for visual recognitionabstractWe have recently shown that: 1) the infomax principle for the organization of perceptual systems leads to visual recognition architectures that are nearly optimal in the minimum Bayes error sense, and 2) a quantity which plays an important role in infomax solutions is the marginal diversity (MD) - the average distance between the class-conditional density of each feature and their mean. Since MD is a discriminant quantity and can be computed with great efficiency, the principle of maximum marginal diversity (MMD) was suggested for discriminant feature selection. In this paper, we study the optimality (in the infomax sense) of the MMD principle and analyze its effectiveness for feature selection in the context of visual recognition. In particular, 1) we derive a close form relation between the optimal infomax and MMD solutions, and 2) show that there is a family of classification problems for which the two are identical. Examination of this family in light of recent studies on the statistics of natural images suggests that the equivalence conditions are likely to hold for the problem of visual recognition. We present experimental evidence supporting the conclusions that: 1) MD is a good predictor for the recognition ability of a given set of features; 2) MMD produces features that are more discriminant than those obtained with currently predominant criteria such as energy compaction; and 3) the extracted features are detectors of visual attributes that are perceptually relevant for low-level image classification. Nuno Vasconcelos |
CVPR (1) | 1 |
| 2003 | A family of information-theoretic algorithms for low-complexity discriminant feature selection in image retrievalabstractFeature selection remains a challenging problem for image retrieval due to the massive amounts of data involved in the retrieval problem and the need to perform learning on-line in response to user-interaction. Existing feature selection techniques have limited ability to satisfy these requirements, due to significant complexity, or dependence on assumptions, e.g. Gaussianity, that are unrealistic for multimedia data. In this paper, we exploit some results connecting information-theoretic feature selection techniques and the minimization of the Bayes classification error to develop a new family of feature selection algorithms. This family is shown to enable the design of discriminant feature spaces with low complexity, and provide explicit control over the trade-off between complexity and optimality in the information-theoretic. Nuno Vasconcelos |
ICIP (3) | 1 |
| 2003 | A Kullback-Leibler Divergence Based Kernel for SVM Classification in Multimedia ApplicationsabstractOver the last years significant efforts have been made to develop kernels that can be applied to sequence data such as DNA, text, speech, video and images. The Fisher Kernel and similar variants have been suggested as good ways to combine an underlying generative model in the feature space and discriminant classifiers such as SVM’s. In this paper we sug- gest an alternative procedure to the Fisher kernel for systematically find- ing kernel functions that naturally handle variable length sequence data in multimedia domains. In particular for domains such as speech and images we explore the use of kernel functions that take full advantage of well known probabilistic models such as Gaussian Mixtures and sin- gle full covariance Gaussian models. We derive a kernel distance based on the Kullback-Leibler (KL) divergence between generative models. In effect our approach combines the best of both generative and discrim- inative methods and replaces the standard SVM kernels. We perform experiments on speaker identification/verification and image classifica- tion tasks and show that these new kernels have the best performance in speaker verification and mostly outperform the Fisher kernel based SVM’s and the generative classifiers in speaker identification and image classification. Pedro J. Moreno 0001, Purdy Ho, Nuno Vasconcelos |
NIPS | 3 |
| 2002 | What Is the Role of Independence for Visual Recognition?
Nuno Vasconcelos, Gustavo Carneiro 0001 |
ECCV (1) | 1 |
| 2002 | Exploiting group structure to improve retrieval accuracy and speed in image databasesabstractMost image retrieval systems perform a linear search over the database to find the closest match to a query. However, databases usually exhibit a natural grouping structure into content classes that can be exploited to improve retrieval precision and speed. We investigate methods that enable search at both the class and image level. It is shown that, through the combination of Bayesian averaging and hierarchical density estimation, it is possible to achieve significant gains in retrieval accuracy and speed, at the cost of a marginal increase in training complexity. The technique is also shown to enable the efficient design of semantic classifiers. Nuno Vasconcelos |
ICIP (1) | 1 |
| 2002 | Feature Selection by Maximum Marginal DiversityabstractWe address the question of feature selection in the context of visual recognition. It is shown that, besides efficient from a computational standpoint, the infomax principle is nearly optimal in the minimum Bayes error sense. The concept of marginal diversity is introduced, lead- ing to a generic principle for feature selection (the principle of maximum marginal diversity) of extreme computational simplicity. The relation- ships between infomax and the maximization of marginal diversity are identified, uncovering the existence of a family of classification proce- dures for which near optimal (in the Bayes error sense) feature selection does not require combinatorial search. Examination of this family in light of recent studies on the statistics of natural images suggests that visual recognition problems are a subset of it. Nuno Vasconcelos |
NIPS | 1 |
| 2001 | Image Indexing with Mixture HierarchiesabstractWe present an image indexing method based on a hierarchical description of the density of each of the image classes in a given database. The method is similar in spirit to traditional agglomerative clustering procedures but produces a complete mixture density, instead of a representative point, at each node of the indexing tree. Estimation of the density at a given node only requires knowledge of the mixture parameters of the children nodes, not the original data. The process is very flexible and efficient, therefore suited to problems involving large databases where existing groupings may have to be combined, or new groupings created, frequently. Experimental results show that the new indexing structure consistently outperforms a linear search when both efficiency and retrieval accuracy are taken into account. Nuno Vasconcelos |
CVPR (1) | 1 |
| 2001 | On the Complexity of Probabilistic Image RetrievalabstractProbabilistic image retrieval approaches can lead to significant gains over standard retrieval techniques. However, this occurs at the cost of a significant increase in computational complexity. In fact, closed-form solutions for probabilistic retrieval are currently available only for simple representations such as the Gaussian and the histogram. We analyze the case of mixture densities and exploit the asymptotic equivalence between likelihood and Kullback-Leibler divergence to derive solutions for these models. In particular, (1) we show that the divergence can be computed exactly for vector quantizers and, (2) has an approximate solution for Gaussian mixtures that introduces no significant degradation of the resulting similarity judgments. In both cases, the new solutions have closed-form and computational complexity equivalent to that of standard retrieval approaches, but significantly better retrieval performance. Nuno Vasconcelos |
ICCV | 1 |
| 2001 | Content-based retrieval from image databases: current solutions and future directionsabstractWe review recent advances in image retrieval. The two fundamental components of a retrieval system, representation and learning, are analyzed. Each component is decomposed into its constituent building blocks: features, feature representation, and similarity function for the representation; short and long-term procedures for learning. We identify a series of requirements for each of the sub-areas, e.g. optimality, invariance, perceptual relevance, computational tractability, and point out various approaches proposed to satisfy them. Several open problems are also identified. Nuno Vasconcelos, Murat Kunt |
ICIP (3) | 1 |
| 2001 | Empirical Bayesian Motion SegmentationabstractWe introduce an empirical Bayesian procedure for the simultaneous segmentation of an observed motion field and estimation of the hyperparameters of a Markov random field prior. The new approach exhibits the Bayesian appeal of incorporating prior beliefs, but requires only a qualitative description of the prior, avoiding the requirement for a quantitative specification of its parameters. This eliminates the need for trial-and-error strategies for the determination of these parameters and leads to better segmentations. Nuno Vasconcelos, Andy Lippman |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | A Probabilistic Architecture for Content-Based Image RetrievalabstractThe design of an effective architecture for content-based retrieval from visual libraries requires careful consideration of the interplay between feature selection, feature representation, and similarity metric. We present a solution where all the modules strive to optimize the same performance criteria: the probability of retrieval error. This solution consists of a Bayesian retrieval criteria (shown to generalize the most prevalent similarity metrics in current use) and an embedded mixture representation over a multiresolution feature space (shown to provide a good trade-off between retrieval accuracy, invariance, perceptual relevance of similarity, and complexity). The new representation extends standard models (histogram and Gaussian) by providing simultaneous support for high-dimensional features and multi-modal densities and performs well on color texture, and generic image databases. Nuno Vasconcelos, Andy Lippman |
CVPR | 1 |
| 2000 | Learning Over Multiple Temporal Scales in Image Databases
Nuno Vasconcelos, Andy Lippman |
ECCV (1) | 1 |
| 2000 | A Unifying View of Image SimilarityabstractWe study solutions to the problem of evaluating image similarity in the context of content-based image retrieval (CBIR). Retrieval is formulated as a classification problem, where the goal is to minimize probability of retrieval error. It is shown that this formulation establishes a common ground for comparing similarity functions, exposes assumptions hidden behind in most commonly used ones, enables a critical analysis of their relative merits, and determines the retrieval scenarios for which each may be most suited. We conclude that most of the current similarity functions are sub-optimal special cases of the Bayesian criteria that results from explicit minimization of error probability. Nuno Vasconcelos, Andy Lippman |
ICPR | 1 |
| 2000 | Bayesian Video Shot SegmentationabstractPrior knowledge about video structure can be used both as a means to improve the peiformance of content analysis and to extract features that allow semantic classification. We introduce statistical models for two important components of this structure, shot duration and activity, and demonstrate the usefulness of these models by introducing a Bayesian formulation for the shot segmentation problem. The new formulations is shown to extend standard thresholding methods in an adaptive and intuitive way, leading to improved segmentation accuracy. Nuno Vasconcelos, Andy Lippman |
NIPS | 1 |
| 2000 | Statistical models of video structure for content analysis and characterizationabstractContent structure plays an important role in the understanding of video. In this paper, we argue that knowledge about structure can be used both as a means to improve the performance of content analysis and to extract features that convey semantic information about the content. We introduce statistical models for two important components of this structure, shot duration and activity, and demonstrate the usefulness of these models with two practical applications. First, we develop a Bayesian formulation for the shot segmentation problem that is shown to extend the standard thresholding model in an adaptive and intuitive way, leading to improved segmentation accuracy. Second, by applying the transformation into the shot duration/activity feature space to a database of movie clips, we also illustrate how the Bayesian model captures semantic properties of the content. We suggest ways in which these properties can be used as a basis for intuitive content-based access to movie libraries. Nuno Vasconcelos, Andy Lippman |
IEEE Trans. Image Process. | 1 |
| 1999 | Learning from User Feedback in Image Retrieval Systems
Nuno Vasconcelos, Andy Lippman |
NIPS | 1 |
| 1998 | A Spatiotemporal Motion Model for Video SummarizationabstractThe compact description of a video sequence through a single image map and a dominant motion has applications in several domains, including video browsing and retrieval, compression, mosaicing, and visual summarization. Building such a representation requires the capability to register all the frames with respect to the dominant object in the scene, a task which has been, in the past, addressed through temporally localized motion estimates. In this paper, we show how the lack of temporal consistency associated with such estimates can undermine the validity of the dominant motion assumption, leading to oscillation between different scene interpretations and poor registration. To avoid this oscillation, we augment the motion model with a generic temporal constraint which increases the robustness against competing interpretations, leading to more meaningful content summarization. Nuno Vasconcelos, Andy Lippman |
CVPR | 1 |
| 1998 | A Bayesian Framework for Semantic Content CharacterizationabstractCurrent systems for content filtering, browsing, and retrieval rely on low-level image descriptors which are unintuitive for most users. In this paper, we propose an alternative framework that exploits the structured nature of most content sources to achieve semantic content characterization, and lead to much more meaningful user interaction. Computationally, this framework is based on the principles of Bayesian inference and can be implemented efficiently with Bayesian networks. As an illustration of its potential we apply it to the domain of movie databases. Nuno Vasconcelos, Andy Lippman |
CVPR | 1 |
| 1998 | A Bayesian Framework for Content-Based Indexing and RetrievalabstractSummary form only given. One of the important requirements for practical retrieval systems is the ability to jointly address the issues of indexing and compression. By formulating query by example as a problem of Bayesian inference and establishing a link between probability density estimation and vector quantization, we have previously introduced a representation that leads to very efficient procedures for indexing and retrieval directly in the compressed domain without compromise of the coding efficiency. In this paper, we build on the potential of the Bayesian formulation to support sophisticated inference, to incorporate this representation in a very flexible indexing and retrieval framework that (1) leads to intuitive retrieval procedures, (2) can integrate different content modalities to eliminate some of the strongest limitations of the query by example paradigm, and (3) supports statistical learning of all the model parameters and can, therefore, be trained automatically. Nuno Vasconcelos, Andy Lippman |
Data Compression Conference | 1 |
| 1998 | Bayesian Modeling of Video Editing and Structure: Semantic Features for Video Summarization and Browsing
Nuno Vasconcelos, Andy Lippman |
ICIP (3) | 1 |
| 1998 | Learning Mixture Hierarchies
Nuno Vasconcelos, Andy Lippman |
NIPS | 1 |
| 1997 | Empirical Bayesian EM-based Motion SegmentationabstractA recent trend in motion-based segmentation has been to rely on statistical procedures derived from expectation-maximization (EM) principles. EM-based approaches have various advantages for segmentation, such as proceeding by taking non-greedy soft decisions regarding the assignment of pixels to regions, or allowing the use of sophisticated priors capable of imposing spatial coherence on the segmentation. A practical difficulty with such priors is, however the determination of appropriate values for their parameters. The authors exploit the fact that the EM framework is itself suited for empirical Bayesian data analysis to develop an algorithm that finds the estimates of the prior parameters which best explain the observed data. Such an approach maintains the Bayesian appeal of incorporating prior beliefs, but requires only a qualitative description of the prior avoiding the requirement of a quantitative specification of its parameters. This eliminates the need for trial-and-error strategies for parameter determination and leads to better segmentation with fewer iterations. Nuno Vasconcelos, Andy Lippman |
CVPR | 1 |
| 1997 | Library-based Coding: a Representation for Efficient Video Compression and RetrievalabstractThe ubiquity of networking and computational capacity associated with the new communications media unveil a universe of new requirements for image representation. Among such requirements is the ability of the representation used for coding to support higher-level tasks such as content-based retrieval. We explore the relationships between probabilistic modeling and data compression to introduce a representation-library-based coding-which, by enabling retrieval in the compressed domain, satisfies this requirement. Because it contains an embedded probabilistic description of the source, this new representation allows the construction of good inference models without compromise of compression efficiency, leads to very efficient procedures for query and retrieval, and provides a framework for higher level tasks such as the analysis and classification of video shots. Nuno Vasconcelos, Andy Lippman |
Data Compression Conference | 1 |
| 1997 | Pre and Post-Filtering for Low Bit-Rate Video CodingabstractWe propose pre and post-filters for low bit rate video coding. The purpose of the former is to make the video sequence easier to encode, whereas the latter aims to remove coding artifacts. The proposed techniques are computationally efficient and lead to scalable architectures which (1) can handle varied types of video content and (2) can be adapted to the available bandwidth and computational resources. Simulation results using H.263 show that significant gains can be achieved by pre and post-filtering. Nuno Vasconcelos, Frédéric Dufaux |
ICIP (1) | 1 |
| 1997 | Towards Semantically Meaningful Feature Spaces for the Characterization of Video ContentabstractEfficient procedures for browsing, filtering, sorting or retrieving pictorial content require accurate content characterization. Of particular interest are representations based on semantically meaningful feature spaces, capable of capturing properties such as violence, sex or profanity. In this work we report on a first step towards this goal, the design of a stochastic model for video editing which provides a transformation from the image space to a low-dimensional feature space where categorization by degree of action can be easily accomplished. Nuno Vasconcelos, Andy Lippman |
ICIP (1) | 1 |
| 1997 | Content-Based Pre-Indexed VideoabstractThe viability of large distributed image databases is strongly dependent on the development of new image representations capable of providing support for extended functionality, directly in the compressed domain. We have previously introduced one such representation (library-based coding) which we now augment with statistical pre-indexing schemes, automatically built at the time of encoding, that provides several layers of content description allowing efficient content-based retrieval and summarization. Nuno Vasconcelos, Andy Lippman |
ICIP (2) | 1 |
| 1997 | Multiresolution Tangent Distance for Affine-invariant Classification
Nuno Vasconcelos, Andy Lippman |
NIPS | 1 |
| 1996 | Frame-free videoabstractCurrent digital video representations emphasize compression efficiency, lacking some of the flexibility required for interactive manipulation of digital bitstreams. We present a video representation which can encompass both space and time, providing a temporally coherent description of video sequences. The video sequence is segmented into its component objects, and the trajectory of each object throughout the sequence is described parametrically, according to a spatiotemporal motion model. Since the motion model is a continuous function of time, the video representation becomes frame-rate independent and the temporal resolution a user-definable parameter. i.e. the traditional sequence of frames, with the temporal structure hardcoded into the bitstream at the time of production, is replaced by a collection of scene snapshots assembled on the fly by the decoder. This enables random access and temporal scalability, the major building blocks for interactivity. Nuno Vasconcelos, Andy Lippman |
ICIP (3) | 1 |
| 1995 | Spatiotemporal model-based optic flow estimationabstractWe introduce a spatiotemporal model-based algorithm capable of providing estimates of optic flow which are coherent along a set of video frames. The algorithm is based on a spatiotemporal motion model that consists of a quadratic constraint in time and an affine constraint in space. Optic flow is computed through a delayed-decision process that incorporates knowledge about both image correlation along time, and the goodness of fit to the underlying motion-model. The temporal coherence and parametric nature of the recovered optic flow can facilitate interactive access to the video stream and improve the efficiency of tasks such as video compression, interpolation or classification. Nuno Vasconcelos, Andy Lippman |
ICIP | 1 |
| 1994 | Library-based image codingabstractWe explore a video coding algorithm that uses long term memory in addition to motion compensated prediction to build a model from video frames for coding. The long term memory, modeled on the cache used in operating systems and searched by a vector quantization formalism, holds recurring visual elements such as cyclically revealed information and actors that persist though several images. This video "library" exploits persistence to gain coding efficiency. Further, the existence of such a library can facilitate interactive and content-based operations on the video sequence such as browsing, searching by content, categorizing and enhancing the pictorial quality of the images.> Nuno Vasconcelos, Andy Lippman |
ICASSP (5) | 1 |