André Araújo 0001

dblp:177/1567-1 · also André F. Araújo 0001, André F. de Araújo 0001 · DBLP profile ↗
← Back
28ranked-venue papers
6as first author
14since 2021 · last 2025
0000-0002-4214-6185ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 17 · 13 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 TIPS: Text-Image Pretraining with Spatial awareness
abstract
While image-text representation learning has become very popular in recent years, existing models tend to lack spatial awareness and have limited direct applicability for dense understanding tasks. For this reason, self-supervised image-only pretraining is still the go-to method for many dense vision applications (e.g. depth estimation, semantic segmentation), despite the lack of explicit supervisory signals. In this paper, we close this gap between image-text and self-supervised learning, by proposing a novel general-purpose image-text model, which can be effectively used off the shelf for dense and global vision tasks. Our method, which we refer to as Text-Image Pretraining with Spatial awareness (TIPS), leverages two simple and effective insights. First, on textual supervision: we reveal that replacing noisy web image captions by synthetically generated textual descriptions boosts dense understanding performance significantly, due to a much richer signal for learning spatially aware representations. We propose an adapted training method that combines noisy and synthetic captions, resulting in improvements across both dense and global understanding tasks. Second, on the learning technique: we propose to combine contrastive image-text learning with self-supervised masked image modeling, to encourage spatial coherence, unlocking substantial enhancements for downstream applications. Building on these two ideas, we scale our model using the transformer architecture, trained on a curated set of public images. Our experiments are conducted on $8$ tasks involving $16$ datasets in total, demonstrating strong off-the-shelf performance on both dense and global understanding, for several image-only and image-text tasks. Code and models are released at https://github.com/google-deepmind/tips .
Kevis-Kokitsi Maninis, Kaifeng Chen, Soham Ghosh 0001, Arjun Karpur, Koert Chen, Bingyi Cao, Daniel Salz, Guangxing Han, Jan Dlabal, Dan Gnanapragasam, Mojtaba Seyedhosseini, Howard Zhou, André Araújo 0001
ICLR14
2025 VESSA: Video-based objEct-centric Self-Supervised Adaptation for Visual Foundation Models
abstract
Foundation models have advanced computer vision by enabling strong performance across diverse tasks through large-scale pretraining and supervised fine-tuning. However, they may underperform in domains with distribution shifts and scarce labels, where supervised fine-tuning may be infeasible. While continued self-supervised learning for model adaptation is common for generative language models, this strategy has not proven effective for vision-centric encoder models. To address this challenge, we introduce a novel formulation of self-supervised fine-tuning for vision foundation models, where the model is adapted to a new domain without requiring annotations, leveraging only short multi-view object-centric videos. Our method is referred to as VESSA: **V**ideo-based obj**E**ct-centric **S**elf-**S**upervised **A**daptation for visual foundation models. VESSA's training technique is based on a self-distillation paradigm, where it is critical to carefully tune prediction heads and deploy parameter-efficient adaptation techniques – otherwise, the model may quickly forget its pretrained knowledge and reach a degraded state. VESSA benefits significantly from multi-view object observations sourced from different frames in an object-centric video, efficiently learning robustness to varied capture conditions, without the need of annotations. Through comprehensive experiments with 3 vision foundation models on 2 datasets, VESSA demonstrates consistent improvements in downstream classification tasks, compared to the base models and previous adaptation methods. Code is publicly available at https://github.com/jesimonbarreto/VESSA.
Jesimon Barreto Santos, Carlos Antônio Caetano Jr., André Araújo 0001, William Robson Schwartz
NeurIPS3
2025 Optimization of Rank Losses for Image Retrieval
abstract
In image retrieval, standard evaluation metrics rely on score ranking, e.g. average precision (AP), recall at k (R@k), normalized discounted cumulative gain (NDCG). In this work, we introduce a general framework for robust and decomposable rank losses optimization. It addresses two major challenges for end-to-end training of deep neural networks with rank losses: non-differentiability and non-decomposability. First, we propose a general surrogate for ranking operator, SupRank, that is amenable to stochastic gradient descent. It provides an upperbound for rank losses and ensures robust training. Second, we use a simple yet effective loss function to reduce the decomposability gap between the averaged batch approximation of ranking losses and their values on the whole training set. We apply our framework to two standard metrics for image retrieval: AP and R@k. Additionally, we apply our framework to hierarchical image retrieval. We introduce an extension of AP, the hierarchical average precision $\mathcal {H}{\mathrm -AP}$H- AP , and optimize it as well as the NDCG. Finally, we create the first hierarchical landmarks retrieval dataset. We use a semi-automatic pipeline to create hierarchical labels, extending the large scale Google Landmarks v2 dataset.
Elias Ramzi, Nicolas Audebert, Clément Rambour, André Araújo 0001, Xavier Bitot, Nicolas Thome
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 LFM-3D: Learnable Feature Matching Across Wide Baselines Using 3D Signals
abstract
Finding localized correspondences across different images of the same object is crucial to understand its geometry. In recent years, this problem has seen remarkable progress with the advent of deep learning-based local image features and learnable matchers. Still, learnable matchers often underperform when there exists only small regions of co-visibility between image pairs (i.e. wide camera baselines). To address this problem, we leverage recent progress in coarse single-view geometry estimation methods. We propose LFM-3D, a Learnable Feature Matching framework that uses models based on graph neural networks and enhances their capabilities by integrating noisy, estimated 3D signals to boost correspondence estimation. When integrating 3D signals into the matcher model, we show that a suitable positional encoding is critical to effectively make use of the low-dimensional 3D information. We experiment with two different 3D signals - normalized object coordinates and monocular depth estimates - and evaluate our method on large-scale (synthetic and real) datasets containing object-centric image pairs across wide baselines. We observe strong feature matching improvements compared to 2D-only methods, with up to +6% total recall and +28% precision at fixed recall. Additionally, we demonstrate that the resulting improved correspondences lead to much higher relative posing accuracy for in-the-wild image pairs - up to 8.6% compared to the 2D-only approach.
Arjun Karpur, Guilherme Perrotta, Ricardo Martin-Brualla, Howard Zhou, André Araújo 0001
3DV5
2024 OmniGlue: Generalizable Feature Matching with Foundation Model Guidance
abstract
The image matching field has been witnessing a contin-uous emergence of novel learnable feature matching techniques, with ever-improving performance on conventional benchmarks. However, our investigation shows that de-spite these gains, their potential for real-world applications is restricted by their limited generalization capabili-ties to novel image domains. In this paper, we introduce OmniGlue, the first learnable image matcher that is de-signed with generalization as a core principle. OmniGlue leverages broad knowledge from a vision foundation model to guide the feature matching process, boosting general-ization to domains not seen at training time. Addition-ally, we propose a novel keypoint position-guided attention mechanism which disentangles spatial and appear-ance information, leading to enhanced matching descrip-tors. We perform comprehensive experiments on a suite of 7 datasets with varied image domains, including scene-level, object-centric and aerial images. OmniGlue's novel components lead to relative gains on unseen domains of 20.9% with respect to a directly comparable reference model, while also outperforming the recent LightGlue method by 9.5% relatively. Code and model can be found at https://hwjiang151o.github.io/OmniGlue.
Hanwen Jiang, Arjun Karpur, Bingyi Cao, Qixing Huang, André Araújo 0001
CVPR5
2024 XFeat: Accelerated Features for Lightweight Image Matching
abstract
We introduce a lightweight and accurate architecture for resource-efficient visual correspondence. Our method, dubbed XFeat (Accelerated Features), revisits fundamen-tal design choices in convolutional neural networks for de-tecting, extracting, and matching local features. Our new model satisfies a critical need for fast and robust algorithms suitable to resource-limited devices. In particular, accu-rate image matching requires sufficiently large image res-olutions -for this reason, we keep the resolution as large as possible while limiting the number of channels in the net-work. Besides, our model is designed to offer the choice of matching at the sparse or semi-dense levels, each of which may be more suitable for different downstream applications, such as visual navigation and augmented reality. Our model is the first to offer semi-dense matching efficiently, leveraging a novel match refinement module that relies on coarse local descriptors. XFeat is versatile and hardware-independent, surpassing current deep learning-based local features in speed (up to 5xfaster) with comparable or better accuracy, proven in pose estimation and visual localization. We showcase it running in real-time on an inexpensive lap-top CPU without specialized hardware optimizations. Code and weights are available at verlab.dcc.ufmg.br/descriptors/xfeat_cvpr24.
Guilherme A. Potje, Felipe C. Chamone, André Araújo 0001, Renato Martins, Erickson R. Nascimento
CVPR3
2024 UDON: Universal Dynamic Online distillatioN for generic image representations
abstract
Universal image representations are critical in enabling real-world fine-grained and instance-level recognition applications, where objects and entities from any domain must be identified at large scale. Despite recent advances, existing methods fail to capture important domain-specific knowledge, while also ignoring differences in data distribution across different domains. This leads to a large performance gap between efficient universal solutions and expensive approaches utilising a collection of specialist models, one for each domain. In this work, we make significant strides towards closing this gap, by introducing a new learning technique, dubbed UDON (Universal Dynamic Online distillatioN). UDON employs multi-teacher distillation, where each teacher is specialized in one domain, to transfer detailed domain-specific knowledge into the student universal embedding. UDON's distillation approach is not only effective, but also very efficient, by sharing most model parameters between the student and all teachers, where all models are jointly trained in an online manner. UDON also comprises a sampling technique which adapts the training process to dynamically allocate batches to domains which are learned slower and require more frequent processing. This boosts significantly the learning of complex domains which are characterised by a large number of classes and long-tail distributions. With comprehensive experiments, we validate each component of UDON, and showcase significant improvements over the state of the art in the recent UnED benchmark. Code: https://github.com/nikosips/UDON.
Nikolaos-Antonios Ypsilantis, Kaifeng Chen, André Araújo 0001, Ondrej Chum
NeurIPS3
2023 Enhancing Deformable Local Features by Jointly Learning to Detect and Describe Keypoints
abstract
Local feature extraction is a standard approach in computer vision for tackling important tasks such as image matching and retrieval. The core assumption of most methods is that images undergo affine transformations, disregarding more complicated effects such as non-rigid deformations. Furthermore, incipient works tailored for non-rigid correspondence still rely on keypoint detectors designed for rigid transformations, hindering performance due to the limitations of the detector. We propose DALF (Deformation-Aware Local Features), a novel deformation-aware network for jointly detecting and describing keypoints, to handle the challenging problem of matching deformable surfaces. All network components work cooperatively through a feature fusion approach that enforces the descriptors' distinctiveness and invariance. Experiments using real deforming objects showcase the superiority of our method, where it delivers 8% improvement in matching scores compared to the previous best results. Our approach also enhances the performance of two real-world applications: deformable object retrieval and non-rigid 3D surface registration. Code for training, inference, and applications are publicly available at verlab.dcc.ufmg.br/descriptors/dalf_cvpr23.
Guilherme A. Potje, Felipe C. Chamone, André Araújo 0001, Renato Martins, Erickson R. Nascimento
CVPR3
2023 Yes, we CANN: Constrained Approximate Nearest Neighbors for local feature-based visual localization
abstract
Large-scale visual localization systems continue to rely on 3D point clouds built from image collections using structure-from-motion. While the 3D points in these models are represented using local image features, directly matching a query image’s local features against the point cloud is challenging due to the scale of the nearest-neighbor search problem. Many recent approaches to visual localization have thus proposed a hybrid method, where first a global (per image) embedding is used to retrieve a small subset of database images, and local features of the query are matched only against those. It seems to have become common belief that global embeddings are critical for said image-retrieval in visual localization, despite the significant downside of having to compute two feature types for each query image. In this paper, we take a step back from this assumption and propose Constrained Approximate Nearest Neighbors (CANN), a joint solution of k-nearest-neighbors across both the geometry and appearance space using only local features. We first derive the theoretical foundation for k-nearest-neighbor retrieval across multiple metrics and then showcase how CANN improves visual localization. Our experiments on public localization benchmarks demonstrate that our method significantly outperforms both state-of-the-art global feature-based retrieval and approaches using local feature aggregation schemes. Moreover, it is an order of magnitude faster in both index and query time than feature aggregation schemes for these datasets. Code will be released.
Dror Aiger, André Araújo 0001, Simon Lynen
ICCV2
2023 Encyclopedic VQA: Visual questions about detailed properties of fine-grained categories
abstract
We propose Encyclopedic-VQA, a large scale visual question answering (VQA) dataset featuring visual questions about detailed properties of fine-grained categories and instances. It contains 221k unique question+answer pairs each matched with (up to) 5 images, resulting in a total of 1M VQA samples. Moreover, our dataset comes with a controlled knowledge base derived from Wikipedia, marking the evidence to support each answer. Empirically, we show that our dataset poses a hard challenge for large vision+language models as they perform poorly on our dataset: PaLI [12] is state-of-the-art on OK-VQA [35], yet it only achieves 13.0% accuracy on our dataset. Moreover, we experimentally show that progress on answering our encyclopedic questions can be achieved by augmenting large models with a mechanism that retrieves relevant information from the knowledge base. An oracle experiment with perfect retrieval achieves 87.0% accuracy on the single-hop portion of our dataset, and an automatic retrieval-augmented prototype yields 48.8%. We believe that our dataset1enables future research on retrieval-augmented vision+language models.
Thomas Mensink, Jasper R. R. Uijlings, Lluís Castrejón, Arushi Goel, Felipe C. Chamone, Howard Zhou, Fei Sha, André Araújo 0001, Vittorio Ferrari
ICCV8
2023 Global Features are All You Need for Image Retrieval and Reranking
abstract
Image retrieval systems conventionally use a two-stage paradigm, leveraging global features for initial retrieval and local features for reranking. However, the scalability of this method is often limited due to the significant storage and computation cost incurred by local feature matching in the reranking stage. In this paper, we present SuperGlobal, a novel approach that exclusively employs global features for both stages, improving efficiency without sacrificing accuracy. SuperGlobal introduces key enhancements to the retrieval system, specifically focusing on the global feature extraction and reranking processes. For extraction, we identify sub-optimal performance when the widely-used ArcFace loss and Generalized Mean (GeM) pooling methods are combined and propose several new modules to improve GeM pooling. In the reranking stage, we introduce a novel method to update the global features of the query and top-ranked images by only considering feature refinement with a small set of images, thus being very compute and memory efficient. Our experiments demonstrate substantial improvements compared to the state of the art in standard benchmarks. Notably, on the Revisited Oxford+1M Hard dataset, our single-stage results improve by 7.1%, while our two-stage gain reaches 3.7% with a strong 64,865× speedup. Our two-stage system surpasses the current single-stage state-of-the-art by 16.3%, offering a scalable, accurate alternative for high-performing image retrieval systems with minimal time overhead.Code: https://github.com/ShihaoShao-GH/SuperGlobal.
Shihao Shao, Kaifeng Chen, Arjun Karpur, Qinghua Cui, André Araújo 0001, Bingyi Cao
ICCV5
2023 Towards Universal Image Embeddings: A Large-Scale Dataset and Challenge for Generic Image Representations
abstract
Fine-grained and instance-level recognition methods are commonly trained and evaluated on specific domains, in a model per domain scenario. Such an approach, however, is impractical in real large-scale applications. In this work, we address the problem of universal image embedding, where a single universal model is trained and used in multiple domains. First, we leverage existing domain-specific datasets to carefully construct a new large-scale public benchmark for the evaluation of universal image embeddings, with 241k query images, 1.4M index images and 2.8M training images across 8 different domains and 349k classes. We define suitable metrics, training and evaluation protocols to foster future research in this area. Second, we provide a comprehensive experimental evaluation on the new dataset, demonstrating that existing approaches and simplistic extensions lead to worse performance than an assembly of models trained for each domain separately. Finally, we conducted a public research competition on this topic, leveraging industrial datasets, which attracted the participation of more than 1k teams world-wide. This exercise generated many interesting research ideas and findings which we present in detail. Project webpage: https://cmp.felk.cvut.cz/univ_emb/
Nikolaos-Antonios Ypsilantis, Kaifeng Chen, Bingyi Cao, Mário Lipovský, Pelin Dogan-Schönberger, Grzegorz Makosa, Boris Bluntschli, Mojtaba Seyedhosseini, Ondrej Chum, André Araújo 0001
ICCV10
2023 NAVI: Category-Agnostic Image Collections with High-Quality 3D Shape and Pose Annotations
abstract
Recent advances in neural reconstruction enable high-quality 3D object reconstruction from casually captured image collections. Current techniques mostly analyze their progress on relatively simple image collections where SfM techniques can provide ground-truth (GT) camera poses. We note that SfM techniques tend to fail on in-the-wild image collections such as image search results with varying backgrounds and illuminations. To enable systematic research progress on 3D reconstruction from casual image captures, we propose `NAVI': a new dataset of category-agnostic image collections of objects with high-quality 3D scans along with per-image 2D-3D alignments providing near-perfect GT camera parameters. These 2D-3D alignments allow us to extract accurate derivative annotations such as dense pixel correspondences, depth and segmentation maps. We demonstrate the use of NAVI image collections on different problem settings and show that NAVI enables more thorough evaluations that were not possible with existing datasets. We believe NAVI is beneficial for systematic research progress on 3D reconstruction and correspondence estimation.
Varun Jampani, Kevis-Kokitsi Maninis, Andreas Engelhardt, Arjun Karpur, Karen Truong, Kyle Sargent, Stefan Popov, André Araújo 0001, Ricardo Martin-Brualla, Kaushal Patel, Daniel Vlasic, Vittorio Ferrari, Ameesh Makadia, Ce Liu 0001, Yuanzhen Li, Howard Zhou
NeurIPS8
2021 Class-Balanced Distillation for Long-Tailed Visual Recognition
Ahmet Iscen, André Araújo 0001, Boqing Gong, Cordelia Schmid
BMVC2
2020 Google Landmarks Dataset v2 - A Large-Scale Benchmark for Instance-Level Recognition and Retrieval
abstract
While image retrieval and instance recognition techniques are progressing rapidly, there is a need for challenging datasets to accurately measure their performance -- while posing novel challenges that are relevant for practical applications. We introduce the Google Landmarks Dataset v2 (GLDv2), a new benchmark for large-scale, fine-grained instance recognition and image retrieval in the domain of human-made and natural landmarks. GLDv2 is the largest such dataset to date by a large margin, including over 5M images and 200k distinct instance labels. Its test set consists of 118k images with ground truth annotations for both the retrieval and recognition tasks. The ground truth construction involved over 800 hours of human annotator work. Our new dataset has several challenging properties inspired by real-world applications that previous datasets did not consider: An extremely long-tailed class distribution, a large fraction of out-of-domain test photos and large intra-class variability. The dataset is sourced from Wikimedia Commons, the world's largest crowdsourced collection of landmark photos. We provide baseline results for both recognition and retrieval tasks based on state-of-the-art methods as well as competitive results from a public challenge. We further demonstrate the suitability of the dataset for transfer learning by showing that image embeddings trained on it achieve competitive retrieval performance on independent datasets. The dataset images, ground-truth and metric scoring code are available at https://github.com/cvdfoundation/google-landmark.
Tobias Weyand, André Araújo 0001, Bingyi Cao, Jack Sim
CVPR2
2020 Unifying Deep Local and Global Features for Image Search
Bingyi Cao, André Araújo 0001, Jack Sim
ECCV (20)2
2019 Detect-To-Retrieve: Efficient Regional Aggregation for Image Search
abstract
Retrieving object instances among cluttered scenes efficiently requires compact yet comprehensive regional image representations. Intuitively, object semantics can help build the index that focuses on the most relevant regions. However, due to the lack of bounding-box datasets for objects of interest among retrieval benchmarks, most recent work on regional representations has focused on either uniform or class-agnostic region selection. In this paper, we first fill the void by providing a new dataset of landmark bounding boxes, based on the Google Landmarks dataset, that includes 94k images with manually curated boxes from 15k unique landmarks. Then, we demonstrate how a trained landmark detector, using our new dataset, can be leveraged to index image regions and improve retrieval accuracy while being much more efficient than existing regional methods. In addition, we introduce a novel regional aggregated selective match kernel (R-ASMK) to effectively combine information from detected regions into an improved holistic image representation. R-ASMK boosts image retrieval accuracy substantially with no dimensionality increase, while even outperforming systems that index image regions independently. Our complete image retrieval system improves upon the previous state-of-the-art by significant margins on the Revisited Oxford and Paris datasets. Code and data will be released.
Marvin Teichmann, André Araújo 0001, Menglong Zhu, Jack Sim
CVPR2
2018 Large-Scale Video Retrieval Using Image Queries
abstract
Retrieving videos from large repositories using image queries is important for many applications, such as brand monitoring or content linking. We introduce a new retrieval architecture, in which the image query can be compared directly with database videos - significantly improving retrieval scalability compared with a baseline system that searches the database on a video frame level. Matching an image to a video is an inherently asymmetric problem. We propose an asymmetric comparison technique for Fisher vectors and systematically explore query or database items with varying amounts of clutter, showing the benefits of the proposed technique. We then propose novel video descriptors that can be compared directly with image descriptors. We start by constructing Fisher vectors for video segments, by exploring different aggregation techniques. For a database of lecture videos, such methods obtain a two orders of magnitude compression gain with respect to a frame-based scheme, with no loss in retrieval accuracy. Then, we consider the design of video descriptors, which combine Fisher embedding with hashing techniques, in a flexible framework based on Bloom filters. Large-scale experiments using three datasets show that this technique enables faster and more memory-efficient retrieval, compared with a frame-based method, with similar accuracy. The proposed techniques are further compared against pre-trained convolutional neural network features, outperforming them on three datasets by a substantial margin.
André Araújo 0001, Bernd Girod
IEEE Trans. Circuits Syst. Video Technol.1
2017 Effective Fisher vector aggregation for 3D object retrieval
abstract
We formulate the task of 3D object retrieval as a visual search problem where a database containing videos of objects captured manually from different viewpoints is queried using a single image. We propose to aggregate visual information of similar views and use the Fisher vector (FV) framework to compactly represent a database of objects. Large-scale experiments on an existing video dataset that we complemented with image queries, shows that our aggregation schemes significantly outperform standard retrieval techniques. When representing our database with only 4 FVs per object, our approach performs with a mean average precision (mAP) of 73.0% on our dataset while the baseline (no aggregation) only reaches a mAP of 43.8%. It can also reach a 72.0% mAP level with a 10× smaller database than the baseline.
Jean-Baptiste Boin, André Araújo 0001, Lamberto Ballan, Bernd Girod
ICASSP2
2017 Large-Scale Image Retrieval with Attentive Deep Local Features
abstract
We propose an attentive local feature descriptor suitable for large-scale image retrieval, referred to as DELE (DEep Local Feature). The new feature is based on convolutional neural networks, which are trained only with image-level annotations on a landmark image dataset. To identify semantically useful local features for image retrieval, we also propose an attention mechanism for key point selection, which shares most network layers with the descriptor. This frame-work can be used for image retrieval as a drop-in replacement for other keypoint detectors and descriptors, enabling more accurate feature matching and geometric verification. Our system produces reliable confidence scores to reject false positives–in particular, it is robust against queries that have no correct match in the database. To evaluate the proposed descriptor, we introduce a new large-scale dataset, referred to as Google-Landmarks dataset, which involves challenges in both database and query such as background clutter, partial occlusion, multiple landmarks, objects in variable scales, etc. We show that DELE outperforms the state-of-the-art global and local descriptors in the large-scale setting by significant margins.
Hyeonwoo Noh, André Araújo 0001, Jack Sim, Tobias Weyand, Bohyung Han
ICCV2
2016 Modeling the impact of keypoint detection errors on local descriptor similarity
abstract
This paper presents a mathematical analysis of the impact of key-point detection errors on the similarity of local image descriptors that are based on histogram of gradients. First, we derive a closed-form expression for the Lp distance between two descriptors, for general translation, scale and orientation detection errors. Second, we introduce a detailed analysis for the special case where translation errors dominate, using the L2 distance. We show that the individual components which form the squared L2 distance can be approximated using Gamma distributions whose parameters are computed in closed-form by our model. We obtain approximate closed-form expressions for the expected squared L2 distances when translation errors are fixed or uniformly distributed. Finally, these models are validated using image patches extracted from two standard image retrieval datasets, by comparing the predicted distributions to the ground-truth.
André Araújo 0001, Haricharan Lakshman, Roland Angst, Bernd Girod
ICIP1
2015 Temporal aggregation for large-scale query-by-image video retrieval
abstract
We address the challenge of using image queries to retrieve video clips from a large database. Using binarized Fisher Vectors as global signatures, we present three novel contributions. First, an asymmetric comparison scheme for binarized Fisher Vectors is shown to boost retrieval performance by 0.27 mean Average Precision, exploiting the fact that query images contain much less clutter than database videos. Second, aggregation of frame-based local features over shots is shown to achieve retrieval performance comparable to aggregation of those local features over single frames, while reducing retrieval latency and memory requirements by more than 3X. Several shot aggregation strategies are compared and results indicate that most perform equally well. Third, aggregation over scenes, in combination with shot signatures, is shown to achieve one order of magnitude faster retrieval at comparable performance. Scene aggregation also outperforms the recently proposed aggregation in random groups.
André Araújo 0001, Jason Chaves, Roland Angst, Bernd Girod
ICIP1
2015 Stanford I2V: a news video dataset for query-by-image experiments
abstract
Reproducible research in the area of visual search depends on the availability of large annotated datasets. In this paper, we address the problem of querying a video database by images that might share some contents with one or more video clips. We present a new large dataset, called Stanford I2V. We have collected more than 3; 800 hours of newscast videos and annotated more than 200 ground-truth queries. In the following, the dataset is described in detail, the collection methodology is outlined and retrieval performance for a benchmark algorithm is presented. These results may serve as a baseline for future research and provide an example of the intended use of the Stanford I2V dataset. The dataset can be downloaded at http://purl.stanford.edu/zx935qw7203.
André Araújo 0001, Jason Chaves, David M. Chen, Roland Angst, Bernd Girod
MMSys1
2014 Interframe Coding of Global Image Signatures for Mobile Augmented Reality
abstract
For mobile augmented reality, an image captured by a mobile device's camera is often compared against a database hosted on a remote server to recognize objects in the image. It is critically important that the amount of data transmitted over the network is as small as possible to reduce the system latency. A low bitrate global signature for still images has been previously shown to achieve high-accuracy image retrieval. In this paper, we develop new methods for interframe coding of a continuous stream of global signatures that can reduce the bitrate by nearly two orders of magnitude compared to independent coding of these global signatures, while achieving the same or better image retrieval accuracy. The global signatures are constructed in an embedded data structure that offers rate scalability. The usage of these new coding methods and the embedded data structure allows the streaming of high-quality global signatures at a bitrate that is less than 2 kbps. Furthermore, a statistical analysis of the retrieval and coding performance is performed to understand the trade off between bitrate and image retrieval accuracy and explain why interframe coding of global signatures substantially outperforms independent coding.
David M. Chen, Mina Makar, André Araújo 0001, Bernd Girod
DCC3
2014 Efficient video search using image queries
abstract
We study the challenges of image-based retrieval when the database consists of videos. This variation of visual search is important for a broad range of applications that require indexing video databases based on their visual contents. We present new solutions to reduce storage requirements, while at the same time improving video search quality. The video database is preprocessed to find different appearances of the same visual elements, and build robust descriptors. Compression algorithms are developed to reduce system's storage requirements. We introduce a dataset of CNN broadcasts and queries that include photos taken with mobile phones and images of objects. Our experiments include pairwise matching and retrieval scenarios. We demonstrate one order of magnitude storage reduction and search quality improvements of up to 12% in mean average precision, compared to a baseline system that does not make use of our techniques.
André Araújo 0001, Mina Makar, Vijay Chandrasekhar 0001, David M. Chen, Sam S. Tsai, Huizhong Chen, Roland Angst, Bernd Girod
ICIP1
2014 Real-time query-by-image video search system
abstract
We demonstrate a novel multimedia system that continuously indexes videos and enables real-time search using images, with a broad range of potential applications. Television shows are recorded and indexed continuously, and iconic images from recent events are discovered automatically. Users can query an uploaded image or an image in the web. When a result is served, the user can play the video clip from the beginning or from the point in time where the retrieved image was found.
André Araújo 0001, David M. Chen, Peter Vajda, Bernd Girod
ACM Multimedia1
2013 EigenNews: a personalized news video delivery platform
abstract
We demonstrate EigenNews, a personalized television news system. Upon visiting the EigenNews website, a user is shown a variety of news videos which have been automatically selected based on her individual preferences. These videos are extracted from 16 continually recorded television programs using a multimodal segmentation algorithm. Relevant metadata for each video are generated by linking videos to online news articles. Selected news videos can be watched in three different layouts and on various devices.
Matt C. Yu, Peter Vajda, David M. Chen, Sam S. Tsai, Maryam Daneshi, André Araújo 0001, Huizhong Chen, Bernd Girod
ACM Multimedia6
2011 Compression of VQM features for low bit-rate video quality monitoring
abstract
Reduced reference video quality assessment techniques provide a practical and convenient way of evaluating the quality of a processed video. In this paper, we propose a method to efficiently compress standardized VQM (Video Quality Model) [1] features to bit-rates that are small relative to the transmitted video. This is achieved through two stages of compression. In the first stage, we remove the redundancy in the features by only transmitting the necessary original video features at the lowest acceptable resolution for the calculation of the final VQM value. The second stage involves using the features of the processed video at the receiver as side-information for efficient entropy coding and reconstruction of the original video features. Experimental results demonstrate that our approach achieves high compression ratios of more than 30× with small error in the final VQM values.
Mina Makar, Yao-Chung Lin, André Araújo 0001, Bernd Girod
MMSP3