EDBT 2026 Demo / reviewers in the wild / expert
Rogério Feris
dblp:f/RogerioSchmidtFeris · also Rogério Schmidt Feris
· DBLP profile ↗
145ranked-venue papers
12as first author
64since 2021 · last 2025
0000-0001-6399-0679ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 110 · 11 first-author · 42 since 2021Artificial intelligence and machine learning · 106 · 5 first-author · 57 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?abstractWe propose Omni-R1 which fine-tunes a recent multi-modal LLM, Qwen2.5-Omni, on an audio question answering dataset with the reinforcement learning method GRPO. This leads to new State-of-the-Art performance on the recent MMAU and MMAR benchmarks. On MMAU, Omni-R1 achieves the highest accuracies on the sounds, music, speech, and overall average categories, both on the Test-mini and Test-full splits. To understand the performance improvement, we tested models both with and without audio and found that much of the performance improvement from GRPO could be attributed to better text-based reasoning. We also made a surprising discovery that fine-tuning without audio on a text-only dataset was effective at improving the audio-based performance. Andrew Rouditchenko, Saurabhchand Bhati, Edson Araujo, Samuel Thomas 0001, Hilde Kuehne, Rogério Feris, James R. Glass |
ASRU | 6 |
| 2025 | CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained AlignmentabstractRecent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations that fail to capture fine-grained temporal correspondences with visual frames. Additionally, existing methods often struggle with conflicting optimization objectives when trying to jointly learn reconstruction and cross-modal alignment. In this work, we propose CAV-MAE Sync as a simple yet effective extension of the original CAV-MAE [14] framework for self-supervised audio-visual learning. We address three key challenges: First, we tackle the granularity mismatch between modalities by treating audio as a temporal sequence aligned with video frames, rather than using global representations. Second, we resolve conflicting optimization goals by separating contrastive and reconstruction objectives through dedicated global tokens. Third, we improve spatial localization by introducing learnable register tokens that reduce the semantic load on patch tokens. We evaluate the proposed approach on AudioSet, VGG Sound, and the ADE20K Sound dataset on zero-shot retrieval, classification, and localization tasks demonstrating state-of-the-art performance and outperforming more complex architectures. Code is available at https://github.com/edsonroteia/cav-mae-sync. Edson Araujo, Andrew Rouditchenko, Yuan Gong 0001, Saurabhchand Bhati, Samuel Thomas 0001, Brian Kingsbury, Leonid Karlinsky, Rogério Feris, James R. Glass, Hilde Kuehne |
CVPR | 8 |
| 2025 | Teaching VLMs to Localize Specific Objects from In-Context ExamplesabstractVision-Language Models (VLMs) have shown remarkable capabilities across diverse visual tasks, including image recognition, video understanding, and Visual Question Answering (VQA) when explicitly trained for these tasks. Despite these advances, we find that present-day VLMs (including the proprietary GPT-4o) lack a fundamental cognitive ability: learning to localize specific objects in a scene by taking into account the context. In this work, we focus on the task of few-shot personalized localization, where a model is given a small set of annotated images (in-context examples) -- each with a category label and bounding box -- and is tasked with localizing the same object type in a query image. Personalized localization can be particularly important in cases of ambiguity of several related objects that can respond to a text or an object that is hard to describe with words. To provoke personalized localization abilities in models, we present a data-centric solution that fine-tunes them using carefully curated data from video object tracking datasets. By leveraging sequences of frames tracking the same object across multiple shots, we simulate instruction-tuning dialogues that promote context awareness. To reinforce this, we introduce a novel regularization technique that replaces object labels with pseudo-names, ensuring the model relies on visual context rather than prior knowledge. Our method significantly enhances the few-shot localization performance of recent VLMs ranging from 7B to 72B in size, without sacrificing generalization, as demonstrated on several benchmarks tailored towards evaluating personalized localization abilities. This work is the first to explore and benchmark personalized few-shot localization for VLMs -- exposing critical weaknesses in present-day VLMs, and laying a foundation for future research in context-driven vision-language applications. Sivan Doveh, Nimrod Shabtay, Eli Schwartz, Hilde Kuehne, Raja Giryes, Rogério Feris, Leonid Karlinsky, James R. Glass, Assaf Arbelle, Shimon Ullman, Muhammad Jehanzeb Mirza |
ICCV | 6 |
| 2025 | BATCLIP: Bimodal Online Test-Time Adaptation for CLIP
Sarthak Kumar Maharana, Baoming Zhang, Leonid Karlinsky, Rogério Feris, Yunhui Guo |
ICCV | 4 |
| 2025 | Enhancing Few-Shot Vision-Language Classification With Large Multimodal Model Features
Chancharik Mitra, Brandon Huang, Tianning Chai, Zhiqiu Lin, Assaf Arbelle, Rogério Feris, Leonid Karlinsky, Trevor Darrell, Deva Ramanan, Roei Herzig |
ICCV | 6 |
| 2025 | Self-MoE: Towards Compositional Large Language Models with Self-Specialized ExpertsabstractWe present Self-MoE, an approach that transforms a monolithic LLM into a compositional, modular system of self-specialized experts, named MiXSE (MiXture of Self-specialized Experts). Our approach leverages self-specialization, which constructs expert modules using self-generated synthetic data, each equipping a shared base LLM with distinct domain-specific capabilities, activated via self-optimized routing. This allows for dynamic and capability-specific handling of various target tasks, enhancing overall capabilities, without extensive human-labeled data and added parameters. Our empirical results reveal that specializing LLMs may exhibit potential trade-offs in performances on non-specialized tasks. On the other hand, our Self-MoE demonstrates substantial improvements (6.5%p on average) over the base LLM across diverse benchmarks such as knowledge, reasoning, math, and coding. It also consistently outperforms other methods, including instance merging and weight merging, while offering better flexibility and interpretability by design with semantic experts and routing. Our findings highlight the critical role of modularity, the applicability of Self-MoE to multiple base LLMs, and the potential of self-improvement in achieving efficient, scalable, and adaptable systems. Junmo Kang, Leonid Karlinsky, Hongyin Luo, Zhen Wang 0041, Jacob A. Hansen, James R. Glass, David D. Cox, Rameswar Panda, Rogério Feris, Alan Ritter |
ICLR | 9 |
| 2025 | M+: Extending MemoryLLM with Scalable Long-Term MemoryabstractEquipping large language models (LLMs) with latent-space memory has attracted increasing attention as they can extend the context window of existing language models. However, retaining information from the distant past remains a challenge. For example, MemoryLLM (Wang et al., 2024a), as a representative work with latent-space memory, compresses past information into hidden states across all layers, forming a memory pool of 1B parameters. While effective for sequence lengths up to 16k tokens, it struggles to retain knowledge beyond 20k tokens. In this work, we address this limitation by introducing M+, a memory-augmented model based on MemoryLLM that significantly enhances long-term information retention. M+ integrates a long-term memory mechanism with a co-trained retriever, dynamically retrieving relevant information during text generation. We evaluate M+ on diverse benchmarks, including long-context understanding and knowledge retention tasks. Experimental results show that M+ significantly outperforms MemoryLLM and recent strong baselines, extending knowledge retention from under 20k to over 160k tokens with similar GPU memory overhead. Yu Wang 0170, Dmitry Krotov, Yifan Gao 0001, Wangchunshu Zhou, Julian J. McAuley, Dan Gutfreund, Rogério Feris, Zexue He |
ICML | 8 |
| 2025 | Special Section on SIBGRAPI 2023
Thales Sehn Körting, Esteban Walter Gonzalez Clua, Rogério Feris, Fernando Vieira Paulovich |
Pattern Recognit. Lett. | 3 |
| 2025 | mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech RecognitionabstractAudio-Visual Speech Recognition (AVSR) combines lip-based video with audio and can improve performance in noise, but most methods are trained only on English data. One limitation is the lack of large-scale multilingual video data, which makes it hard to train models from scratch. In this work, we propose mWhisper-Flamingo for multilingual AVSR which combines the strengths of a pre-trained audio model (Whisper) and video model (AV-HuBERT). To enable better multi-modal integration and improve the noisy multilingual performance, we introduce decoder modality dropout where the model is trained both on paired audio-visual inputs and separate audio/visual inputs. mWhisper-Flamingo achieves state-of-the-art WER on MuAViC, an AVSR dataset of 9 languages. Audio-visual mWhisper-Flamingo consistently outperforms audio-only Whisper on all languages in noisy conditions. Andrew Rouditchenko, Samuel Thomas 0001, Hilde Kuehne, Rogério Feris, James R. Glass |
IEEE Signal Process. Lett. | 4 |
| 2024 | What, When, and Where? Self-Supervised Spatio- Temporal Grounding in Untrimmed Multi-Action Videos from Narrated InstructionsabstractSpatio-temporal grounding describes the task of localizing events in space and time, e.g., in video data, based on verbal descriptions only. Models for this task are usually trained with human-annotated sentences and bounding box supervision. This work addresses this task from a multimodal supervision perspective, proposing a framework for spatio-temporal action grounding trained on loose video and subtitle supervision only, without human annotation. To this end, we combine local representation learning, which focuses on leveraging fine-grained spatial information, with a global representation encoding that captures higher-level representations and incorporates both in a joint approach. To evaluate this challenging task in a real-life setting, a new benchmark dataset is proposed, providing dense spatio-temporal grounding annotations in long, untrimmed, multi-action instructional videos for over 5K events. We evaluate the proposed approach and other methods on the proposed and standard downstream tasks, showing that our method improves over current baselines in various settings, including spatial, temporal, and untrimmed multi-action spatio-temporal grounding. Brian Chen 0001, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann, Samuel Thomas 0001, Shih-Fu Chang, Rogério Feris, James R. Glass, Hilde Kuehne |
CVPR | 7 |
| 2024 | Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
Andrew Rouditchenko, Yuan Gong 0001, Samuel Thomas 0001, Leonid Karlinsky, Hilde Kuehne, Rogério Feris, James R. Glass |
INTERSPEECH | 6 |
| 2024 | Large Scale Generative AI Text Applied to Sports and MusicabstractWe address the problem of scaling up the production of media content, including commentary and personalized news stories, for large-scale sports and music events worldwide. Our approach relies on generative AI models to transform a large volume of multimodal data (e.g., videos, articles, real-time scoring feeds, statistics, and fact sheets) into coherent and fluent text. Based on this approach, we introduce, for the first time, an AI commentary system, which was deployed to produce automated narrations for highlight packages at the 2023 US Open, Wimbledon, and Masters tournaments. In the same vein, our solution was extended to create personalized content for ESPN Fantasy Football and stories about music artists for the GRAMMY awards. These applications were built using a common software architecture achieved a 15x speed improvement with an average Rouge-L of 82.00 and perplexity of 6.6. Our work was successfully deployed at the aforementioned events, supporting 90 million fans around the world with 8 billion page views, continuously pushing the bounds on what is possible at the intersection of sports, entertainment, and AI. Aaron K. Baughman, Eduardo Morales, Rahul Agarwal, Gozde Akay, Rogério Feris, Tony Johnson, Stephen Hammer, Leonid Karlinsky |
KDD | 5 |
| 2024 | ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMsabstractCompositional Reasoning (CR) entails grasping the significance of attributes, relations, and word order. Recent Vision-Language Models (VLMs), comprising a visual encoder and a Large Language Model (LLM) decoder, have demonstrated remarkable proficiency in such reasoning tasks. This prompts a crucial question: have VLMs effectively tackled the CR challenge? We conjecture that existing CR benchmarks may not adequately push the boundaries of modern VLMs due to the reliance on an LLM only negative text generation pipeline. Consequently, the negatives produced either appear as outliers from the natural language distribution learned by VLMs' LLM decoders or as improbable within the corresponding image context. To address these limitations, we introduce ConMe\footnote{ConMe is an abbreviation for Confuse Me.} -- a compositional reasoning benchmark and a novel data generation pipeline leveraging VLMs to produce `hard CR Q&A'. Through a new concept of VLMs conversing with each other to collaboratively expose their weaknesses, our pipeline autonomously generates, evaluates, and selects challenging compositional reasoning questions, establishing a robust CR benchmark, also subsequently validated manually. Our benchmark provokes a noteworthy, up to 33%, decrease in CR performance compared to preceding benchmarks, reinstating the CR challenge even for state-of-the-art VLMs. Irene Huang, Wei Lin 0019, Muhammad Jehanzeb Mirza, Jacob A. Hansen, Sivan Doveh, Victor Ion Butoi, Roei Herzig, Assaf Arbelle, Hilde Kuehne, Trevor Darrell, Chuang Gan 0001, Aude Oliva, Rogério Feris, Leonid Karlinsky |
NeurIPS | 13 |
| 2024 | Trans-LoRA: towards data-free Transferable Parameter Efficient Finetuning
Runqian Wang, Soumya Ghosh, David D. Cox, Diego Antognini, Aude Oliva, Rogério Feris, Leonid Karlinsky |
NeurIPS | 6 |
| 2024 | DASS: Distilled Audio State Space Models are Stronger and More Duration-Scalable LearnersabstractState-space models (SSMs) have emerged as an alternative to Transformers for audio modeling due to their high computational efficiency with long inputs. While recent efforts on Audio SSMs have reported encouraging results, two main limitations remain: First, in 10 -second short audio tagging tasks, Audio SSMs still underperform compared to Transformer-based models such as Audio Spectrogram Transformer (AST). Second, although Audio SSMs theoretically support long audio inputs, their actual performance with long audio has not been thoroughly evaluated. To address these limitations, in this paper, 1) We applied knowledge distillation in audio space model training, resulting in a model called Knowledge Distilled Audio SSM (DASS). To the best of our knowledge, it is the first SSM that outperforms the Transformers on AudioSet and achieves an mAP of 47.6; and 2) We designed a new test called Audio Needle In A Haystack (Audio NIAH). We find that DASS, trained with only 10 -second audio clips, can retrieve sound events in audio recordings up to 2.5 hours long, while the AST model fails when the input is just 50 seconds, demonstrating state-space models are indeed more duration scalable. Saurabhchand Bhati, Yuan Gong 0001, Leonid Karlinsky, Hilde Kuehne, Rogério Feris, James R. Glass |
SLT | 5 |
| 2024 | Improved Techniques for Quantizing Deep Networks with Adaptive Bit-WidthsabstractQuantizing deep networks with adaptive bit-widths is a promising technique for efficient inference across many devices and resource constraints. In contrast to static methods that repeat the quantization process and train different models for different constraints, adaptive quantization enables us to flexibly adjust the bit-widths of a single deep network during inference for instant adaptation in different scenarios. While existing research shows encouraging results on common image classification benchmarks, this paper investigates how to train such adaptive networks more effectively. Specifically, we present two novel techniques for quantizing deep neural networks with adaptive bit-widths of weights and activations. First, we propose a collaborative strategy to choose a high-precision "teacher" for transferring knowledge to the low-precision "student" while jointly optimizing the model with all bit-widths. Second, to effectively transfer knowledge, we develop a dynamic block swapping method by randomly replacing the blocks in the lower-precision student network with the corresponding blocks in the higher-precision teacher network. Extensive experiments on multiple image and video classification datasets, well demonstrate the efficacy of our approach over state-of-the-art methods. Ximeng Sun, Rameswar Panda, Chun-Fu Chen 0001, Naigang Wang, Bowen Pan, Aude Oliva, Rogério Feris, Kate Saenko |
WACV | 7 |
| 2023 | Teaching Structured Vision & Language Concepts to Vision & Language ModelsabstractVision and Language ($VL$) models have demonstrated remarkable zero-shot performance in a variety of tasks. However, some aspects of complex language understanding still remain a challenge. We introduce the collective notion of Structured Vision & Language Concepts (SVLC) which includes object attributes, relations, and states which are present in the text and visible in the image. Recent studies have shown that even the best$VL$models struggle with SVLC. A possible way of fixing this issue is by collecting dedicated datasets for teaching each SVLC type, yet this might be expensive and time-consuming. Instead, we propose a more elegant data-driven approach for enhancing$VL$models' understanding of SVLCs that makes more effective use of existing$VL$pre-training datasets and does not require any additional data. While automatic understanding of image structure still remains largely unsolved, language structure is much better modeled and understood, allowing for its effective utilization in teaching$VL$models. In this paper, we propose various techniques based on language structure understanding that can be used to manipulate the textual part of off-the-shelf paired$VL$datasets.$VL$models trained with the updated data exhibit a significant improvement of up to 15% in their SVLC understanding with only a mild degradation in their zero-shot capabilities both when training from scratch or fine-tuning a pre-trained model. Our code and pretrained models are available at: https://github.com/SivanDoveh/TSVLC Sivan Doveh, Assaf Arbelle, Sivan Harary, Eli Schwartz, Roei Herzig, Raja Giryes, Rogério Feris, Rameswar Panda, Shimon Ullman, Leonid Karlinsky |
CVPR | 7 |
| 2023 | ConStruct-VL: Data-Free Continual Structured VL Concepts LearningabstractRecently, large-scale pre-trained Vision-and-Language (VL) foundation models have demonstrated remarkable capabilities in many zero-shot downstream tasks, achieving competitive results for recognizing objects defined by as little as short text prompts. However, it has also been shown that VL models are still brittle in Structured VL Concept (SVLC) reasoning, such as the ability to recognize object attributes, states, and inter-object relations. This leads to reasoning mistakes, which need to be corrected as they occur by teaching VL models the missing SVLC skills; often this must be done using private data where the issue was found, which naturally leads to a data-free continual (no task-id) VL learning setting. In this work, we introduce the first Continual Data-Free Structured VL Concepts Learning (ConStruct-VL) benchmark11Our code is publicly available at https://github.com/jamessealesmith/ConStruct-VL and show it is challenging for many existing data-free CL strategies. We, therefore, propose a data-free method comprised of a new approach of Adversarial Pseudo-Replay (APR) which generates adversarial reminders of past tasks from past task models. To use this method efficiently, we also propose a continual parameter-efficient Layered-LoRA (LaLo) neural architecture allowing no-memory-cost access to all past models at train time. We show this approach outperforms all data-free methods by as much as ~ 7% while even matching some levels of experience-replay (prohibitive for applications where data-privacy must be preserved). James Seale Smith, Paola Cascante-Bonilla, Assaf Arbelle, Donghyun Kim 0006, Rameswar Panda, David D. Cox, Diyi Yang, Zsolt Kira, Rogério Feris, Leonid Karlinsky |
CVPR | 9 |
| 2023 | CODA-Prompt: COntinual Decomposed Attention-Based Prompting for Rehearsal-Free Continual LearningabstractComputer vision models suffer from a phenomenon known as catastrophic forgetting when learning novel concepts from continuously shifting training data. Typical solutions for this continual learning problem require extensive rehearsal of previously seen data, which increases memory costs and may violate data privacy. Recently, the emergence of large-scale pre-trained vision transformer models has enabled prompting approaches as an alternative to data-rehearsal. These approaches rely on a key-query mechanism to generate prompts and have been found to be highly resistant to catastrophic forgetting in the well-established rehearsal-free continual learning setting. However, the key mechanism of these methods is not trained end-to-end with the task sequence. Our experiments show that this leads to a reduction in their plasticity, hence sacrificing new task accuracy, and inability to benefit from expanded parameter capacity. We instead propose to learn a set of prompt components which are assembled with input-conditioned weights to produce input-conditioned prompts, resulting in a novel attention-based end-to-end key-query scheme. Our experiments show that we outperform the current SOTA method DualPrompt on established benchmarks by as much as 4.5% in average final accuracy. We also outperform the state of art by as much as 4.4% accuracy on a continual learning benchmark which contains both class-incremental and domain-incremental task shifts, corresponding to many practical settings. Our code is available at https://github.com/GT-RIPL/CODA-Prompt James Seale Smith, Leonid Karlinsky, Vyshnavi Gutta, Paola Cascante-Bonilla, Donghyun Kim 0006, Assaf Arbelle, Rameswar Panda, Rogério Feris, Zsolt Kira |
CVPR | 8 |
| 2023 | Incorporating Structured Representations into Pretrained Vision & Language Models Using Scene GraphsabstractRoei Herzig, Alon Mendelson, Leonid Karlinsky, Assaf Arbelle, Rogerio Feris, Trevor Darrell, Amir Globerson. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Roei Herzig, Alon Mendelson, Leonid Karlinsky, Assaf Arbelle, Rogério Feris, Trevor Darrell, Amir Globerson |
EMNLP | 5 |
| 2023 | C2KD: Cross-Lingual Cross-Modal Knowledge Distillation for Multilingual Text-Video RetrievalabstractMultilingual text-video retrieval methods have improved significantly in recent years, but the performance for languages other than English still lags. We propose a Cross-Lingual Cross-Modal Knowledge Distillation method to improve multilingual text-video retrieval. Inspired by the fact that English text-video retrieval outperforms other languages, we train a student model using input text in different languages to match the cross-modal predictions from teacher models using input text in English. We propose a cross entropy based objective which forces the distribution over the student’s text-video similarity scores to be similar to those of the teacher models. We introduce a new multilingual video dataset, Multi-YouCook2, by translating the English captions in the YouCook2 video dataset to 8 other languages. Our method improves multilingual text-video retrieval performance on Multi-YouCook2 and several other datasets such as Multi-MSRVTT and VATEX. We also conducted an analysis on the effectiveness of different multilingual text models as teachers. Andrew Rouditchenko, Yung-Sung Chuang, Nina Shvetsova, Samuel Thomas 0001, Rogério Feris, Brian Kingsbury, Leonid Karlinsky, David F. Harwath, Hilde Kuehne, James R. Glass |
ICASSP | 5 |
| 2023 | MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language KnowledgeabstractLarge scale Vision Language (VL) models have shown tremendous success in aligning representations between visual and text modalities. This enables remarkable progress in zero-shot recognition, image generation & editing, and many other exciting tasks. However, VL models tend to over-represent objects while paying much less attention to verbs, and require additional tuning on video data for best zero-shot action recognition performance. While previous work relied on large-scale, fully-annotated data, in this work we propose an unsupervised approach. We adapt a VL model for zero-shot and few-shot action recognition using a collection of unlabeled videos and an unpaired action dictionary. Based on that, we leverage Large Language Models and VL models to build a text bag for each unlabeled video via matching, text expansion and captioning. We use those bags in a Multiple Instance Learning setup to adapt an image-text backbone to video data. Although finetuned on unlabeled video data, our resulting models demonstrate high transferability to numerous unseen zero-shot downstream tasks, improving the base VL model performance by up to 14%, and even comparing favorably to fully-supervised baselines in both zero-shot and few-shot video recognition transfer. The code is released at https://github.com/wlin-at/MAXI. Wei Lin 0019, Leonid Karlinsky, Nina Shvetsova, Horst Possegger, Mateusz Kozinski, Rameswar Panda, Rogério Feris, Hilde Kuehne, Horst Bischof |
ICCV | 7 |
| 2023 | Going Beyond Nouns With Vision & Language Models Using Synthetic DataabstractLarge-scale pre-trained Vision & Language (VL) models have shown remarkable performance in many applications, enabling replacing a fixed set of supported classes with zero-shot open vocabulary reasoning over (almost arbitrary) natural language prompts. However, recent works have uncovered a fundamental weakness of these models. For example, their difficulty to understand Visual Language Concepts (VLC) that go ‘beyond nouns’ such as the meaning of non-object words (e.g., attributes, actions, relations, states, etc.), or difficulty in performing compositional reasoning such as understanding the significance of the order of the words in a sentence. In this work, we investigate to which extent purely synthetic data could be leveraged to teach these models to overcome such shortcomings without compromising their zero-shot capabilities. We contribute Synthetic Visual Concepts (SyViC) - a million-scale synthetic dataset and data generation codebase allowing to generate additional suitable data to improve VLC understanding and compositional reasoning of VL models. Additionally, we propose a general VL finetuning strategy for effectively leveraging SyViC towards achieving these improvements. Our extensive experiments and ablations on VL-Checklist, Winoground, and ARO benchmarks demonstrate that it is possible to adapt strong pre-trained VL models with synthetic data significantly enhancing their VLC understanding (e.g. by 9.9% on ARO and 4.3% on VL-Checklist) with under 1% drop in their zero-shot accuracy. Paola Cascante-Bonilla, Khaled Shehada, James Seale Smith, Sivan Doveh, Donghyun Kim 0006, Rameswar Panda, Gül Varol, Aude Oliva, Vicente Ordonez, Rogério Feris, Leonid Karlinsky |
ICCV | 10 |
| 2023 | CDAC: Cross-domain Attention Consistency in Transformer for Domain Adaptive Semantic SegmentationabstractWhile transformers have greatly boosted performance in semantic segmentation, domain adaptive transformers are not yet well explored. We identify that the domain gap can cause discrepancies in self-attention. Due to this gap, the transformer attends to spurious regions or pixels, which deteriorates accuracy on the target domain. We propose Cross-Domain Attention Consistency (CDAC), to perform adaptation on attention maps using cross-domain attention layers that share features between source and target domains. Specifically, we impose consistency between predictions from cross-domain attention and self-attention modules to encourage similar distributions across domains in both the attention and output of the model, i.e., attention-level and output-level alignment. We also enforce consistency in attention maps between different augmented views to further strengthen the attention-based alignment. Combining these two components, CDAC mitigates the discrepancy in attention maps across domains and further boosts the performance of the transformer under unsupervised domain adaptation settings. Our method is evaluated on various widely used benchmarks and outperforms the state-of-the-art baselines, including GTAV-to-Cityscapes by 1.3 and 1.5 percent point (pp) and Synthia-to-Cityscapes by 0.6 pp and 2.9 pp when combining with two competitive Transformer-based backbones, respectively. Our code will be publicly available at https://github.com/wangkaihong/CDAC. Kaihong Wang, Donghyun Kim 0006, Rogério Feris, Margrit Betke |
ICCV | 3 |
| 2023 | Learning to Grow Pretrained Models for Efficient Transformer Training
Peihao Wang, Rameswar Panda, Lucas Torroba Hennigen, Philip Greengard, Leonid Karlinsky, Rogério Feris, David D. Cox, Zhangyang Wang |
ICLR | 6 |
| 2023 | Multitask Prompt Tuning Enables Parameter-Efficient Transfer Learning
Zhen Wang 0041, Rameswar Panda, Leonid Karlinsky, Rogério Feris, Huan Sun 0001 |
ICLR | 4 |
| 2023 | Comparison of Multilingual Self-Supervised and Weakly-Supervised Speech Pre-Training for Adaptation to Unseen Languages
Andrew Rouditchenko, Sameer Khurana, Samuel Thomas 0001, Rogério Feris, Leonid Karlinsky, Hilde Kuehne, David F. Harwath, Brian Kingsbury, James R. Glass |
INTERSPEECH | 4 |
| 2023 | Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL ModelsabstractVision and Language (VL) models offer an effective method for aligning representation spaces of images and text allowing for numerous applications such as cross-modal retrieval, visual and multi-hop question answering, captioning, and many more. However, the aligned image-text spaces learned by all the popular VL models are still suffering from the so-called 'object bias' - their representations behave as 'bags of nouns' mostly ignoring or downsizing the attributes, relations, and states of objects described/appearing in texts/images. Although some great attempts at fixing these `compositional reasoning' issues were proposed in the recent literature, the problem is still far from being solved. In this paper, we uncover two factors limiting the VL models' compositional reasoning performance. These two factors are properties of the paired VL dataset used for finetuning (or pre-training) the VL model: (i) the caption quality, or in other words 'image-alignment', of the texts; and (ii) the 'density' of the captions in the sense of mentioning all the details appearing on the image. We propose a fine-tuning approach for automatically treating these factors on a standard collection of paired VL data (CC3M). Applied to CLIP, we demonstrate its significant compositional reasoning performance increase of up to $\sim27$\% over the base model, up to $\sim20$\% over the strongest baseline, and by $6.7$\% on average. Our code is provided in the Supplementary and would be released upon acceptance. Sivan Doveh, Assaf Arbelle, Sivan Harary, Roei Herzig, Donghyun Kim 0006, Paola Cascante-Bonilla, Amit Alfassy, Rameswar Panda, Raja Giryes, Rogério Feris, Shimon Ullman, Leonid Karlinsky |
NeurIPS | 10 |
| 2023 | LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image CollectionsabstractRecently, large-scale pre-trained Vision and Language (VL) models have set a new state-of-the-art (SOTA) in zero-shot visual classification enabling open-vocabulary recognition of potentially unlimited set of categories defined as simple language prompts. However, despite these great advances, the performance of these zero-shot classifiers still falls short of the results of dedicated (closed category set) classifiers trained with supervised fine-tuning. In this paper we show, for the first time, how to reduce this gap without any labels and without any paired VL data, using an unlabeled image collection and a set of texts auto-generated using a Large Language Model (LLM) describing the categories of interest and effectively substituting labeled visual instances of those categories. Using our label-free approach, we are able to attain significant performance improvements over the zero-shot performance of the base VL model and other contemporary methods and baselines on a wide variety of datasets, demonstrating absolute improvement of up to $11.7\%$ ($3.8\%$ on average) in the label-free setting. Moreover, despite our approach being label-free, we observe $1.3\%$ average gains over leading few-shot prompting baselines that do use 5-shot supervision. Muhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin 0019, Horst Possegger, Mateusz Kozinski, Rogério Feris, Horst Bischof |
NeurIPS | 6 |
| 2023 | Learning Human Action Recognition Representations Without Real HumansabstractPre-training on massive video datasets has become essential to achieve high action recognition performance on smaller downstream datasets. However, most large-scale video datasets contain images of people and hence are accompanied with issues related to privacy, ethics, and data protection, often preventing them from being publicly shared for reproducible research. Existing work has attempted to alleviate these problems by blurring faces, downsampling videos, or training on synthetic data. On the other hand, analysis on the {\em transferability} of privacy-preserving pre-trained models to downstream tasks has been limited. In this work, we study this problem by first asking the question: can we pre-train models for human action recognition with data that does not include real humans? To this end, we present, for the first time, a benchmark that leverages real-world videos with {\em humans removed} and synthetic data containing virtual humans to pre-train a model. We then evaluate the transferability of the representation learned on this data to a diverse set of downstream action recognition benchmarks. Furthermore, we propose a novel pre-training strategy, called Privacy-Preserving MAE-Align, to effectively combine synthetic data and human-removed real data. Our approach outperforms previous baselines by up to 5\% and closes the performance gap between human and no-human action recognition representations on downstream tasks, for both linear probing and fine-tuning. Our benchmark, code, and models are available at https://github.com/howardzh01/PPMA. Howard Zhong, Samarth Mishra, Donghyun Kim 0006, SouYoung Jin, Rameswar Panda, Hilde Kuehne, Leonid Karlinsky, Venkatesh Saligrama, Aude Oliva, Rogério Feris |
NeurIPS | 10 |
| 2023 | Addressing Feature Suppression in Unsupervised Visual RepresentationsabstractContrastive learning is one of the fastest growing research areas in machine learning due to its ability to learn useful representations without labeled data. However, contrastive learning is susceptible to feature suppression – i.e., it may discard important information relevant to the task of interest, and learn irrelevant features. Past work has addressed this limitation via handcrafted data augmentations that eliminate irrelevant information. This approach however does not work across all datasets and tasks. Further, data augmentations fail in addressing feature suppression in multi-attribute classification when one attribute can suppress features relevant to other attributes. In this paper, we analyze the objective function of contrastive learning and formally prove that it is vulnerable to feature suppression. We then present Predictive Contrastive Learning (PrCL), a framework for learning unsupervised representations that are robust to feature suppression. The key idea is to force the learned representation to predict the input, and hence prevent it from discarding important information. Extensive experiments verify that PrCL is robust to feature suppression and outperforms state-of-the-art contrastive learning methods on a variety of datasets and tasks. Tianhong Li, Lijie Fan, Yuan Yuan 0002, Hao He 0011, Yonglong Tian, Rogério Feris, Piotr Indyk, Dina Katabi |
WACV | 6 |
| 2023 | Select, Label, and Mix: Learning Discriminative Invariant Feature Representations for Partial Domain AdaptationabstractPartial domain adaptation which assumes that the unknown target label space is a subset of the source label space has attracted much attention in computer vision. Despite recent progress, existing methods often suffer from three key problems: negative transfer, lack of discriminability, and domain invariance in the latent space. To alleviate the above issues, we develop a novel ‘Select, Label, and Mix’ (SLM) framework that aims to learn discriminative invariant feature representations for partial domain adaptation. First, we present an efficient "select" module that automatically filters out the outlier source samples to avoid negative transfer while aligning distributions across both domains. Second, the "label" module iteratively trains the classifier using both the labeled source domain data and the generated pseudo-labels for the target domain to enhance the discriminability of the latent space. Finally, the "mix" module utilizes domain mixup regularization jointly with the other two modules to explore more intrinsic structures across domains leading to a domain-invariant latent space for partial domain adaptation. Extensive experiments on several benchmark datasets for partial domain adaptation demonstrate the superiority of our proposed framework over state-of-the-art methods. Project page: https://cvir.github.io/projects/slm. Aadarsh Sahoo, Rameswar Panda, Rogério Feris, Kate Saenko, Abir Das |
WACV | 3 |
| 2022 | Sim VQA: Exploring Simulated Environments for Visual Question AnsweringabstractExisting work on VQA explores data augmentation to achieve better generalization by perturbing images in the dataset or modifying existing questions and answers. While these methods exhibit good performance, the diversity of the questions and answers are constrained by the available images. In this work we explore using synthetic computer-generated data to fully control the visual and language space, allowing us to provide more diverse scenarios. We quantify the effectiveness of leveraging synthetic data for real-world VQA. By exploiting 3D and physics simulation platforms, we provide a pipeline to generate synthetic data to expand and replace type-specific questions and answers without risking exposure of sensitive or personal data that might be present in real images. We offer a comprehensive analysis while expanding existing hyper-realistic datasets to be usedfor VQA. We also propose Feature Swapping (F-SWAP) - where we randomly switch object-level features during training to make a VQA model more domain invariant. We show that F-SWAP is effective for improving VQA models on real images without compromising on their accuracy to answer existing questions in the dataset. Paola Cascante-Bonilla, Hui Wu 0009, Letao Wang, Rogério Feris, Vicente Ordonez |
CVPR | 4 |
| 2022 | Unsupervised Domain Generalization by Learning a Bridge Across DomainsabstractThe ability to generalize learned representations across significantly different visual domains, such as between real photos, clipart, paintings, and sketches, is a fundamental capacity of the human visual system. In this paper, different from most cross-domain works that utilize some (or full) source domain supervision, we approach a relatively new and very practical Unsupervised Domain Generalization (UDG) setup of having no training supervision in neither source nor target domains. Our approach is based on self-supervised learning of a Bridge Across Domains (BrAD) - an auxiliary bridge domain accompanied by a set of semantics preserving visual (image-to-image) mappings to BrAD from each of the training domains. The BrAD and mappings to it are learned jointly (end-to-end) with a contrastive self-supervised representation model that semantically aligns each of the domains to its BrAD-projection, and hence implicitly drives all the domains (seen or unseen) to semantically align to each other. In this work, we show how using an edge-regularized BrAD our approach achieves significant gains across multiple benchmarks and a range of tasks, including UDG, Few-shot UDA, and unsupervised generalization across multi-domain datasets (including generalization to unseen domains and classes). Sivan Harary, Eli Schwartz, Assaf Arbelle, Peter W. J. Staar, Shady Abu-Hussein, Elad Amrani, Roei Herzig, Amit Alfassy, Raja Giryes, Hilde Kuehne, Dina Katabi, Kate Saenko, Rogério Feris, Leonid Karlinsky |
CVPR | 13 |
| 2022 | Targeted Supervised Contrastive Learning for Long-Tailed RecognitionabstractReal-world data often exhibits long tail distributions with heavy class imbalance, where the majority classes can dominate the training process and alter the decision bound-aries of the minority classes. Recently, researchers have in-vestigated the potential of supervised contrastive learning for long-tailed recognition, and demonstrated that it provides a strong performance gain. In this paper, we show that while supervised contrastive learning can help improve performance, past baselines suffer from poor uniformity brought in by imbalanced data distribution. This poor uni-formity manifests in samples from the minority class having poor separability in the feature space. To address this problem, we propose targeted supervised contrastive learning (TSC), which improves the uniformity of the feature distribution on the hypersphere. TSC first generates a set of targets uniformly distributed on a hypersphere. It then makes the features of different classes converge to these distinct and uniformly distributed targets during training. This forces all classes, including minority classes, to main-tain a uniform distribution in the feature space, improves class boundaries, and provides better generalization even in the presence of long-tail data. Experiments on multi-ple datasets show that TSC achieves state-of-the-art performance on long-tailed recognition tasks. Tianhong Li, Yuan Yuan 0002, Lijie Fan, Yuzhe Yang 0003, Rogério Feris, Piotr Indyk, Dina Katabi |
CVPR | 6 |
| 2022 | VALHALLA: Visual Hallucination for Machine TranslationabstractDesigning better machine translation systems by considering auxiliary inputs such as images has attracted much attention in recent years. While existing methods show promising performance over the conventional text-only translation systems, they typically require paired text and image as input during inference, which limits their applicability to real-world scenarios. In this paper, we introduce a visual hallucination framework, called VALHALLA, which requires only source sentences at inference time and instead uses hallucinated visual representations for multi-modal machine translation. In particular, given a source sentence an autoregressive hallucination transformer is used to predict a discrete visual representation from the input text, and the combined text and hallucinated representations are utilized to obtain the target translation. We train the hallucination transformer jointly with the translation transformer using standard backpropagation with crossentropy losses while being guided by an additional loss that encourages consistency between predictions using either groundtruth or hallucinated visual representations. Extensive experiments on three standard translation datasets with a diverse set of language pairs demonstrate the effectiveness of our approach over both text-only baselines and state-of-the-art methods. Project page: http://www.svcl.ucsd.jects/valhalla.edu/pro. Yi Li 0051, Rameswar Panda, Chun-Fu Chen 0001, Rogério Feris, David D. Cox, Nuno Vasconcelos |
CVPR | 5 |
| 2022 | Task2Sim: Towards Effective Pre-training and Transfer from Synthetic DataabstractPre-training models on Imagenet or other massive datasets of real images has led to major advances in Computer vision, albeit accompanied with shortcomings related to curation cost, privacy, usage rights, and ethical issues. In this paper, for the first time, we study the transferability of pre-trained models based on synthetic data generated by graphics simulators to downstream tasks from very different domains. In using such synthetic data for pre-training, we find that downstream performance on different tasks are fa-vored by different configurations of simulation parameters (e.g. lighting, object pose, backgrounds, etc.), and that there is no one-size-fits-all solution. It is thus better to tailor syn-thetic pre-training data to a specific downstream task, for best performance. We introduce Task2Sim, a unified model mapping downstream task representations to optimal sim-ulation parameters to generate synthetic pre-training data for them. Task2Sim learns this mapping by training to find the set of best parameters on a set of “seen” tasks. Once trained, it can then be used to predict best simulation pa-rameters for novel “unseen” tasks in one shot, without re-quiring additional training. Given a budget in number of images per class, our extensive experiments with 20 di-verse downstream tasks show Task2Sim's task-adaptive pre-training data results in significantly better downstream per-formance than non-adaptively choosing simulation param-eters on both seen and unseen tasks. It is even competitive with pre-training on real images from Imagenet. Samarth Mishra, Rameswar Panda, Cheng Perng Phoo, Chun-Fu Chen 0001, Leonid Karlinsky, Kate Saenko, Venkatesh Saligrama, Rogério Feris |
CVPR | 8 |
| 2022 | Everything at Once - Multi-modal Fusion Transformer for Video RetrievalabstractMulti-modal learning from video data has seen increased attention recently as it allows training of semantically meaningful embeddings without human annotation, enabling tasks like zero-shot retrieval and action localization. In this work, we present a multi-modal, modality agnostic fusion transformer that learns to exchange information between multiple modalities, such as video, audio, and text, and integrate them into a fused representation in a joined multi-modal embedding space. We propose to train the system with a combinatorial loss on everything at once – any combination of input modalities, such as single modalities as well as pairs of modalities, explicitly leaving out any add-ons such as position or modality encoding. At test time, the resulting model can process and fuse any number of input modalities. Moreover, the implicit properties of the transformer allow to process inputs of different lengths. To evaluate the proposed approach, we train the model on the large scale HowTo100M dataset and evaluate the resulting embedding space on four challenging benchmark datasets obtaining state-of-the-art results in zero-shot video retrieval and zero-shot video action localization. Our code for this work is also available.11https://github.com/ninatu/everything_at_once Nina Shvetsova, Brian Chen 0001, Andrew Rouditchenko, Samuel Thomas 0001, Brian Kingsbury, Rogério Feris, David F. Harwath, James R. Glass, Hilde Kuehne |
CVPR | 6 |
| 2022 | A Maximal Correlation Approach to Imposing Fairness in Machine LearningabstractAs machine learning algorithms grow in popularity and diversify to many industries, ethical and legal concerns regarding their fairness have become increasingly relevant. We explore the problem of algorithmic fairness, taking an information-theoretic view. The maximal correlation framework is introduced for expressing fairness constraints and shown to be capable of deriving regularizers that enforce independence and separation-based fairness criteria, which admit optimization algorithms that are more computationally efficient than existing algorithms. We show that these algorithms provide smooth performance-fairness tradeoff curves and perform competitively with state-of-the-art methods on the Communities and Crimes dataset. Joshua K. Lee, Yuheng Bu, Prasanna Sattigeri, Rameswar Panda, Gregory W. Wornell, Leonid Karlinsky, Rogério Feris |
ICASSP | 7 |
| 2022 | FETA: Towards Specializing Foundational Models for Expert Task ApplicationsabstractFoundational Models (FMs) have demonstrated unprecedented capabilities including zero-shot learning, high fidelity data synthesis, and out of domain generalization. However, the parameter capacity of FMs is still limited, leading to poor out-of-the-box performance of FMs on many expert tasks (e.g. retrieval of car manuals technical illustrations from language queries), data for which is either unseen or belonging to a long-tail part of the data distribution of the huge datasets used for FM pre-training. This underlines the necessity to explicitly evaluate and finetune FMs on such expert tasks, arguably ones that appear the most in practical real-world applications. In this paper, we propose a first of its kind FETA benchmark built around the task of teaching FMs to understand technical documentation, via learning to match their graphical illustrations to corresponding language descriptions. Our FETA benchmark focuses on text-to-image and image-to-text retrieval in public car manuals and sales catalogue brochures. FETA is equipped with a procedure for completely automatic annotation extraction (code would be released upon acceptance), allowing easy extension of FETA to more documentation types and application domains in the future. Our automatic annotation leads to an automated performance metric shown to be consistent with metrics computed on human-curated annotations (also released). We provide multiple baselines and analysis of popular FMs on FETA leading to several interesting findings that we believe would be very valuable to the FM community, paving the way towards real-world application of FMs for many practical expert tasks currently being `overlooked' by standard benchmarks focusing on common objects. Amit Alfassy, Assaf Arbelle, Oshri Halimi, Sivan Harary, Roei Herzig, Eli Schwartz, Rameswar Panda, Michele Dolfi, Christoph Auer, Peter W. J. Staar, Kate Saenko, Rogério Feris, Leonid Karlinsky |
NeurIPS | 12 |
| 2022 | Procedural Image Programs for Representation LearningabstractLearning image representations using synthetic data allows training neural networks without some of the concerns associated with real images, such as privacy and bias. Existing work focuses on a handful of curated generative processes which require expert knowledge to design, making it hard to scale up. To overcome this, we propose training with a large dataset of twenty-one thousand programs, each one generating a diverse set of synthetic images. These programs are short code snippets, which are easy to modify and fast to execute using OpenGL. The proposed dataset can be used for both supervised and unsupervised representation learning, and reduces the gap between pre-training with real and procedurally generated images by 38%. Manel Baradad Jurjo, Chun-Fu Chen 0001, Jonas Wulff, Tongzhou Wang 0001, Rogério Feris, Antonio Torralba 0001, Phillip Isola |
NeurIPS | 5 |
| 2022 | How Transferable are Video Representations Based on Synthetic Data?abstractAction recognition has improved dramatically with massive-scale video datasets. Yet, these datasets are accompanied with issues related to curation cost, privacy, ethics, bias, and copyright. Compared to that, only minor efforts have been devoted toward exploring the potential of synthetic video data. In this work, as a stepping stone towards addressing these shortcomings, we study the transferability of video representations learned solely from synthetically-generated video clips, instead of real data. We propose SynAPT, a novel benchmark for action recognition based on a combination of existing synthetic datasets, in which a model is pre-trained on synthetic videos rendered by various graphics simulators, and then transferred to a set of downstream action recognition datasets, containing different categories than the synthetic data. We provide an extensive baseline analysis on SynAPT revealing that the simulation-to-real gap is minor for datasets with low object and scene bias, where models pre-trained with synthetic data even outperform their real data counterparts. We posit that the gap between real and synthetic action representations can be attributed to contextual bias and static objects related to the action, instead of the temporal dynamics of the action itself. The SynAPT benchmark is available at https://github.com/mintjohnkim/SynAPT. Yo-whan Kim, Samarth Mishra, SouYoung Jin, Rameswar Panda, Hilde Kuehne, Leonid Karlinsky, Venkatesh Saligrama, Kate Saenko, Aude Oliva, Rogério Feris |
NeurIPS | 10 |
| 2022 | Multi-Moments in Time: Learning and Interpreting Models for Multi-Action Video UnderstandingabstractVideos capture events that typically contain multiple sequential, and simultaneous, actions even in the span of only a few seconds. However, most large-scale datasets built to train models for action recognition in video only provide a single label per video. Consequently, models can be incorrectly penalized for classifying actions that exist in the videos but are not explicitly labeled and do not learn the full spectrum of information present in each video in training. Towards this goal, we present the Multi-Moments in Time dataset (M-MiT) which includes over two million action labels for over one million three second videos. This multi-label dataset introduces novel challenges on how to train and analyze models for multi-action detection. Here, we present baseline results for multi-action recognition using loss functions adapted for long tail multi-label learning, provide improved methods for visualizing and interpreting models trained for multi-label action detection and show the strength of transferring models trained on M-MiT to smaller datasets. Mathew Monfort, Bowen Pan, Kandan Ramakrishnan, Alex Andonian, Barry A. McNamara, Alex Lascelles, Quanfu Fan, Dan Gutfreund, Rogério Feris, Aude Oliva |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2022 | Baby steps towards few-shot learning with multiple semantics
Eli Schwartz, Leonid Karlinsky, Rogério Feris, Raja Giryes, Alexander M. Bronstein |
Pattern Recognit. Lett. | 3 |
| 2021 | StarNet: towards Weakly Supervised Few-Shot Object DetectionabstractFew-shot detection and classification have advanced significantly in recent years. Yet, detection approaches require strong annotation (bounding boxes) both for pre-training and for adaptation to novel classes, and classification approaches rarely provide localization of objects in the scene. In this paper, we introduce StarNet - a few-shot model featuring an end-to-end differentiable non-parametric star-model detection and classification head. Through this head, the backbone is meta-trained using only image-level labels to produce good features for jointly localizing and classifying previously unseen categories of few-shot test tasks using a star-model that geometrically matches between the query and support images (to find corresponding object instances). Being a few-shot detector, StarNet does not require any bounding box annotations, neither during pre-training nor for novel classes adaptation. It can thus be applied to the previously unexplored and challenging task of Weakly Supervised Few-Shot Object Detection (WS-FSOD), where it attains significant improvements over the baselines. In addition, StarNet shows significant gains on few-shot classification benchmarks that are less cropped around the objects (where object localization is key). Leonid Karlinsky, Joseph Shtok, Amit Alfassy, Moshe Lichtenstein, Sivan Harary, Eli Schwartz, Sivan Doveh, Prasanna Sattigeri, Rogério Feris, Alexander M. Bronstein, Raja Giryes |
AAAI | 9 |
| 2021 | NASTransfer: Analyzing Architecture Transferability in Large Scale Neural Architecture SearchabstractNeural Architecture Search (NAS) is an open and challenging problem in machine learning. While NAS offers great promise, the prohibitive computational demand of most of the existing NAS methods makes it difficult to directly search the architectures on large-scale tasks. The typical way of conducting large scale NAS is to search for an architectural building block on a small dataset (either using a proxy set from the large dataset or a completely different small scale dataset) and then transfer the block to a larger dataset. Despite a number of recent results that show the promise of transfer from proxy datasets, a comprehensive evaluation of different NAS methods studying the impact of different source datasets has not yet been addressed. In this work, we propose to analyze the architecture transferability of different NAS methods by performing a series of experiments on large scale benchmarks such as ImageNet1K and ImageNet22K. We find that: (i) The size and domain of the proxy set does not seem to influence architecture performance on the target dataset. On average, transfer performance of architectures searched using completely different small datasets (e.g., CIFAR10) perform similarly to the architectures searched directly on proxy target datasets. However, design of proxy sets has considerable impact on rankings of different NAS methods. (ii) While different NAS methods show similar performance on a source dataset (e.g., CIFAR10), they significantly differ on the transfer performance to a large dataset (e.g., ImageNet1K). (iii) Even on large datasets, random sampling baseline is very competitive, but the choice of the appropriate combination of proxy set and search strategy can provide significant improvement over it. We believe that our extensive empirical analysis will prove useful for future design of NAS algorithms. Rameswar Panda, Michele Merler, Mayoore S. Jaiswal, Hui Wu 0009, Kandan Ramakrishnan, Ulrich Finkler, Chun-Fu Chen 0001, Minsik Cho, Rogério Feris, David S. Kung 0001, Bishwaranjan Bhattacharjee |
AAAI | 9 |
| 2021 | Fine-Grained Angular Contrastive Learning With Coarse LabelsabstractFew-shot learning methods offer pre-training techniques optimized for easier later adaptation of the model to new classes (unseen during training) using one or a few examples. This adaptivity to unseen classes is especially important for many practical applications where the pre-trained label space cannot remain fixed for effective use and the model needs to be "specialized" to support new categories on the fly. One particularly interesting scenario, essentially overlooked by the few-shot literature, is Coarse-to-Fine Few-Shot (C2FS), where the training classes (e.g. animals) are of much ‘coarser granularity’ than the target (test) classes (e.g. breeds). A very practical example of C2FS is when the target classes are sub-classes of the training classes. Intuitively, it is especially challenging as (both regular and few-shot) supervised pre-training tends to learn to ignore intra-class variability which is essential for separating sub-classes. In this paper, we introduce a novel ’Angular normalization’ module that allows to effectively combine supervised and self-supervised contrastive pre-training to approach the proposed C2FS task, demonstrating significant gains in a broad study over multiple baselines and datasets. We hope that this work will help to pave the way for future research on this new, challenging, and very practical topic of C2FS classification. Guy Bukchin, Eli Schwartz, Kate Saenko, Ori Shahar, Rogério Feris, Raja Giryes, Leonid Karlinsky |
CVPR | 5 |
| 2021 | Deep Analysis of CNN-Based Spatio-Temporal Representations for Action RecognitionabstractIn recent years, a number of approaches based on 2D or 3D convolutional neural networks (CNN) have emerged for video action recognition, achieving state-of-the-art results on several large-scale benchmark datasets. In this paper, we carry out in-depth comparative analysis to better understand the differences between these approaches and the progress made by them. To this end, we develop an unified framework for both 2D-CNN and 3D-CNN action models, which enables us to remove bells and whistles and provides a common ground for fair comparison. We then conduct an effort towards a large-scale analysis involving over 300 action recognition models. Our comprehensive analysis reveals that a) a significant leap is made in efficiency for action recognition, but not in accuracy; b) 2D-CNN and 3D-CNN models behave similarly in terms of spatio-temporal representation abilities and transferability. Our codes are available at https://github.com/IBM/action-recognition-pytorch. Chun-Fu Chen 0001, Rameswar Panda, Kandan Ramakrishnan, Rogério Feris, John Cohn, Aude Oliva, Quanfu Fan |
CVPR | 4 |
| 2021 | Spoken Moments: Learning Joint Audio-Visual Representations From Video DescriptionsabstractWhen people observe events, they are able to abstract key information and build concise summaries of what is happening. These summaries include contextual and semantic information describing the important high-level details (what, where, who and how) of the observed event and exclude background information that is deemed unimportant to the observer. With this in mind, the descriptions people generate for videos of different dynamic events can greatly improve our understanding of the key information of interest in each video. These descriptions can be captured in captions that provide expanded attributes for video labeling (e.g. actions/objects/scenes/sentiment/etc.) while allowing us to gain new insight into what people find important or necessary to summarize specific events. Existing caption datasets for video understanding are either small in scale or restricted to a specific domain. To address this, we present the Spoken Moments (S-MiT) dataset of 500k spoken captions each attributed to a unique short video depicting a broad range of different events. We collect our descriptions using audio recordings to ensure that they remain as natural and concise as possible while allowing us to scale the size of a large classification dataset. In order to utilize our proposed dataset, we present a novel Adaptive Mean Margin (AMM) approach to contrastive learning and evaluate our models on video/caption retrieval on multiple datasets. We show that our AMM approach consistently improves our results and that models trained on our Spoken Moments dataset generalize better than those trained on other video-caption datasets.http://moments.csail.mit.edu/spoken.html Mathew Monfort, SouYoung Jin, Alexander H. Liu, David F. Harwath, Rogério Feris, James R. Glass, Aude Oliva |
CVPR | 5 |
| 2021 | Semi-Supervised Action Recognition With Temporal Contrastive LearningabstractLearning to recognize actions from only a handful of labeled videos is a challenging problem due to the scarcity of tediously collected activity labels. We approach this problem by learning a two-pathway temporal contrastive model using unlabeled videos at two different speeds lever-aging the fact that changing video speed does not change an action. Specifically, we propose to maximize the similarity between encoded representations of the same video at two different speeds as well as minimize the similarity between different videos played at different speeds. This way we use the rich supervisory information in terms of ‘time’ that is present in otherwise unsupervised pool of videos. With this simple yet effective strategy of manipulating video playback rates, we considerably outperform video extensions of sophisticated state-of-the-art semi-supervised image recognition methods across multiple diverse bench-mark datasets and network architectures. Interestingly, our proposed approach benefits from out-of-domain unlabeled videos showing generalization and robustness. We also per-form rigorous ablations and analysis to validate our approach. Project page: https://cvir.github.io/TCL/. Omprakash Chakraborty, Ashutosh Varshney, Rameswar Panda, Rogério Feris, Kate Saenko, Abir Das |
CVPR | 5 |
| 2021 | Separating Skills and Concepts for Novel Visual Question AnsweringabstractGeneralization to out-of-distribution data has been a problem for Visual Question Answering (VQA) models. To measure generalization to novel questions, we propose to separate them into "skills" and "concepts". "Skills" are visual tasks, such as counting or attribute recognition, and are applied to "concepts" mentioned in the question, such as objects and people. VQA methods should be able to compose skills and concepts in novel ways, regardless of whether the specific composition has been seen in training, yet we demonstrate that existing models have much to improve upon towards handling new compositions. We present a novel method for learning to compose skills and concepts that separates these two factors implicitly within a model by learning grounded concept representations and disentangling the encoding of skills from that of concepts. We enforce these properties with a novel contrastive learning procedure that does not rely on external annotations and can be learned from unlabeled image-question pairs. Experiments demonstrate the effectiveness of our approach for improving compositional and grounding performance.1 Spencer Whitehead, Hui Wu 0009, Heng Ji 0001, Rogério Feris, Kate Saenko |
CVPR | 4 |
| 2021 | Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language FeedbackabstractConversational interfaces for the detail-oriented retail fashion domain are more natural, expressive, and user friendly than classical keyword-based search interfaces. In this paper, we introduce the Fashion IQ dataset to support and advance research on interactive fashion image retrieval. Fashion IQ is the first fashion dataset to provide human-generated captions that distinguish similar pairs of garment images together with side-information consisting of real-world product descriptions and derived visual attribute labels for these images. We provide a detailed analysis of the characteristics of the Fashion IQ data, and present a transformer-based user simulator and interactive image retriever that can seamlessly integrate visual attributes with image features, user feedback, and dialog history, leading to improved performance over the state of the art in dialog-based image retrieval. We believe that our dataset will encourage further work on developing more natural and real-world applicable conversational shopping assistants.1 Hui Wu 0009, Yupeng Gao, Ziad Al-Halah, Steven Rennie, Kristen Grauman, Rogério Feris |
CVPR | 7 |
| 2021 | Detector-Free Weakly Supervised Grounding by SeparationabstractNowadays, there is an abundance of data involving images and surrounding free-form text weakly corresponding to those images. Weakly Supervised phrase-Grounding (WSG) deals with the task of using this data to learn to localize (or to ground) arbitrary text phrases in images without any additional annotations. However, most recent SotA methods for WSG assume an existence of a pre-trained object detector, relying on it to produce the ROIs for localization. In this work, we focus on the task of Detector-Free WSG (DF-WSG) to solve WSG without relying on a pre-trained detector. The key idea behind our proposed Grounding by Separation (GbS) method is synthesizing ‘text to image-regions’ associations by random alpha-blending of arbitrary image pairs and using the corresponding texts of the pair as conditions to recover the alpha map from the blended image via a segmentation network. At test time, this allows using the query phrase as a condition for a non-blended query image, thus interpreting the test image as a composition of a region corresponding to the phrase and the complement region. Our GbS shows an 8.5% accuracy improvement over previous DF-WSG SotA, for a range of benchmarks including Flickr30K, Visual Genome, and ReferIt, as well as a complementary improvement (above 7%) over the detector-based approaches for WSG. Assaf Arbelle, Sivan Doveh, Amit Alfassy, Joseph Shtok, Guy Lev, Eli Schwartz, Hilde Kuehne, Hila Levi, Prasanna Sattigeri, Rameswar Panda, Chun-Fu Chen 0001, Alexander M. Bronstein, Kate Saenko, Shimon Ullman, Raja Giryes, Rogério Feris, Leonid Karlinsky |
ICCV | 16 |
| 2021 | Multimodal Clustering Networks for Self-supervised Learning from Unlabeled VideosabstractMultimodal self-supervised learning is getting more and more attention as it allows not only to train large networks without human supervision but also to search and retrieve data across various modalities. In this context, this paper proposes a framework that, starting from a pre-trained backbone, learns a common multimodal embedding space that, in addition to sharing representations across different modalities, enforces a grouping of semantically similar instances. To this end, we extend the concept of instance-level contrastive learning with a multimodal clustering step in the training pipeline to capture semantic similarities across modalities. The resulting embedding space enables retrieval of samples across all modalities, even from unseen datasets and different domains. To evaluate our approach, we train our model on the HowTo100M dataset and evaluate its zero-shot retrieval capabilities in two challenging domains, namely text-to-video retrieval, and temporal action localization, showing state-of-the-art results on four different datasets. Brian Chen 0001, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne, Samuel Thomas 0001, Angie W. Boggust, Rameswar Panda, Brian Kingsbury, Rogério Feris, David F. Harwath, James R. Glass, Michael Picheny, Shih-Fu Chang |
ICCV | 9 |
| 2021 | A Broad Study on the Transferability of Visual Representations with Contrastive LearningabstractTremendous progress has been made in visual representation learning, notably with the recent success of self-supervised contrastive learning methods. Supervised contrastive learning has also been shown to outperform its cross-entropy counterparts by leveraging labels for choosing where to contrast. However, there has been little work to explore the transfer capability of contrastive learning to a different domain. In this paper, we conduct a comprehensive study on the transferability of learned representations of different contrastive approaches for linear evaluation, full-network transfer, and few-shot recognition on 12 downstream datasets from different domains, and object detection tasks on MSCOCO and VOC0712. The results show that the contrastive approaches learn representations that are easily transferable to a different downstream task. We further observe that the joint objective of self-supervised contrastive loss with cross-entropy/supervised-contrastive loss leads to better transferability of these models over their supervised counterparts. Our analysis reveals that the representations learned from the contrastive approaches contain more low/mid-level semantics than cross-entropy models, which enables them to quickly adapt to a new task. Our codes and models will be publicly available to facilitate future research on transferability of visual representations.1 Ashraful Islam, Chun-Fu Chen 0001, Rameswar Panda, Leonid Karlinsky, Richard J. Radke, Rogério Feris |
ICCV | 6 |
| 2021 | AdaMML: Adaptive Multi-Modal Learning for Efficient Video RecognitionabstractMulti-modal learning, which focuses on utilizing various modalities to improve the performance of a model, is widely used in video recognition. While traditional multi-modal learning offers excellent recognition results, its computational expense limits its impact for many real-world applications. In this paper, we propose an adaptive multi-modal learning framework, called AdaMML, that selects on-the-fly the optimal modalities for each segment conditioned on the input for efficient video recognition. Specifically, given a video segment, a multi-modal policy net-work is used to decide what modalities should be used for processing by the recognition model, with the goal of improving both accuracy and efficiency. We efficiently train the policy network jointly with the recognition model using standard back-propagation. Extensive experiments on four challenging diverse datasets demonstrate that our proposed adaptive approach yields 35% − 55% reduction in computation when compared to the traditional baseline that simply uses all the modalities irrespective of the in-put, while also achieving consistent improvements in accuracy over the state-of-the-art methods. Project page: https://rpand002.github.io/adamml.html. Rameswar Panda, Chun-Fu Chen 0001, Quanfu Fan, Ximeng Sun, Kate Saenko, Aude Oliva, Rogério Feris |
ICCV | 7 |
| 2021 | Dynamic Network Quantization for Efficient Video InferenceabstractDeep convolutional networks have recently achieved great success in video recognition, yet their practical realization remains a challenge due to the large amount of computational resources required to achieve robust recognition. Motivated by the effectiveness of quantization for boosting efficiency, in this paper, we propose a dynamic network quantization framework, that selects optimal precision for each frame conditioned on the input for efficient video recognition. Specifically, given a video clip, we train a very lightweight network in parallel with the recognition network, to produce a dynamic policy indicating which numerical precision to be used per frame in recognizing videos. We train both networks effectively using standard backpropagation with a loss to achieve both competitive performance and resource efficiency required for video recognition. Extensive experiments on four challenging diverse benchmark datasets demonstrate that our proposed approach provides significant savings in computation and memory usage while outperforming the existing state-of-the-art methods. Project page: https://cs-people.bu.edu/sunxm/VideoIQ/project.html. Ximeng Sun, Rameswar Panda, Chun-Fu Chen 0001, Aude Oliva, Rogério Feris, Kate Saenko |
ICCV | 5 |
| 2021 | AdaFuse: Adaptive Temporal Fusion Network for Efficient Action Recognition
Rameswar Panda, Chung-Ching Lin, Prasanna Sattigeri, Leonid Karlinsky, Kate Saenko, Aude Oliva, Rogério Feris |
ICLR | 8 |
| 2021 | VA-RED2: Video Adaptive Redundancy Reduction
Bowen Pan, Rameswar Panda, Camilo Fosco, Chung-Ching Lin, Alex Andonian, Kate Saenko, Aude Oliva, Rogério Feris |
ICLR | 9 |
| 2021 | Cascaded Multilingual Audio-Visual Learning from VideosabstractIn this paper, we explore self-supervised audio-visual models that learn from instructional videos.Prior work has shown that these models can relate spoken words and sounds to visual content after training on a large-scale dataset of videos, but they were only trained and evaluated on videos in English.To learn multilingual audio-visual representations, we propose a cascaded approach that leverages a model trained on English videos and applies it to audio-visual data in other languages, such as Japanese videos.With our cascaded approach, we show an improvement in retrieval performance of nearly 10x compared to training on the Japanese videos solely.We also apply the model trained on English videos to Japanese and Hindi spoken captions of images, achieving state-of-the-art performance. Andrew Rouditchenko, Angie W. Boggust, David F. Harwath, Samuel Thomas 0001, Hilde Kuehne, Brian Chen 0001, Rameswar Panda, Rogério Feris, Brian Kingsbury, Michael Picheny, James R. Glass |
Interspeech | 8 |
| 2021 | AVLnet: Learning Audio-Visual Language Representations from Instructional VideosabstractCurrent methods for learning visually grounded language from videos often rely on text annotation, such as human generated captions or machine generated automatic speech recognition (ASR) transcripts. In this work, we introduce the Audio-Video Language Network (AVLnet), a self-supervised network that learns a shared audio-visual embedding space directly from raw video inputs. To circumvent the need for text annotation, we learn audio-visual representations from randomly segmented video clips and their raw audio waveforms. We train AVLnet on HowTo100M, a large corpus of publicly available instructional videos, and evaluate on image retrieval and video retrieval tasks, achieving state-of-the-art performance. We perform analysis of AVLnet's learned representations, showing our model utilizes speech and natural sounds to learn audio-visual concepts. Further, we propose a tri-modal model that jointly processes raw audio, video, and text captions from videos to learn a multi-modal semantic embedding space useful for text-video retrieval. Our code, data, and trained models will be released at avlnet.csail.mit.edu Andrew Rouditchenko, Angie W. Boggust, David F. Harwath, Brian Chen 0001, Dhiraj Joshi, Samuel Thomas 0001, Kartik Audhkhasi, Hilde Kuehne, Rameswar Panda, Rogério Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba 0001, James R. Glass |
Interspeech | 10 |
| 2021 | Dynamic Distillation Network for Cross-Domain Few-Shot Recognition with Unlabeled DataabstractMost existing works in few-shot learning rely on meta-learning the network on a large base dataset which is typically from the same domain as the target dataset. We tackle the problem of cross-domain few-shot learning where there is a large shift between the base and target domain. The problem of cross-domain few-shot recognition with unlabeled target data is largely unaddressed in the literature. STARTUP was the first method that tackles this problem using self-training. However, it uses a fixed teacher pretrained on a labeled base dataset to create soft labels for the unlabeled target samples. As the base dataset and unlabeled dataset are from different domains, projecting the target images in the class-domain of the base dataset with a fixed pretrained model might be sub-optimal. We propose a simple dynamic distillation-based approach to facilitate unlabeled images from the novel/base dataset. We impose consistency regularization by calculating predictions from the weakly-augmented versions of the unlabeled images from a teacher network and matching it with the strongly augmented versions of the same images from a student network. The parameters of the teacher network are updated as exponential moving average of the parameters of the student network. We show that the proposed network learns representation that can be easily adapted to the target domain even though it has not been trained with target-specific classes during the pretraining phase. Our model outperforms the current state-of-the art method by 4.4% for 1-shot and 3.6% for 5-shot classification in the BSCD-FSL benchmark, and also shows competitive performance on traditional in-domain few-shot learning task. Ashraful Islam, Chun-Fu Chen 0001, Rameswar Panda, Leonid Karlinsky, Rogério Feris, Richard J. Radke |
NeurIPS | 5 |
| 2021 | IA-RED$^2$: Interpretability-Aware Redundancy Reduction for Vision TransformersabstractThe self-attention-based model, transformer, is recently becoming the leading backbone in the field of computer vision. In spite of the impressive success made by transformers in a variety of vision tasks, it still suffers from heavy computation and intensive memory costs. To address this limitation, this paper presents an Interpretability-Aware REDundancy REDuction framework (IA-RED$^2$). We start by observing a large amount of redundant computation, mainly spent on uncorrelated input patches, and then introduce an interpretable module to dynamically and gracefully drop these redundant patches. This novel framework is then extended to a hierarchical structure, where uncorrelated tokens at different stages are gradually removed, resulting in a considerable shrinkage of computational cost. We include extensive experiments on both image and video tasks, where our method could deliver up to 1.4x speed-up for state-of-the-art models like DeiT and TimeSformer, by only sacrificing less than 0.7% accuracy. More importantly, contrary to other acceleration approaches, our method is inherently interpretable with substantial visual evidence, making vision transformer closer to a more human-understandable architecture while being lighter. We demonstrate that the interpretability that naturally emerged in our framework can outperform the raw attention learned by the original visual transformer, as well as those generated by off-the-shelf interpretation methods, with both qualitative and quantitative results. Project Page: http://people.csail.mit.edu/bpan/ia-red/. Bowen Pan, Rameswar Panda, Yifan Jiang 0001, Zhangyang Wang, Rogério Feris, Aude Oliva |
NeurIPS | 5 |
| 2021 | MetAdapt: Meta-learned task-adaptive architecture for few-shot classification
Sivan Doveh, Eli Schwartz, Rogério Feris, Alexander M. Bronstein, Raja Giryes, Leonid Karlinsky |
Pattern Recognit. Lett. | 4 |
| 2020 | Video Instance Segmentation Tracking With a Modified VAE ArchitectureabstractWe propose a modified variational autoencoder (VAE) architecture built on top of Mask R-CNN for instance-level video segmentation and tracking. The method builds a shared encoder and three parallel decoders, yielding three disjoint branches for predictions of future frames, object detection boxes, and instance segmentation masks. To effectively solve multiple learning tasks, we introduce a Gaussian Process model to enhance the statistical representation of VAE by relaxing the prior strong independent and identically distributed (iid) assumption of conventional VAEs and allowing potential correlations among extracted latent variables. The network learns embedded spatial interdependence and motion continuity in video data and creates a representation that is effective to produce high-quality segmentation masks and track multiple instances in diverse and unstructured videos. Evaluation on a variety of recently introduced datasets shows that our model outperforms previous methods and achieves the new best in class performance. Chung-Ching Lin, Ying Hung, Rogério Feris, Linglin He |
CVPR | 3 |
| 2020 | Differential Treatment for Stuff and Things: A Simple Unsupervised Domain Adaptation Method for Semantic SegmentationabstractWe consider the problem of unsupervised domain adaptation for semantic segmentation by easing the domain shift between the source domain (synthetic data) and the target domain (real data) in this work. State-of-the-art approaches prove that performing semantic-level alignment is helpful in tackling the domain shift issue. Based on the observation that stuff categories usually share similar appearances across images of different domains while things (i.e. object instances) have much larger differences, we propose to improve the semantic-level alignment with different strategies for stuff regions and for things: 1) for the stuff categories, we generate feature representation for each class and conduct the alignment operation from the target domain to the source domain; 2) for the thing categories, we generate feature representation for each individual instance and encourage the instance in the target domain to align with the most similar one in the source domain. In this way, the individual differences within thing categories will also be considered to alleviate over-alignment. In addition to our proposed method, we further reveal the reason why the current adversarial loss is often unstable in minimizing the distribution discrepancy and show that our method can help ease this issue by minimizing the most similar stuff and instance features between the source and the target domains. We conduct extensive experiments in two unsupervised domain adaptation tasks, i.e. GTA5 → Cityscapes and SYNTHIA → Cityscapes, and achieve the new state-of-the-art segmentation accuracy. Zhonghao Wang 0001, Mo Yu, Yunchao Wei, Rogério Feris, Jinjun Xiong, Wen-Mei W. Hwu, Thomas S. Huang, Humphrey Shi |
CVPR | 4 |
| 2020 | OnlineAugment: Online Data Augmentation with Less Domain Knowledge
Zhiqiang Tang 0001, Yunhe Gao, Leonid Karlinsky, Prasanna Sattigeri, Rogério Feris, Dimitris N. Metaxas |
ECCV (7) | 5 |
| 2020 | We Have So Much in Common: Modeling Semantic Relational Set Abstractions in Videos
Alex Andonian, Camilo Fosco, Mathew Monfort, Allen Lee, Rogério Feris, Carl Vondrick, Aude Oliva |
ECCV (18) | 5 |
| 2020 | A Broader Study of Cross-Domain Few-Shot Learning
Yunhui Guo, Noel Codella, Leonid Karlinsky, James V. Codella, John R. Smith, Kate Saenko, Tajana Rosing, Rogério Feris |
ECCV (27) | 8 |
| 2020 | TAFSSL: Task-Adaptive Feature Sub-Space Learning for Few-Shot Classification
Moshe Lichtenstein, Prasanna Sattigeri, Rogério Feris, Raja Giryes, Leonid Karlinsky |
ECCV (7) | 3 |
| 2020 | AR-Net: Adaptive Frame Resolution for Efficient Action Recognition
Chung-Ching Lin, Rameswar Panda, Prasanna Sattigeri, Leonid Karlinsky, Aude Oliva, Kate Saenko, Rogério Feris |
ECCV (7) | 8 |
| 2020 | AdaShare: Learning What To Share For Efficient Deep Multi-Task LearningabstractMulti-task learning is an open and challenging problem in computer vision. The typical way of conducting multi-task learning with deep neural networks is either through handcrafted schemes that share all initial layers and branch out at an adhoc point, or through separate task-specific networks with an additional feature sharing/fusion mechanism. Unlike existing methods, we propose an adaptive sharing approach, calledAdaShare, that decides what to share across which tasks to achieve the best recognition accuracy, while taking resource efficiency into account. Specifically, our main idea is to learn the sharing pattern through a task-specific policy that selectively chooses which layers to execute for a given task in the multi-task network. We efficiently optimize the task-specific policy jointly with the network weights, using standard back-propagation. Experiments on several challenging and diverse benchmark datasets with a variable number of tasks well demonstrate the efficacy of our approach over state-of-the-art methods. Project page: https://cs-people.bu.edu/sunxm/AdaShare/project.html Ximeng Sun, Rameswar Panda, Rogério Feris, Kate Saenko |
NeurIPS | 3 |
| 2019 | Improving Object Detection from Scratch via Gated Feature Reuse
Humphrey Shi, NhatHai Phan, Rogério Feris, Liangliang Cao, Ding Liu 0001, Xinchao Wang, Thomas S. Huang, Marios Savvides |
BMVC | 5 |
| 2019 | LaSO: Label-Set Operations Networks for Multi-Label Few-Shot LearningabstractExample synthesis is one of the leading methods to tackle the problem of few-shot learning, where only a small number of samples per class are available. However, current synthesis approaches only address the scenario of a single category label per image. In this work, we propose a novel technique for synthesizing samples with multiple labels for the (yet unhandled) multi-label few-shot classification scenario. We propose to combine pairs of given examples in feature space, so that the resulting synthesized feature vectors will correspond to examples whose label sets are obtained through certain set operations on the label sets of the corresponding input pairs. Thus, our method is capable of producing a sample containing the intersection, union or set-difference of labels present in two input samples. As we show, these set operations generalize to labels unseen during training. This enables performing augmentation on examples of novel categories, thus, facilitating multi-label few-shot classifier learning. We conduct numerous experiments showing promising results for the label-set manipulation capabilities of the proposed approach, both directly (using the classification and retrieval metrics), and in the context of performing data augmentation for multi-label few-shot learning. We propose a benchmark for this new and challenging task and show that our method compares favorably to all the common baselines. Amit Alfassy, Leonid Karlinsky, Amit Aides, Joseph Shtok, Sivan Harary, Rogério Feris, Raja Giryes, Alexander M. Bronstein |
CVPR | 6 |
| 2019 | SpotTune: Transfer Learning Through Adaptive Fine-TuningabstractTransfer learning, which allows a source task to affect the inductive bias of the target task, is widely used in computer vision. The typical way of conducting transfer learning with deep neural networks is to fine-tune a model pretrained on the source task using data from the target task. In this paper, we propose an adaptive fine-tuning approach, called SpotTune, which finds the optimal fine-tuning strategy per instance for the target data. In SpotTune, given an image from the target task, a policy network is used to make routing decisions on whether to pass the image through the fine-tuned layers or the pre-trained layers. We conduct extensive experiments to demonstrate the effectiveness of the proposed approach. Our method outperforms the traditional fine-tuning approach on 12 out of 14 standard datasets. We also compare SpotTune with other state-of-the-art fine-tuning strategies, showing superior performance. On the Visual Decathlon datasets, our method achieves the highest score across the board without bells and whistles. Yunhui Guo, Humphrey Shi, Abhishek Kumar 0001, Kristen Grauman, Tajana Rosing, Rogério Feris |
CVPR | 6 |
| 2019 | RepMet: Representative-Based Metric Learning for Classification and Few-Shot Object DetectionabstractDistance metric learning (DML) has been successfully applied to object classification, both in the standard regime of rich training data and in the few-shot scenario, where each category is represented by only a few examples. In this work, we propose a new method for DML that simultaneously learns the backbone network parameters, the embedding space, and the multi-modal distribution of each of the training categories in that space, in a single end-to-end training process. Our approach outperforms state-of-the-art methods for DML-based object classification on a variety of standard fine-grained datasets. Furthermore, we demonstrate the effectiveness of our approach on the problem of few-shot object detection, by incorporating the proposed DML architecture as a classification head into a standard object detection model. We achieve the best results on the ImageNet-LOC dataset compared to strong baselines, when only a few training examples are available. We also offer the community a new episodic benchmark based on the ImageNet dataset for the few-shot object detection task. Leonid Karlinsky, Joseph Shtok, Sivan Harary, Eli Schwartz, Amit Aides, Rogério Feris, Raja Giryes, Alexander M. Bronstein |
CVPR | 6 |
| 2019 | Learning Motion in Feature Space: Locally-Consistent Deformable Convolution Networks for Fine-Grained Action DetectionabstractFine-grained action detection is an important task with numerous applications in robotics and human-computer interaction. Existing methods typically utilize a two-stage approach including extraction of local spatio-temporal features followed by temporal modeling to capture long-term dependencies. While most recent papers have focused on the latter (long-temporal modeling), here, we focus on producing features capable of modeling fine-grained motion more efficiently. We propose a novel locally-consistent deformable convolution, which utilizes the change in receptive fields and enforces a local coherency constraint to capture motion information effectively. Our model jointly learns spatio-temporal features (instead of using independent spatial and temporal streams). The temporal component is learned from the feature space instead of pixel space, e.g. optical flow. The produced features can be flexibly used in conjunction with other long-temporal modeling networks, e.g. ST-CNN, DilatedTCN, and ED-TCN. Overall, our proposed approach robustly outperforms the original long-temporal models on two fine-grained action datasets: 50 Salads and GTEA, achieving F1 scores of 80.22% and 75.39% respectively. Khoi-Nguyen C. Mac, Dhiraj Joshi, Raymond A. Yeh, Jinjun Xiong, Rogério Feris, Minh N. Do |
ICCV | 5 |
| 2019 | Big-Little Net: An Efficient Multi-Scale Feature Representation for Visual and Speech Recognition
Chun-Fu Chen 0001, Quanfu Fan, Neil Mallinar, Tom Sercu, Rogério Feris |
ICLR (Poster) | 5 |
| 2019 | Automatic Curation of Sports Highlights Using Multimodal Excitement FeaturesabstractThe production of sports highlight packages summarizing a game's most exciting moments is an essential task for broadcast media. Yet, it requires labor-intensive video editing. We propose a novel approach for auto-curating sports highlights, and demonstrate it to create a first of a kind, real-world system for the editorial aid of golf and tennis highlight reels. Our method fuses information from the players’ reactions (action recognition such as high-fives and fist pumps), players’ expressions (aggressive, tense, smiling, and neutral), spectators (crowd cheering), commentator (tone of the voice and word analysis), and game analytics to determine the most interesting moments of a game. We accurately identify the start and end frames of key shot highlights with additional metadata, such as the player's name and the whole number, or analysts input allowing personalized content summarization and retrieval. In addition, we introduce new techniques for learning our classifiers with reduced manual training data annotation by exploiting the correlation of different modalities. Our work has been demonstrated at a major golf tournament (2017 Masters) and two major international tennis tournaments (2017 Wimbledon and U.S. Open), successfully extracting highlights through the course of the sporting events. For the 2017 Masters, 54% of the clips selected by our system overlapped with the official highlights reels. Furthermore, user studies showed that 90% of the non-overlapping ones were of the same quality of the official clips for the 2017 Masters, while the automatic selection of clips for highlights of 2017 Wimbledon and 2017 US Open agreed with human preferences 80% and 84.2% of the time, respectively. Michele Merler, Khoi-Nguyen C. Mac, Dhiraj Joshi, Quoc-Bao Nguyen, Stephen Hammer, John Kent, Jinjun Xiong, Minh N. Do, John R. Smith, Rogério Feris |
IEEE Trans. Multim. | 10 |
| 2018 | Jointly Optimize Data Augmentation and Network Training: Adversarial Data Augmentation in Human Pose EstimationabstractRandom data augmentation is a critical technique to avoid overfitting in training deep models. Yet, data augmentation and network training are often two isolated processes in most settings, yielding to a suboptimal training. Why not jointly optimize the two? We propose adversarial data augmentation to address this limitation. The key idea is to design a generator (e.g. an augmentation network) that competes against a discriminator (e.g. a target network) by generating hard examples online. The generator explores weaknesses of the discriminator, while the discriminator learns from hard augmentations to achieve better performance. A reward/penalty strategy is also proposed for efficient joint training. We investigate human pose estimation and carry out comprehensive ablation studies to validate our method. The results prove that our method can effectively improve state-of-the-art models without additional data effort. Xi Peng 0005, Zhiqiang Tang 0001, Fei Yang 0001, Rogério Feris, Dimitris N. Metaxas |
CVPR | 4 |
| 2018 | BlockDrop: Dynamic Inference Paths in Residual NetworksabstractVery deep convolutional neural networks offer excellent recognition results, yet their computational expense limits their impact for many real-world applications. We introduce BlockDrop, an approach that learns to dynamically choose which layers of a deep network to execute during inference so as to best reduce total computation without degrading prediction accuracy. Exploiting the robustness of Residual Networks (ResNets) to layer dropping, our framework selects on-the-fly which residual blocks to evaluate for a given novel image. In particular, given a pretrained ResNet, we train a policy network in an associative reinforcement learning setting for the dual reward of utilizing a minimal number of blocks while preserving recognition accuracy. We conduct extensive experiments on CIFAR and ImageNet. The results provide strong quantitative and qualitative evidence that these learned policies not only accelerate inference but also encode meaningful visual information. Built upon a ResNet-101 model, our method achieves a speedup of 20% on average, going as high as 36% for some images, while maintaining the same 76.4% top-1 accuracy on ImageNet. Zuxuan Wu, Tushar Nagarajan, Abhishek Kumar 0001, Steven Rennie, Larry Davis 0001, Kristen Grauman, Rogério Feris |
CVPR | 7 |
| 2018 | Revisiting RCNN: On Awakening the Classification Power of Faster RCNN
Bowen Cheng, Yunchao Wei, Humphrey Shi, Rogério Feris, Jinjun Xiong, Thomas S. Huang |
ECCV (15) | 4 |
| 2018 | Learning to Separate Object Sounds by Watching Unlabeled Video
Ruohan Gao, Rogério Feris, Kristen Grauman |
ECCV (3) | 2 |
| 2018 | Dialog-based Interactive Image RetrievalabstractExisting methods for interactive image retrieval have demonstrated the merit of integrating user feedback, improving retrieval results. However, most current systems rely on restricted forms of user feedback, such as binary relevance responses, or feedback based on a fixed set of relative attributes, which limits their impact. In this paper, we introduce a new approach to interactive image search that enables users to provide feedback via natural language, allowing for more natural and effective interaction. We formulate the task of dialog-based interactive image retrieval as a reinforcement learning problem, and reward the dialog system for improving the rank of the target image during each dialog turn. To mitigate the cumbersome and costly process of collecting human-machine conversations as the dialog system learns, we train our system with a user simulator, which is itself trained to describe the differences between target and candidate images. The efficacy of our approach is demonstrated in a footwear retrieval application. Experiments on both simulated and real-world data show that 1) our proposed learning framework achieves better accuracy than other supervised and reinforcement learning baselines and 2) user feedback based on natural language rather than pre-specified attributes leads to more effective retrieval results, and a more natural and expressive communication interface. Hui Wu 0009, Yu Cheng 0001, Steven Rennie, Gerald Tesauro, Rogério Feris |
NeurIPS | 6 |
| 2018 | Co-regularized Alignment for Unsupervised Domain AdaptationabstractDeep neural networks, trained with large amount of labeled data, can fail to generalize well when tested with examples from a target domain whose distribution differs from the training data distribution, referred as the source domain. It can be expensive or even infeasible to obtain required amount of labeled data in all possible domains. Unsupervised domain adaptation sets out to address this problem, aiming to learn a good predictive model for the target domain using labeled examples from the source domain but only unlabeled examples from the target domain. Domain alignment approaches this problem by matching the source and target feature distributions, and has been used as a key component in many state-of-the-art domain adaptation methods. However, matching the marginal feature distributions does not guarantee that the corresponding class conditional distributions will be aligned across the two domains. We propose co-regularized domain alignment for unsupervised domain adaptation, which constructs multiple diverse feature spaces and aligns source and target distributions in each of them individually, while encouraging that alignments agree with each other with regard to the class predictions on the unlabeled target examples. The proposed method is generic and can be used to improve any domain adaptation method which uses domain alignment. We instantiate it in the context of a recent state-of-the-art method and observe that it provides significant performance improvements on several domain adaptation benchmarks. Abhishek Kumar 0001, Prasanna Sattigeri, Kahini Wadhawan, Leonid Karlinsky, Rogério Feris, William T. Freeman, Gregory W. Wornell |
NeurIPS | 5 |
| 2018 | Delta-encoder: an effective sample synthesis method for few-shot object recognitionabstractLearning to classify new categories based on just one or a few examples is a long-standing challenge in modern computer vision. In this work, we propose a simple yet effective method for few-shot (and one-shot) object recognition. Our approach is based on a modified auto-encoder, denoted delta-encoder, that learns to synthesize new samples for an unseen category just by seeing few examples from it. The synthesized samples are then used to train a classifier. The proposed approach learns to both extract transferable intra-class deformations, or "deltas", between same-class pairs of training examples, and to apply those deltas to the few provided examples of a novel class (unseen during training) in order to efficiently synthesize samples from that new class. The proposed method improves the state-of-the-art of one-shot object-recognition and performs comparably in the few-shot case. Eli Schwartz, Leonid Karlinsky, Joseph Shtok, Sivan Harary, Mattias Marder, Abhishek Kumar 0001, Rogério Feris, Raja Giryes, Alexander M. Bronstein |
NeurIPS | 7 |
| 2018 | RED-Net: A Recurrent Encoder-Decoder Network for Video-Based Face Alignment
Xi Peng 0005, Rogério Feris, Xiaoyu Wang 0002, Dimitris N. Metaxas |
Int. J. Comput. Vis. | 2 |
| 2017 | Fully-Adaptive Feature Sharing in Multi-Task Networks with Applications in Person Attribute ClassificationabstractMulti-task learning aims to improve generalization performance of multiple prediction tasks by appropriately sharing relevant information across them. In the context of deep neural networks, this idea is often realized by hand-designed network architectures with layers that are shared across tasks and branches that encode task-specific features. However, the space of possible multi-task deep architectures is combinatorially large and often the final architecture is arrived at by manual exploration of this space, which can be both error-prone and tedious. We propose an automatic approach for designing compact multi-task deep learning architectures. Our approach starts with a thin multi-layer network and dynamically widens it in a greedy manner during training. By doing so iteratively, it creates a tree-like deep architecture, on which similar tasks reside in the same branch until at the top layers. Evaluation on person attributes classification tasks involving facial and clothing attributes suggests that the models produced by the proposed method are fast, compact and can closely match or exceed the state-of-the-art accuracy from strong baselines by much more expensive models. Yongxi Lu, Abhishek Kumar 0001, Shuangfei Zhai, Yu Cheng 0001, Tara Javidi, Rogério Feris |
CVPR | 6 |
| 2017 | S3Pool: Pooling with Stochastic Spatial Sampling
Shuangfei Zhai, Hui Wu 0009, Abhishek Kumar 0001, Yu Cheng 0001, Yongxi Lu, Zhongfei Zhang, Rogério Feris |
CVPR | 7 |
| 2017 | IBM High-Five: Highlights From Intelligent Video EngineabstractWe introduce a novel multi-modal system for auto-curating golf highlights that fuses information from players' reactions (celebration actions), spectators (crowd cheering), and commentator (tone of the voice and word analysis) to determine the most interesting moments of a game. The start of a highlight is determined with additional metadata (player's name and the hole number), allowing personalized content summarization and retrieval. Our system was demonstrated at Masters 2017, a major golf tournament, generating real-time highlights from four live video streams over four days. Dhiraj Joshi, Michele Merler, Quoc-Bao Nguyen, Stephen Hammer, John Kent, John R. Smith, Rogério Feris |
ACM Multimedia | 7 |
| 2016 | Walk and Learn: Facial Attribute Representation Learning from Egocentric Video and Contextual DataabstractThe way people look in terms of facial attributes (ethnicity, hair color, facial hair, etc.) and the clothes or accessories they wear (sunglasses, hat, hoodies, etc.) is highly dependent on geo-location and weather condition, respectively. This work explores, for the first time, the use of this contextual information, as people with wearable cameras walk across different neighborhoods of a city, in order to learn a rich feature representation for facial attribute classification, without the costly manual annotation required by previous methods. By tracking the faces of casual walkers on more than 40 hours of egocentric video, we are able to cover tens of thousands of different identities and automatically extract nearly 5 million pairs of images connected by or from different face tracks, along with their weather and location context, under pose and lighting variations. These image pairs are then fed into a deep network that preserves similarity of images connected by the same track, in order to capture identity-related attribute features, and optimizes for location and weather prediction to capture additional facial attribute features. Finally, the network is fine-tuned with manually annotated samples. We perform an extensive experimental analysis on wearable data and two standard benchmark datasets based on web images (LFWA and CelebA). Our method outperforms by a large margin a network trained from scratch. Moreover, even without using manually annotated identity labels for pre-training as in previous methods, our approach achieves results that are better than the state of the art. Jing Wang 0069, Yu Cheng 0001, Rogério Feris |
CVPR | 3 |
| 2016 | A Unified Multi-scale Deep Convolutional Neural Network for Fast Object Detection
Zhaowei Cai, Quanfu Fan, Rogério Feris, Nuno Vasconcelos |
ECCV (4) | 3 |
| 2016 | A Recurrent Encoder-Decoder Network for Sequential Face Alignment
Xi Peng 0005, Rogério Feris, Xiaoyu Wang 0002, Dimitris N. Metaxas |
ECCV (1) | 2 |
| 2016 | Edge-Guided Single Depth Image Super ResolutionabstractRecently, consumer depth cameras have gained significant popularity due to their affordable cost. However, the limited resolution and the quality of the depth map generated by these cameras are still problematic for several applications. In this paper, a novel framework for the single depth image superresolution is proposed. In our framework, the upscaling of a single depth image is guided by a high-resolution edge map, which is constructed from the edges of the low-resolution depth image through a Markov random field optimization in a patch synthesis based manner. We also explore the self-similarity of patches during the edge construction stage, when limited training data are available. With the guidance of the high-resolution edge map, we propose upsampling the high-resolution depth image through a modified joint bilateral filter. The edge-based guidance not only helps avoiding artifacts introduced by direct texture prediction, but also reduces jagged artifacts and preserves the sharp edges. Experimental results demonstrate the effectiveness of our method both qualitatively and quantitatively compared with the state-of-the-art methods. Rogério Feris, Ming-Ting Sun |
IEEE Trans. Image Process. | 2 |
| 2015 | Efficient 24/7 object detection in surveillance videosabstractWe address the problem of 24/7 object detection in urban surveillance videos, which presents unique challenges due to significant object appearance variations caused by lighting effects such as shadows and specular reflections, object pose variation, multiple weather conditions, and different times of the day. Rather than training a generic detector and adapting its parameters over time to handle all these variations, we rely on a large set of complementary and extremely efficient detector models, covering multiple overlapping appearance subspaces. At run time, our method continuously selects the most suitable detectors for a given scene and condition, using a novel approach inspired by parametric background modeling algorithms. We provide a comprehensive experimental analysis to show the effectiveness of our approach, considering traffic monitoring as our application domain. Our system runs at 100 frames per second on a standard laptop computer. Rogério Feris, Russell Bobbitt, Sharath Pankanti, Ming-Ting Sun |
AVSS | 1 |
| 2015 | Deep domain adaptation for describing people based on fine-grained clothing attributesabstractWe address the problem of describing people based on fine-grained clothing attributes. This is an important problem for many practical applications, such as identifying target suspects or finding missing people based on detailed clothing descriptions in surveillance videos or consumer photos. We approach this problem by first mining clothing images with fine-grained attribute labels from online shopping stores. A large-scale dataset is built with about one million images and fine-detailed attribute sub-categories, such as various shades of color (e.g., watermelon red, rosy red, purplish red), clothing types (e.g., down jacket, denim jacket), and patterns (e.g., thin horizontal stripes, houndstooth). As these images are taken in ideal pose/lighting/background conditions, it is unreliable to directly use them as training data for attribute prediction in the domain of unconstrained images captured, for example, by mobile phones or surveillance cameras. In order to bridge this gap, we propose a novel double-path deep domain adaptation network to model the data from the two domains jointly. Several alignment cost layers placed inbetween the two columns ensure the consistency of the two domain features and the feasibility to predict unseen attribute categories in one of the domains. Finally, to achieve a working system with automatic human body alignment, we trained an enhanced RCNN-based detector to localize human bodies in images. Our extensive experimental evaluation demonstrates the effectiveness of the proposed approach for describing people based on fine-grained clothing attributes. Qiang Chen 0007, Junshi Huang, Rogério Feris, Lisa M. Brown, Jian Dong 0011, Shuicheng Yan |
CVPR | 3 |
| 2015 | An Exploration of Parameter Redundancy in Deep Networks with Circulant ProjectionsabstractWe explore the redundancy of parameters in deep neural networks by replacing the conventional linear projection in fully-connected layers with the circulant projection. The circulant structure substantially reduces memory footprint and enables the use of the Fast Fourier Transform to speed up the computation. Considering a fully-connected neural network layer with d input nodes, and d output nodes, this method improves the time complexity from O(d2) to O(dlogd) and space complexity from O(d2) to O(d). The space savings are particularly important for modern deep convolutional neural network architectures, where fully-connected layers typically contain more than 90% of the network parameters. We further show that the gradient computation and optimization of the circulant projections can be performed very efficiently. Our experiments on three standard datasets show that the proposed approach achieves this significant gain in storage and efficiency with minimal increase in error rate compared to neural networks with unstructured projections. Yu Cheng 0001, Felix X. Yu, Rogério Feris, Sanjiv Kumar, Alok N. Choudhary, Shih-Fu Chang |
ICCV | 3 |
| 2015 | Cross-Domain Image Retrieval with a Dual Attribute-Aware Ranking NetworkabstractWe address the problem of cross-domain image retrieval, considering the following practical application: given a user photo depicting a clothing image, our goal is to retrieve the same or attribute-similar clothing items from online shopping stores. This is a challenging problem due to the large discrepancy between online shopping images, usually taken in ideal lighting/pose/background conditions, and user photos captured in uncontrolled conditions. To address this problem, we propose a Dual Attribute-aware Ranking Network (DARN) for retrieval feature learning. More specifically, DARN consists of two sub-networks, one for each domain, whose retrieval feature representations are driven by semantic attribute learning. We show that this attribute-guided learning is a key factor for retrieval accuracy improvement. In addition, to further align with the nature of the retrieval problem, we impose a triplet visual similarity constraint for learning to rank across the two subnetworks. Another contribution of our work is a large-scale dataset which makes the network learning feasible. We exploit customer review websites to crawl a large set of online shopping images and corresponding offline user photos with fine-grained clothing attributes, i.e., around 450,000 online shopping images and about 90,000 exact offline counterpart images of those online ones. All these images are collected from real-world consumer websites reflecting the diversity of the data modality, which makes this dataset unique and rare in the academic community. We extensively evaluate the retrieval performance of networks in different configurations. The top-20 retrieval accuracy is doubled when using the proposed DARN other than the current popular solution using pre-trained CNN features only (0.570 vs. 0.268). Junshi Huang, Rogério Feris, Qiang Chen 0007, Shuicheng Yan |
ICCV | 2 |
| 2015 | Automated Axon Segmentation from Highly Noisy Microscopic VideosabstractWe present a novel method for automated segmentation of axons in extremely noisy videos obtained via two-photon microscopy in awake mice. We formulate segmentation as a pixel-wise classification problem in which a pixel is classified into "axon" or "non-axon" based on its feature vector. In order to deal with high levels of noise, the features of our classifier are derived from spatio-temporal Independent Component Analysis (stICA) which effectively isolates noise from signal components while leveraging temporal coherence from the video. We fit parametric models to represent the distribution of the extracted features and apply a probabilistic classifier over stICA components to determine the label of each pixel. Finally, we show compelling qualitative and quantitative results from very challenging two-photon microscopic, demonstrating the usefulness of our approach. An example time-series of two-photon images with our automated ROI extraction over layed is available with the supplemental materials. John Bowler, Rogério Feris, Liangliang Cao |
WACV | 2 |
| 2015 | Fine registration of 3D point clouds fusing structural and photometric information using an RGB-D camera
Yu-Feng Hsu, Rogério Feris, Ming-Ting Sun |
J. Vis. Commun. Image Represent. | 3 |
| 2015 | Joint Super Resolution and Denoising From a Single Depth ImageabstractThis paper describes a new algorithm for depth image super resolution and denoising using a single depth image as input. A robust coupled dictionary learning method with locality coordinate constraints is introduced to reconstruct the corresponding high resolution depth map. The local constraints effectively reduce the prediction uncertainty and prevent the dictionary from over-fitting. We also incorporate an adaptively regularized shock filter to simultaneously reduce the jagged noise and sharpen the edges. Furthermore, a joint reconstruction and smoothing framework is proposed with an L0 gradient smooth constraint, making the reconstruction more robust to noise. Experimental results demonstrate the effectiveness of our proposed algorithm compared to previously reported methods. Rogério Feris, Shiaw-Shian Yu, Ming-Ting Sun |
IEEE Trans. Multim. | 2 |
| 2014 | Fusing well-crafted feature descriptors for efficient fine-grained classificationabstractAs citizen science projects become more popular and engage an increasing number of volunteers, smartphones are turning into commonly used sensors in the biodiversity environment. In this paper, we propose a novel approach for classification of subordinate categories such as plant and insect species that is fast and suitable for use in mobile devices. In particular, we show that a combination of carefully designed features, including a robust shape descriptor to capture fine morphological structures of objects, as well as traditional color and texture features, is essential for obtaining good performance. A novel weighting technique assigns different costs to each feature, taking into account the inter-class and intra-class variation between species. We tested our proposed method in the popular Oxford Flower Dataset and in the Leeds Butterfly Dataset. We are able to achieve state-of-the-art accuracy while proposing an efficient approach that is suitable for mobile applications and can be applied to different species. Andrea Britto Mottos, Rogério Feris |
ICIP | 2 |
| 2014 | Edge guided single depth image super resolutionabstractRecently, consumer depth cameras have gained significant popularity due to their affordable cost. However, the limited resolution and quality of the depth map generated by these cameras are still problems for several applications. In this paper, we propose a novel framework for single depth image super resolution guided by a high resolution edge map constructed from the edges in the low resolution depth image via a Markov Random Field (MRF) optimization. With the guidance of the high resolution edge map, the high resolution depth image is up-sampled via a joint bilateral filter. The edge guidance not only helps avoid artifacts introduced by direct texture prediction, but also reduces the jagged artifacts and preserves the sharp edges. Experimental results demonstrate the effectiveness of our proposed algorithm compared to previously reported methods. Rogério Feris, Ming-Ting Sun |
ICIP | 2 |
| 2014 | RiskWheel: Interactive visual analytics for surveillance event detectionabstractDetecting human behaviors in vast amounts of video is a challenging task in a variety of real-world applications. Thus an interactive tool designed to support this task with human in the loop is of significance in various domains including public safety and security. In this paper, we design and develop an interactive visual analytics system, RiskWheel, that enables effective analysis of detection results and utilization of user feedback to improve surveillance event detection. In particular, we propose 1) an interactive approach to visualize data with temporal relations and 2) a novel risk ranking method to differentiate detection results and present more informative ones to the user for better interaction. In our experiments, we demonstrate RiskWheel through a case study on TRECVID Surveillance Event Detection (SED) task [1]. The experimental results quantitatively show that RiskWheel outperforms multiple baselines, demonstrating the power of the risk ranking technique. Yu Cheng 0001, Lisa M. Brown, Qibin Fan, Rogério Feris, Sharath Pankanti, Tao Zhang 0006 |
ICME | 4 |
| 2014 | Single depth image super resolution and denoising via coupled dictionary learning with local constraints and shock filteringabstractRecently, consumer depth cameras have gained significant popularity due to their affordable cost. However, the limited resolution and quality of the depth map generated by these cameras are still problems for several applications. In this paper, we propose a new algorithm for depth image super resolution using a single depth image as input. We reconstruct the corresponding high resolution depth map through a robust coupled dictionary learning algorithm with local coordinate constraints. The local constraints remove the prediction uncertainty and prevent the dictionary from over-fitting. We also incorporate an adaptively regularized Shock filter to simultaneously reduce the noise and sharpen the edges. Experimental results demonstrate the effectiveness of our proposed algorithm compared to previously reported methods. Cheng-Chuan Chou, Rogério Feris, Ming-Ting Sun |
ICME | 3 |
| 2014 | Temporal Non-maximum Suppression for Pedestrian Detection Using Self-CalibrationabstractIn this paper, we show how pedestrian detection accuracy and efficiency can be improved for static surveillance cameras using scene context and temporal non-maximal suppression. First, using the geometry of the scene, we derive the relationship between height in the image with respect to height in the real-world. This relationship is used to learn the range of scales to evaluate for potential detections. The geometry can also be used to estimate distances on the ground plane and to predict the image distance for pedestrian movement. Secondly, we show the error inherent in the standard non-maximum suppression method and demonstrate how the error can be reduced using a small temporal window. The scene context information and temporal non-maximum suppression (tNMS) can be applied to any detection algorithm. We evaluate the accuracy of each method across seven different videos. We show the scene context information and temporal non-maximum suppression improve accuracy on the order of ten and five percent respectively and over fifteen percent combined for seven different scenes for two publicly available detectors. Furthermore, both these approaches significantly reduce computational cost. Lisa M. Brown, Rogério Feris, Sharath Pankanti |
ICPR | 2 |
| 2014 | Appearance-Based Object Detection Under Varying Environmental ConditionsabstractPractical surveillance systems deployed in urban scenarios need to operate 24/7 under a wide range of environmental conditions. As modern video analytics shift from blob-based to object-centered architectures, appearance-based object detection under different weather conditions and lighting effects emerges as a critical yet largely unaddressed problem. This paper investigates this research topic, using as a case study the problem of vehicle detection in urban surveillance environments. In particular, we show that a simple and efficient Winsorized lighting correction technique improves performance significantly when outliers due to shadows, specularities, headlights, and occluders are present. Moreover, we demonstrate that a self-training mechanism utilizing a balanced training set automatically acquired from the target domain yields superior performance. Our experimental results are carried out on a novel dataset of vehicle images collected from a public traffic camera and categorized according to multiple environmental conditions. Rogério Feris, Lisa M. Brown, Sharath Pankanti, Ming-Ting Sun |
ICPR | 1 |
| 2014 | Attribute-based People Search: Lessons Learnt from a Practical Surveillance SystemabstractWe address the problem of attribute-based people search in real surveillance environments. The system we developed is capable of answering user queries such as "show me all people with a beard and sunglasses, wearing a white hat and a patterned blue shirt, from all metro cameras in the downtown area, from 2pm to 4pm last Saturday". In this paper, we describe the lessons we learned from practical deployments of our system, and how we made our algorithms achieve the accuracy and efficiency required by many police departments around the world. In particular, we show that a novel set of multimodal integral filters and proper normalization of attribute scores are critical to obtain good performance. We conduct a comprehensive experimental analysis on video footage captured from a large set of surveillance cameras monitoring metro chokepoints, in both crowded and normal activity periods. Moreover, we show impressive results using images from the recent Boston marathon bombing event, where our system can rapidly retrieve the two suspects based on their attributes from a database containing more than one thousand people present at the event. Rogério Feris, Russell Bobbitt, Lisa M. Brown, Sharath Pankanti |
ICMR | 1 |
| 2014 | Flower Classification for a Citizen Science Mobile AppabstractThis work describes an efficient approach for flower classification that is suitable for deployment in mobile devices, allowing its use in a citizen science application for biodiversity monitoring. In the proposed system, geo-located images are uploaded by the user and segmented semi-automatically. We propose a classification method based on histogram comparison of color, shape and texture cues, using metric learning for feature weighting. Our method is tested on the Oxford Flower Dataset and we are able to achieve state-of-the-art accuracy, while proposing an approach that can run efficiently in mobile devices. Andréa Britto Mattos, Ricardo Herrmann, Kelly Shigeno, Rogério Feris |
ICMR | 4 |
| 2014 | A spatial-color layout feature for representing galaxy imagesabstractWe propose a spatial-color layout feature specially designed for galaxy images. Inspired by findings on galaxy formation and evolution from Astronomy, the proposed feature captures both global and local morphological information of galaxies. In addition, our feature is scale and rotation invariant. By developing a hashing-based approach with the proposed feature, we implemented an efficient galaxy image retrieval system on a dataset with more than 280 thousand galaxy images from the Sloan Digital Sky Survey project. Given a query image, the proposed system can rank-order all galaxies from the dataset according to relevance in only 35 milliseconds on a single PC. To the best of our knowledge, this is one of the first works on galaxy-specific feature design and large-scale galaxy image retrieval. We evaluated the performance of the proposed feature and the galaxy image retrieval system using web user annotations, showing that the proposed feature outperforms other classic features, including HOG, Gist, LBP, and Color-histograms. The success of our retrieval system demonstrates the advantages of leveraging computer vision techniques in Astronomy problems. Yin Cui, Yongzhou Xiang, Kun Rong, Rogério Feris, Liangliang Cao |
WACV | 4 |
| 2013 | Efficient Maximum Appearance Search for Large-Scale Object DetectionabstractIn recent years, efficiency of large-scale object detection has arisen as an important topic due to the exponential growth in the size of benchmark object detection datasets. Most current object detection methods focus on improving accuracy of large-scale object detection with efficiency being an afterthought. In this paper, we present the Efficient Maximum Appearance Search (EMAS) model which is an order of magnitude faster than the existing state-of-the-art large-scale object detection approaches, while maintaining comparable accuracy. Our EMAS model consists of representing an image as an ensemble of densely sampled feature points with the proposed Point wise Fisher Vector encoding method, so that the learnt discriminative scoring function can be applied locally. Consequently, the object detection problem is transformed into searching an image sub-area for maximum local appearance probability, thereby making EMAS an order of magnitude faster than the traditional detection methods. In addition, the proposed model is also suitable for incorporating global context at a negligible extra computational cost. EMAS can also incorporate fusion of multiple features, which greatly improves its performance in detecting multiple object categories. Our experiments show that the proposed algorithm can perform detection of 1000 object classes in less than one minute per image on the Image Net ILSVRC2012 dataset and for 107 object classes in less than 5 seconds per image for the SUN09 dataset using a single CPU. Qiang Chen 0007, Rogério Feris, Ankur Datta, Liangliang Cao, ZhongYang Huang, Shuicheng Yan |
CVPR | 3 |
| 2013 | Designing Category-Level Attributes for Discriminative Visual RecognitionabstractAttribute-based representation has shown great promises for visual recognition due to its intuitive interpretation and cross-category generalization property. However, human efforts are usually involved in the attribute designing process, making the representation costly to obtain. In this paper, we propose a novel formulation to automatically design discriminative "category-level attributes", which can be efficiently encoded by a compact category-attribute matrix. The formulation allows us to achieve intuitive and critical design criteria (category-separability, learn ability) in a principled way. The designed attributes can be used for tasks of cross-category knowledge transfer, achieving superior performance over well-known attribute dataset Animals with Attributes (AwA) and a large-scale ILSVRC2010 dataset (1.2M images). This approach also leads to state-of-the-art performance on the zero-shot learning task on AwA. Felix X. Yu, Liangliang Cao, Rogério Feris, John R. Smith, Shih-Fu Chang |
CVPR | 3 |
| 2013 | Shape Analysis Using the Spectral Graph Wavelet TransformabstractThe present work describes a framework for morphological characterization of galaxies based on the Spectral Graph Wavelet Transform. A galaxy image is sampled with a number of points randomly chosen, whose Delaunay triangulation results in an arbitrary graph. The average intensity value in a 5 × 5 vicinity of a pixel related to a graph vertex is assigned to the corresponding graph vertex. A weight inversely proportional to the photometric distance between each pair of vertices is assigned to the respective graph edge. The Spectral Graph Wavelet Transform is computed from this weighted graph with real-valued vertices yielding a high-dimensional feature vector, which is reduced to a two dimensional vector through Principal Component Analysis. The proposed framework has been assessed through two case studies, namely, the case study of analyzing (i) 2D binary images from shapes and preliminary results of (ii) 2D gray tone images from galaxies. The obtained results imply the suitability of this framework for the characterization of galaxies images. Jorge J. G. Leandro, Roberto Marcondes Cesar Junior, Rogério Feris |
e-Science | 3 |
| 2013 | Fast Face Detector Training Using Tailored ViewsabstractFace detection is an important task in computer vision and often serves as the first step for a variety of applications. State-of-the-art approaches use efficient learning algorithms and train on large amounts of manually labeled imagery. Acquiring appropriate training images, however, is very time-consuming and does not guarantee that the collected training data is representative in terms of data variability. Moreover, available data sets are often acquired under controlled settings, restricting, for example, scene illumination or 3D head pose to a narrow range. This paper takes a look into the automated generation of adaptive training samples from a 3D morphable face model. Using statistical insights, the tailored training data guarantees full data variability and is enriched by arbitrary facial attributes such as age or body weight. Moreover, it can automatically adapt to environmental constraints, such as illumination or viewing angle of recorded video footage from surveillance cameras. We use the tailored imagery to train a new many-core implementation of Viola Jones' AdaBoost object detection framework. The new implementation is not only faster but also enables the use of multiple feature channels such as color features at training time. In our experiments we trained seven view-dependent face detectors and evaluate these on the Face Detection Data Set and Benchmark (FDDB). Our experiments show that the use of tailored training imagery outperforms state-of-the-art approaches on this challenging dataset. Kristina Scherbaum, James Petterson, Rogério Feris, Volker Blanz, Hans-Peter Seidel |
ICCV | 3 |
| 2013 | Fine registration of 3D point clouds with iterative closest point using an RGB-D cameraabstractWe address the problem of accurate and efficient alignment of 3D point clouds captured by an RGB-D (Kinect-style) camera from different viewpoints. Our approach introduces a new cost function for the iterative closest point (ICP) algorithm that balances the significance of structural and photometric features with dynamically adjusted weights to improve the error minimization process. We also enhance the algorithm with a novel outlier rejection method, which relies on adaptive thresholding at each ICP iteration, using both the structural information of the object and the spatial distances of sparse SIFT feature pairs. The effectiveness of our proposed approach is demonstrated in challenging scenarios, involving objects lacking structural features, and significant camera view and lighting changes. We obtained superior registration accuracy than existing related methods while requiring low computational processing. Yu-Feng Hsu, Rogério Feris, Ming-Ting Sun |
ISCAS | 3 |
| 2013 | Spatio-temporal fisher vector coding for surveillance event detectionabstractWe present a generic event detection system evaluated in the Surveillance Event Detection (SED) task of TRECVID 2012. We investigate a statistical approach with spatio-temporal features applied to seven event classes, which were defined by the SED task. This approach is based on local spatio-temporal descriptors, called MoSIFT and generated by pair-wise video frames. A Gaussian Mixture Model(GMM) is learned to model the distribution of the low level features. Then for each sliding window, the Fisher vector encoding [improvedFV] is used to generate the sample representation. The model is learnt using a Linear SVM for each event. The main novelty of our system is the introduction of Fisher vector encoding into video event detection. Fisher vector encoding has demonstrated great success in image classification. The key idea is to model the low level visual features as a Gaussian Mixture Model and to generate an intermediate vector representation for bag of features. FV encoding uses higher order statistics in place of histograms in the standard BoW. FV has several good properties: (a) it can naturally separate the video specific information from the noisy local features and (b) we can use a linear model for this representation. We build an efficient implementation for FV encoding which can attain a 10 times speed-up over real-time. We also take advantage of non-trivial object localization techniques to feed into the video event detection, e.g. multi-scale detection and non-maximum suppression. This approach outperformed the results of all other teams submissions in TRECVID SED 2012 on four of the seven event types. Qiang Chen 0007, Yang Cai 0002, Lisa M. Brown, Ankur Datta, Quanfu Fan, Rogério Feris, Shuicheng Yan, Alex Hauptmann 0001, Sharath Pankanti |
ACM Multimedia | 6 |
| 2013 | Boosting object detection performance in crowded surveillance videosabstractWe present a novel approach to automatically create efficient and accurate object detectors tailored to work well on specific video surveillance cameras (specific-domain detectors), using samples acquired with the help of a more expensive, general-domain detector (trained using images from multiple cameras). Our method requires no manual labels from the target domain. We automatically collect training data using tracking over short periods of time from high-confidence samples selected by the general-domain detector. In this context, a novel confidence measure is proposed for detectors based on a cascade of classifiers, which are frequently adopted for computer vision applications that require real-time processing. We demonstrate our proposed approach on the problem of vehicle detection in crowded surveillance videos, showing that an automatically generated detector significantly outperforms the original general-domain detector with much less feature computations. Rogério Feris, Ankur Datta, Sharath Pankanti, Ming-Ting Sun |
WACV | 1 |
| 2013 | Domain adaptive object detectionabstractWe study the use of domain adaptation and transfer learning techniques as part of a framework for adaptive object detection. Unlike recent applications of domain adaptation work in computer vision, which generally focus on image classification, we explore the problem of extreme class imbalance present when performing domain adaptation for object detection. The main difficulty caused by this imbalance is that test images contain millions or billions of negative image subwindows but just a few image subwindows containing positive instances, which makes it difficult to adapt to changes in the positive classes present new domains by simple techniques such as random sampling. We propose an initial approach to addressing this problem and apply our technique to vehicle detection in a challenging urban surveillance dataset, demonstrating the performance of our approach with various amounts of supervision, including the fully unsupervised case. Fatemeh Mirrashed, Vlad I. Morariu, Behjat Siddiquie, Rogério Feris, Larry Davis 0001 |
WACV | 4 |
| 2012 | Learning Detectors from Large Datasets for Object Retrieval in Video SurveillanceabstractWe address the problem of learning robust and efficient multi-view object detectors for surveillance video indexing and retrieval. Our philosophy is that effective solutions for this problem can be obtained by learning detectors from huge amounts of training data. Along this research direction, we propose a novel approach that consists of strategically partitioning the training set and learning a large array of complementary, compact, deep cascade detectors. At test time, given a video sequence captured by a fixed camera, a small number of detectors is automatically selected per image location. We demonstrate our approach on the problem of vehicle detection in challenging surveillance scenarios, using a large training dataset composed of around one million images. Our system runs at an impressive average rate of 125 frames per second on a conventional laptop computer. Rogério Feris, Sharath Pankanti, Behjat Siddiquie |
ICME | 1 |
| 2012 | Appearance modeling for person re-identification using Weighted Brightness Transfer Functions
Ankur Datta, Lisa M. Brown, Rogério Feris, Sharath Pankanti |
ICPR | 3 |
| 2012 | Unsupervised model selection for view-invariant object detection in surveillance environments
Behjat Siddiquie, Rogério Feris, Ankur Datta, Larry Davis 0001 |
ICPR | 2 |
| 2012 | Large-Scale Vehicle Detection, Indexing, and Search in Urban Surveillance VideosabstractWe present a novel approach for visual detection and attribute-based search of vehicles in crowded surveillance scenes. Large-scale processing is addressed along two dimensions: 1) large-scale indexing, where hundreds of billions of events need to be archived per month to enable effective search and 2) learning vehicle detectors with large-scale feature selection, using a feature pool containing millions of feature descriptors. Our method for vehicle detection also explicitly models occlusions and multiple vehicle types (e.g., buses, trucks, SUVs, cars), while requiring very few manual labeling. It runs quite efficiently at an average of 66 Hz on a conventional laptop computer. Once a vehicle is detected and tracked over the video, fine-grained attributes are extracted and ingested into a database to allow future search queries such as “Show me all blue trucks larger than 7 ft. length traveling at high speed northbound last Saturday, from 2 pm to 5 pm”. We perform a comprehensive quantitative analysis to validate our approach, showing its usefulness in realistic urban surveillance settings. Rogério Feris, Behjat Siddiquie, James Petterson, Yun Zhai, Ankur Datta, Lisa M. Brown, Sharath Pankanti |
IEEE Trans. Multim. | 1 |
| 2011 | Image ranking and retrieval based on multi-attribute queriesabstractWe propose a novel approach for ranking and retrieval of images based on multi-attribute queries. Existing image retrieval methods train separate classifiers for each word and heuristically combine their outputs for retrieving multiword queries. Moreover, these approaches also ignore the interdependencies among the query terms. In contrast, we propose a principled approach for multi-attribute retrieval which explicitly models the correlations that are present between the attributes. Given a multi-attribute query, we also utilize other attributes in the vocabulary which are not present in the query, for ranking/retrieval. Furthermore, we integrate ranking and retrieval within the same formulation, by posing them as structured prediction problems. Extensive experimental evaluation on the Labeled Faces in the Wild(LFW), FaceTracer and PASCAL VOC datasets show that our approach significantly outperforms several state-of-the-art ranking and retrieval methods. Behjat Siddiquie, Rogério Feris, Larry Davis 0001 |
CVPR | 2 |
| 2011 | Hierarchical ranking of facial attributesabstractWe propose a novel hierarchical structured prediction approach for ranking images of faces based on attributes. We view ranking as a bipartite graph matching problem; learning to rank under this setting can be achieved through structured prediction techniques that directly optimize the matching measures. Our key contribution is a novel model that combines structured predictors for different feature descriptors in a hierarchical fashion, enabling accurate ranking. We demonstrate our method on an important application which consists of searching for people over short intervals of time based on facial attributes. Given queries containing physical traits of a person (e.g., red hat, beard, and sunglasses), and an input database of face images, our system ranks the images in the database according to the query. Experiments show that our proposed hierarchical ranking approach poses significant enhancements in terms of accuracy over the non-hierarchical baseline. Ankur Datta, Rogério Feris, Daniel A. Vaquero |
FG | 2 |
| 2011 | Attribute-based vehicle search in crowded surveillance videosabstractWe present a novel application for searching for vehicles in surveillance videos based on semantic attributes. At the interface, the user specifies a set of vehicle characteristics (such as color, direction of travel, speed, length, height, etc.) and the system automatically retrieves video events that match the provided description. A key differentiating aspect of our system is the ability to handle challenging urban conditions such as high volumes of activity and environmental factors. This is achieved through a novel multi-view vehicle detection approach which relies on what we call motionlet classifiers, i.e. classifiers that are learned with vehicle samples clustered in the motion configuration space. We employ massively parallel feature selection to learn compact and accurate motionlet detectors. Moreover, in order to deal with different vehicle types (buses, trucks, SUVs, cars), we learn the motionlet detectors in a shape-free appearance space, where all training samples are resized to the same aspect ratio, and then during test time the aspect ratio of the sliding window is changed to allow the detection of different vehicle types. Once a vehicle is detected and tracked over the video, fine-grained attributes are extracted and ingested into a database to allow future search queries such as "Show me all blue trucks larger than 7ft length traveling at high speed northbound last Saturday, from 2pm to 5pm". Rogério Feris, Behjat Siddiquie, Yun Zhai, James Petterson, Lisa M. Brown, Sharath Pankanti |
ICMR | 1 |
| 2011 | Large-scale vehicle detection in challenging urban surveillance environmentsabstractWe present a novel approach for vehicle detection in urban surveillance videos, capable of handling unstructured and crowded environments with large occlusions, different vehicle shapes, and environmental conditions such as lighting changes, rain, shadows, and reflections. This is achieved with virtually no manual labeling efforts. The system runs quite efficiently at an average of 66Hz on a conventional laptop computer. Our proposed approach relies on three key contributions: (1) a co-training scheme where data is automatically captured based on motion and shape cues and used to train a detector based on appearance information; (2) an occlusion handling technique based on synthetically generated training samples obtained through Poisson image reconstruction from image gradients; (3) massively parallel feature selection over multiple feature planes which allows the final detector to be more accurate and more efficient. We perform a comprehensive quantitative analysis to validate our approach, showing its usefulness in realistic urban surveillance settings. Rogério Feris, James Petterson, Behjat Siddiquie, Lisa M. Brown, Sharath Pankanti |
WACV | 1 |
| 2011 | Robust Detection of Abandoned and Removed Objects in Complex Surveillance VideosabstractTracking-based approaches for abandoned object detection often become unreliable in complex surveillance videos due to occlusions, lighting changes, and other factors. We present a new framework to robustly and efficiently detect abandoned and removed objects based on background subtraction (BGS) and foreground analysis with complement of tracking to reduce false positives. In our system, the background is modeled by three Gaussian mixtures. In order to handle complex situations, several improvements are implemented for shadow removal, quick-lighting change adaptation, fragment reduction, and keeping a stable update rate for video streams with different frame rates. Then, the same Gaussian mixture models used for BGS are employed to detect static foreground regions without extra computation cost. Furthermore, the types of the static regions (abandoned or removed) are determined by using a method that exploits context information about the foreground masks, which significantly outperforms previous edge-based techniques. Based on the type of the static regions and user-defined parameters (e.g., object size and abandoned time), a matching method is proposed to detect abandoned and removed objects. A person-detection process is also integrated to distinguish static objects from stationary people. The robustness and efficiency of the proposed method is tested on IBM Smart Surveillance Solutions for public safety applications in big cities and evaluated by several public databases, such as The Image library for intelligent detection systems (i-LIDS) and IEEE Performance Evaluation of Tracking and Surveillance Workshop (PETS) 2006 datasets. The test and evaluation demonstrate our method is efficient to run in real-time, while being robust to quick-lighting changes and occlusions in complex environments. Yingli Tian, Rogério Feris, Arun Hampapur, Ming-Ting Sun |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2010 | Unsupervised action classification using space-time link analysisabstractIn this paper we address the problem of unsupervised discovery of action classes in video data. Different from all existing methods thus far proposed for this task, we present a space-time link analysis approach which matches the performance of traditional unsupervised action categorization methods in a standard dataset. Our method is inspired by the recent success of link analysis techniques in the image domain. By applying these techniques in the space-time domain, we are able to naturally take into account the spatio-temporal relationships between the video features, while leveraging the power of graph matching for action classification. We present an experiment to demonstrate that our approach is capable of handling cluttered backgrounds, activities with subtle movements, and video data from moving cameras. Rogério Feris, Volker Krüger, Ming-Ting Sun |
ISCAS | 2 |
| 2009 | Video Analytics in Urban EnvironmentsabstractUrban environments present unique challenges from the perspective of surveillance and security. Threat activity in urban environments tends to be very similar to background activity, while the volume of activity is often very high. The widespread geographical area presents issues from the perspective of response. These characteristics of urban environments create challenges to traditional applications of video analytics technologies and opens up opportunities for novel approaches. This paper explores the applicability of video analytics in various scenarios presented in urban surveillance situations. We also describe novel technical solutions to some of the challenges of urban surveillance. Arun Hampapur, Russell Bobbitt, Lisa M. Brown, Mike Desimone, Rogério Feris, Rick Kjeldsen, Max Lu, Carl Mercier, Chris Milite, Stephen Russo, Chiao-Fe Shu, Yun Zhai |
AVSS | 5 |
| 2009 | Shape classification through structured learning of matching measuresabstractMany traditional methods for shape classification involve establishing point correspondences between shapes to produce matching scores, which are in turn used as similarity measures for classification. Learning techniques have been applied only in the second stage of this process, after the matching scores have been obtained. In this paper, instead of simply taking for granted the scores obtained by matching and then learning a classifier, we learn the matching scores themselves so as to produce shape similarity scores that minimize the classification loss. The solution is based on a max-margin formulation in the structured prediction setting. Experiments in shape databases reveal that such an integrated learning algorithm substantially improves on existing methods. Longbin Chen, Julian J. McAuley, Rogério Feris, Tibério S. Caetano, Matthew Turk 0001 |
CVPR | 3 |
| 2009 | A projector-camera setup for geometry-invariant frequency demultiplexingabstractConsider a projector-camera setup where a sinusoidal pattern is projected onto the scene, and an image of the objects imprinted with the pattern is captured by the camera. In this configuration, the local frequency of the sinusoidal pattern as seen by the camera is a function of both the frequency of the projected sinusoid and the local geometry of objects in the scene. We observe that, by strategically placing the projector and the camera in canonical configuration and projecting sinusoidal patterns aligned with the epipolar lines, the frequency of the sinusoids seen in the image becomes invariant to the local object geometry. This property allows us to design systems composed of a camera and multiple projectors, which can be used to capture a single image of a scene illuminated by all projectors at the same time, and then demultiplex the frequencies generated by each individual projector separately. We show how imaging systems like those can be used to segment, from a single image, the shadows cast by each individual projector - an application that we call coded shadow photography. The method is useful to extend the applicability of techniques that rely on the analysis of shadows cast by multiple light sources placed at different positions, as the individual shadows captured at distinct instants of time now can be obtained from a single shot, enabling the processing of dynamic scenes. Daniel A. Vaquero, Ramesh Raskar, Rogério Feris, Matthew Turk 0001 |
CVPR | 3 |
| 2009 | Attribute-based people search in surveillance environmentsabstractWe propose a novel framework for searching for people in surveillance environments. Rather than relying on face recognition technology, which is known to be sensitive to typical surveillance conditions such as lighting changes, face pose variation, and low-resolution imagery, we approach the problem in a different way: we search for people based on a parsing of human parts and their attributes, including facial hair, eyewear, clothing color, etc. These attributes can be extracted using detectors learned from large amounts of training data. A complete system that implements our framework is presented. At the interface, the user can specify a set of personal characteristics, and the system then retrieves events that match the provided description. For example, a possible query is ¿show me the bald people who entered a given building last Saturday wearing a red shirt and sunglasses.¿ This capability is useful in several applications, such as finding suspects or missing people. To evaluate the performance of our approach, we present extensive experiments on a set of images collected from the Internet, on infrared imagery, and on two-and-a-half months of video from a real surveillance environment. We are not aware of any similar surveillance system capable of automatically finding people in video based on their fine-grained body parts and attributes. Daniel A. Vaquero, Rogério Feris, Duan Tran, Lisa M. Brown, Arun Hampapur, Matthew Turk 0001 |
WACV | 2 |
| 2008 | An Integrated System for Moving Object Classification in Surveillance VideosabstractMoving object classification in far-field video is a key component of smart surveillance systems. In this paper, we propose a reliable system for person-vehicle classification which works well in challenging real-word conditions, including the presence of shadows, low resolution imagery, perspective distortions, arbitrary camera viewpoints, and groups of people. Our system runsin real-time (30 Hz) on conventional machines and has low memory consumption. We achieved accurate results by relying on powerful discriminative features, including a novel measure of object deformation based on differences of histograms of oriented gradients. We also provide an interactive user interface, enabling users to specify regions of interest for each class and correct for perspective distortions by specifying different sizes indifferent positions of the camera view. Finally, we use anautomatic adaptation process to continuously update the parameters of the system so that its performance increases for a particular environment. Experimental results demonstrate the effectiveness of our system in standard dataset and a variety of video clips captured with our surveillance cameras. Longbin Chen, Rogério Feris, Yun Zhai, Lisa M. Brown, Arun Hampapur |
AVSS | 2 |
| 2008 | Characterizing the shadow space of camera-light pairsabstractWe present a theoretical analysis for characterizing the shadows cast by a point light source given its relative position to the camera. In particular, we analyze the epipolar geometry of camera-light pairs, including unusual camera-light configurations such as light sources aligned with the camerapsilas optical axis as well as convenient arrangements such as lights placed in the camera plane. A mathematical characterization of the shadows is derived to determine the orientations and locations of depth discontinuities when projected onto the image plane that could potentially be associated with cast shadows. The resulting theory is applied to compute a lower bound on the number of lights needed to extract all depth discontinuities from a general scene using a multiflash camera. We also provide a characterization of which discontinuities are missed and which are correctly detected by the algorithm, and a foundation for choosing an optimal light placement. Experiments with depth edges computed using two-flash setups and a four-flash setup illustrate the theory, and an additional configuration with a flash at the camerapsilas center of projection is exploited as a solution for some degenerate cases. Daniel A. Vaquero, Rogério Feris, Matthew Turk 0001, Ramesh Raskar |
CVPR | 2 |
| 2008 | Facial image analysis using local feature adaptation prior to learningabstractMany facial image analysis methods rely on learning-based techniques such as Adaboost or SVMs to project classifiers based on the selection of local image filters (e.g., Haar and Gabor filters) from large sets of training data. In general, the learning process consists of selecting discriminative image filters from a large feature pool that contains filters uniformly sampled from the parameter space. In this paper, we argue that we are able to improve these methods by incorporating a local feature adaptation technique prior to learning, which generates a more compact and meaningful pool of image filters, consequently reducing both learning and detection/recognition computational costs, while at the same time improving accuracies. In the first stage of our approach, local feature adaptation is carried out by a nonlinear optimization method that determines image filter parameters (such as position, orientation and scale) in order to match the geometrical structure of each training sample. In the second stage, Adaboost feature selection technique is applied to the adapted feature pool to obtain the final set of discriminative local image filters. We demonstrate the effectiveness and efficiency of the proposed framework in the face detection domain. In the experiments, we have applied our method using a pool of wavelet features, including Haar and Gabor filters. The results showed that with local feature adaptation, significant improvements in terms of detection accuracy and computational cost reduction are achieved over learning based on the same features sampled uniformly from the parameter space. Rogério Feris, Yingli Tian, Yun Zhai, Arun Hampapur |
FG | 1 |
| 2008 | Multiflash Stereopsis: Depth-Edge-Preserving Stereo with Small Baseline IlluminationabstractTraditional stereo matching algorithms are limited in their ability to produce accurate results near depth discontinuities, due to partial occlusions and violation of smoothness constraints. In this paper, we use small baseline multi-flash illumination to produce a rich set of feature maps that enable acquisition of discontinuity preserving point correspondences. First, from a single multi-flash camera, we formulate a qualitative depth map using a gradient domain method that encodes object relative distances. Then, in a multiview setup, we exploit shadows created by light sources to compute an occlusion map. Finally, we demonstrate the usefulness of these feature maps by incorporating them into two different dense stereo correspondence algorithms, the first based on local search and the second based on belief propagation. Experimental results show that our enhanced stereo algorithms are able to extract high quality, discontinuity preserving correspondence maps from scenes that are extremely challenging for conventional stereo methods. We also demonstrate that small baseline illumination can be useful to handle specular reflections in stereo imagery. Different from most existing active illumination techniques, our method is simple, inexpensive, compact, and requires no calibration of light sources. Rogério Feris, Ramesh Raskar, Longbin Chen, Kar-Han Tan, Matthew Turk 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Searching surveillance videoabstractSurveillance video is used in two key modes, watching for known threats in real-time and searching for events of interest after the fact. Typically, real-time alerting is a localized function, e.g. airport security center receives and reacts to a “perimeter breach alert”, while investigations often tend to encompass a large number of geographically distributed cameras like the London bombing, or Washington sniper incidents. Enabling effective search of surveillance video for investigation & preemption, involves indexing the video along multiple dimensions. This paper presents a framework for surveillance search which includes, video parsing, indexing and query mechanisms. It explores video parsing techniques which automatically extract index data from video, indexing which stores data in relational tables, retrieval which uses SQL queries to retrieve events of interest and the software architecture that integrates these technologies. Arun Hampapur, Lisa M. Brown, Rogério Feris, Andrew W. Senior, Chiao-Fe Shu, Yingli Tian, Yun Zhai, Max Lu |
AVSS | 3 |
| 2007 | Video analytics for retailabstractWe describe a set of tools for retail analytics based on a combination of video understanding and transaction-log. Tools are provided for loss prevention (returns fraud and cashier fraud), store operations (customer counting) and merchandising (display effectiveness). Results are presented on returns fraud and customer counting. Andrew W. Senior, Lisa M. Brown, Arun Hampapur, Chiao-Fe Shu, Yun Zhai, Rogério Feris, Yingli Tian, Sergio Borger, Christopher R. Carlson |
AVSS | 6 |
| 2007 | Capturing People in Surveillance VideoabstractThis paper presents reliable techniques for detecting, tracking, and storing keyframes of people in surveillance video. The first component of our system is a novel face detector algorithm, which is based on first learning local adaptive features for each training image, and then using Adaboost learning to select the most general features for detection. This method provides a powerful mechanism for combining multiple features, allowing faster training time and better detection rates. The second component is a face tracking algorithm that interleaves multiple view-based classifiers along the temporal domain in a video sequence. This interleaving technique, combined with a correlation-based tracker, enables fast and robust face tracking over time. Finally, the third component of our system is a keyframe selection method that combines a person classifier with a face classifier. The basic idea is to generate a person keyframe in case the face is not visible, in order to reduce the number of false negatives. We performed quantitatively evaluation of our techniques on standard datasets and on surveillance videos captured by a camera over several days. Rogério Feris, Yingli Tian, Arun Hampapur |
CVPR | 1 |
| 2006 | Manifold based analysis of facial expression
Ya Chang, Changbo Hu, Rogério Feris, Matthew Turk 0001 |
Image Vis. Comput. | 3 |
| 2006 | Local approach for face verification in polar frequency domain
Yossi Zana, Roberto Marcondes Cesar Junior, Rogério Feris, Matthew Turk 0001 |
Image Vis. Comput. | 3 |
| 2005 | Discontinuity Preserving Stereo with Small Baseline Multi-Flash IlluminationabstractCurrently, sharp discontinuities in depth and partial occlusions in multiview imaging systems pose serious challenges for many dense correspondence algorithms. However, it is important for 3D reconstruction methods to preserve depth edges as they correspond to important shape features like silhouettes which are critical for understanding the structure of a scene. In this paper, we show how active illumination algorithms can produce a rich set of feature maps that are useful in dense 3D reconstruction. We start by showing a method to compute a qualitative depth map from a single camera, which encodes object relative distances and can be used as a prior for stereo. In a multiview setup, we show that along with depth edges, binocular half-occluded pixels can also be explicitly and reliably labeled. To demonstrate the usefulness of these feature maps, we show how they can be used in two different algorithms for dense stereo correspondence. Our experimental results show that our enhanced stereo algorithms are able to extract high quality, discontinuity preserving correspondence maps from scenes that are extremely challenging for conventional stereo methods. Rogério Feris, Ramesh Raskar, Longbin Chen, Kar-Han Tan, Matthew Turk 0001 |
ICCV | 1 |
| 2004 | Shape-Enhanced Surgical Visualizations and Medical Illustrations with Multi-flash Imaging
Kar-Han Tan, James Kobler, Paul H. Dietz, Ramesh Raskar, Rogério Feris |
MICCAI (2) | 5 |
| 2004 | Non-photorealistic camera: depth edge detection and stylized rendering using multi-flash imagingabstractWe present a non-photorealistic rendering approach to capture and convey shape features of real-world scenes. We use a camera with multiple flashes that are strategically positioned to cast shadows along depth discontinuities in the scene. The projective-geometric relationship of the camera-flash setup is then exploited to detect depth discontinuities and distinguish them from intensity edges due to material discontinuities.We introduce depiction methods that utilize the detected edge features to generate stylized static and animated images. We can highlight the detected features, suppress unnecessary details or combine features from multiple images. The resulting images more clearly convey the 3D structure of the imaged scenes.We take a very different approach to capturing geometric features of a scene than traditional approaches that require reconstructing a 3D model. This results in a method that is both surprisingly simple and computationally efficient. The entire hardware/software setup can conceivably be packaged into a self-contained device no larger than existing digital cameras. Ramesh Raskar, Kar-Han Tan, Rogério Feris, Jingyi Yu 0001, Matthew Turk 0001 |
ACM Trans. Graph. | 3 |
| 2003 | Active Wavelet Networks for Face AlignmentabstractThe active appearance model (AAM) algorithm has proved to be a successful method for face alignment and synthesis. By elegantly combining both shape and texture models, AAM allows fast and robust deformable image matching. However, the method is sensitive to partial occlusions and illumination changes. In such cases, the PCA-based texture model causes the reconstruction error to be globally spread over the image. In this paper, we propose a new method for face alignment called active wavelet networks (AWN), which replaces the AAM texture model by a wavelet network representation. Since we consider spatially localized wavelets for modeling texture, our method shows more robustness against partial occlusions and some illumination changes. 1 Changbo Hu, Rogério Feris, Matthew Turk 0001 |
BMVC | 2 |