VLDB 2026 Research / reviewers in the wild / expert
Yogesh S. Rawat
dblp:148/2258 · also Yogesh Singh Rawat
· DBLP profile ↗
54ranked-venue papers
7as first author
39since 2021 · last 2026
0000-0003-4052-6798ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 7 first-author · 26 since 2021Artificial intelligence and machine learning · 36 · 29 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ACuRE: Accurate Continuity-Regularized SpO2 Estimation Using Liquid Time-Constant NetworksabstractBlood oxygen saturation (SpO2) is a vital measure of respiratory and circulatory health, essential for detecting hypoxemia in conditions like chronic obstructive pulmonary disease and heart failure. Current non-contact SpO2estimation methods using remote photoplethysmography (rPPG) struggle with motion artifacts, illumination variability, and limited temporal modeling, hindering their practical use. We propose ACuRE, a novel framework that integrates a two-branch 3D-ResNet-18 for AC/DC signal separation, Liquid Time-Constant (LTC) networks for continuous-time dynamics, and a physics-informed partial differential equation (PDE) loss based on mass conservation. ACuRE overcomes these challenges by isolating pulsatile (AC) and baseline (DC) signals for enhanced robustness, using LTC networks to capture nonlinear physiological dynamics, and applying PDE regularization to ensure signal continuity. This achieves a significant reduction in mean absolute error compared to baselines, with strong performance under motion and illumination stress. Evaluated across multiple datasets, ACuRE demonstrates robust accuracy and generalization, offering a scalable solution for video-based health monitoring in telemedicine and low-resource settings. Code available at: https://github.com/Shahzadnit/ACURE_WACV Shahzad Ahmad 0002, Divya Mishra, Sania Bano, Sukalpa Chanda, Yogesh S. Rawat |
WACV | 5 |
| 2026 | RobustGait: Robustness Analysis for Appearance Based Gait RecognitionabstractAppearance-based gait recognition has achieved strong performance on controlled datasets, yet systematic evaluation of its robustness to real-world corruptions and silhouette variability remains lacking. We present RobustGait, a framework for fine-grained robustness evaluation of appearance-based gait recognition systems. RobustGait evaluation spans four dimensions: the type of perturbation (digital, environmental, temporal, occlusion), the silhouette extraction method (segmentation and parsing networks), the architectural capacities of gait recognition models, and various deployment scenarios. The benchmark introduces 15 corruption types at 5 severity levels across CASIA-B, CCPG, and SUSTech1K, with in-the-wild validation on MEVID, and evaluates six state-of-the-art gait systems. We came across several exciting insights. First, applying noise at the RGB level better reflects real-world degradation, and reveals how distortions propagate through silhouette extraction to the downstream gait recognition systems. Second, gait accuracy is highly sensitive to silhouette extractor biases, revealing an overlooked source of benchmark bias. Third, robustness is dependent on both the type of perturbation and the architectural design. Finally, we explore robustness-enhancing strategies, showing that noise-aware training and knowledge distillation improve performance and move toward deployment-ready systems. Reeshoon Sayera, Akash Kumar 0016, Sirshapan Mitra, Prudvi Kamtam, Yogesh S. Rawat |
WACV | 5 |
| 2025 | Stable Mean Teacher for Semi-supervised Video Action DetectionabstractIn this work, we focus on semi-supervised learning for video action detection. Video action detection requires spatio-temporal localization in addition to classification, and a limited amount of labels makes the model prone to unreliable predictions. We present Stable Mean Teacher, a simple end-to-end student-teacher-based framework that benefits from improved and temporally consistent pseudo labels. It relies on a novel ErrOr Recovery (EoR) module, which learns from students' mistakes on labeled samples and transfers this to the teacher to improve pseudo labels for unlabeled samples. Moreover, existing spatio-temporal losses do not take temporal coherency into account and are prone to temporal inconsistencies. To overcome this, we present Difference of Pixels (DoP), a simple and novel constraint focused on temporal consistency, which leads to coherent temporal detections. We evaluate our approach on four different spatio-temporal detection benchmarks: UCF101-24, JHMDB21, AVA, and Youtube-VOS. Our approach outperforms the supervised baselines for action detection by an average margin of 23.5% on UCF101-24, 16% on JHMDB21, and 3.3% on AVA. Using merely 10% and 20% of data, it provides a competitive performance compared to the supervised baseline trained on 100% annotations on UCF101-24 and JHMDB21 respectively. We further evaluate its effectiveness on AVA for scaling to large-scale datasets and Youtube-VOS for video object segmentation, demonstrating its generalization capability to other tasks in the video domain. Akash Kumar 0016, Sirshapan Mitra, Yogesh S. Rawat |
AAAI | 3 |
| 2025 | HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video UnderstandingabstractDespite advancements in multimodal large language models (MLLMs), current approaches struggle in medium-to-long video understanding due to frame and context length limitations. As a result, these models often depend on frame sampling, which risks missing key information over time and lacks task-specific relevance. To address these challenges, we introduce HierarQ, a task-aware hierarchical Q-Former based framework that sequentially processes frames to bypass the need for frame sampling, while avoiding LLM’s context length limitations. We introduce a lightweight two-stream language-guided feature modulator to incorporate task awareness in video understanding, with the entity stream capturing frame-level object information within a short context and the scene stream identifying their broader interactions over longer period of time. Each stream is supported by dedicated memory banks which enables our proposed Hierarchical Querying transformer (HierarQ) to effectively capture short and long-term context. Extensive evaluations on 10 video benchmarks across video understanding, question answering, and captioning tasks demonstrate HierarQ’s state-of-the-art performance across most datasets, proving its robustness and efficiency for comprehensive video analysis. Shehreen Azad, Vibhav Vineet, Yogesh S. Rawat |
CVPR | 3 |
| 2025 | STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal GroundingabstractIn this work we study Weakly Supervised Spatio-Temporal Video Grounding (WSTVG), a challenging task of localizing subjects spatio-temporally in videos using only textual queries and no bounding box supervision. Inspired by recent advances in vision-language foundation models, we investigate their utility for WSTVG, leveraging their zero-shot grounding capabilities. However, we find that a simple adaptation lacks essential spatio-temporal grounding abilities. To bridge this gap, we introduce Tubelet Referral Grounding (TRG), which connects textual queries to tubelets to enable spatio-temporal predictions. Despite its promise, TRG struggles with compositional action understanding and dense scene scenarios. To address these limitations, we propose STPro, a novel progressive learning framework with two key modules: (1) Sub-Action Temporal Curriculum Learning (SA-TCL), which incrementally builds compositional action understanding, and (2) Congestion-Guided Spatial Curriculum Learning (CG-SCL), which adapts the model to complex scenes by spatially increasing task difficulty. STPro achieves state-of-the-art results on three benchmark datasets, with improvements of 1.0% on VidSTG-Declarative and 3.0% on HCSTVG-v1. Aaryan Garg, Akash Kumar 0016, Yogesh S. Rawat |
CVPR | 3 |
| 2025 | DIFFER: Disentangling Identity Features via Semantic Cues for Clothes-Changing Person Re-IDabstractClothes-changing person re-identification (CC-ReID) aims to recognize individuals under different clothing scenarios. Current CC-ReID approaches either concentrate on modeling body shape using additional modalities including silhouette, pose, and body mesh, potentially causing the model to overlook other critical biometric traits such as gender, age, and style, or they incorporate supervision through additional labels that the model tries to disregard or emphasize, such as clothing or personal attributes. However, these annotations are discrete in nature and do not capture comprehensive descriptions.In this work, we propose DIFFER: Disentangle Identity Features From Entangled Representations, a novel adversarial learning method that leverages textual descriptions to disentangle identity features. Recognizing that image features inherently mix inseparable information, DIFFER introduces NBDetach, a mechanism designed for feature disentanglement by leveraging the separable nature of text descriptions as supervision. It partitions the feature space into distinct subspaces and, through gradient reversal layers, effectively separates identity-related features from non-biometric features. We evaluate DIFFER on 4 different benchmark datasets (LTCC, PRCC, CelebreID-Light, and CCVID) to demonstrate its effectiveness and provide state-of-the-art performance across all the benchmarks. DIFFER consistently outperforms the baseline method, with improvements in top-1 accuracy of 3.6% on LTCC, 3.4% on PRCC, 2.5% on CelebReID-Light, and 1% on CCVID. Our code can be found here. Yogesh S. Rawat |
CVPR | 2 |
| 2025 | Punching Bag vs. Punching Person: Motion Transferability in VideosabstractAction recognition models demonstrate strong generalization, but can they effectively transfer high-level motion concepts across diverse contexts, even within similar distributions? For example, can a model recognize the broad action "punching" when presented with an unseen variation such as "punching person"? To explore this, we introduce a motion transferability framework with three datasets: (1) Syn-TA, a synthetic dataset with 3D object motions; (2) Kinetics400-TA; and (3) Something-Something-v2-TA, both adapted from natural video datasets. We evaluate 13 state-of-the-art models on these benchmarks and observe a significant drop in performance when recognizing high-level actions in novel contexts. Our analysis reveals: 1) Multimodal models struggle more with fine-grained unknown actions than with coarse ones; 2) The bias-free Syn-TA proves as challenging as real-world datasets, with models showing greater performance drops in controlled settings; 3) Larger models improve transferability when spatial cues dominate but struggle with intensive temporal reasoning, while reliance on object and background cues hinders generalization. We further explore how disentangling coarse and fine motions can improve recognition in temporally challenging datasets. We believe this study establishes a crucial benchmark for assessing motion transferability in action recognition. Datasets and relevant code: https://github.com/raiyaan-abdullah/Motion-Transfer. Raiyaan Abdullah, Jared Claypoole, Michael Cogswell, Ajay Divakaran, Yogesh S. Rawat |
ICCV | 5 |
| 2025 | DisenQ: Disentangling Q-Former for Activity-BiometricsabstractIn this work, we address activity-biometrics, which involves identifying individuals across diverse set of activities. Unlike traditional person identification, this setting introduces additional challenges as identity cues become entangled with motion dynamics and appearance variations, making biometrics feature learning more complex. While additional visual data like pose and/or silhouette help, they often struggle from extraction inaccuracies. To overcome this, we propose a multimodal language-guided framework that replaces reliance on additional visual data with structured textual supervision. At its core, we introduce \textbf{DisenQ} (\textbf{Disen}tangling \textbf{Q}-Former), a unified querying transformer that disentangles biometrics, motion, and non-biometrics features by leveraging structured language guidance. This ensures identity cues remain independent of appearance and motion variations, preventing misidentifications. We evaluate our approach on three activity-based video benchmarks, achieving state-of-the-art performance. Additionally, we demonstrate strong generalization to complex real-world scenario with competitive performance on a traditional video-based identification benchmark, showing the effectiveness of our framework. Shehreen Azad, Yogesh S. Rawat |
ICCV | 2 |
| 2025 | Colors See Colors Ignore: Clothes Changing ReID with Color DisentanglementabstractClothes-Changing Re-Identification (CC-ReID) aims to recognize individuals across different locations and times, irrespective of clothing. Existing methods often rely on additional models or annotations to learn robust, clothing-invariant features, making them resource-intensive. In contrast, we explore the use of color - specifically foreground and background colors - as a lightweight, annotation-free proxy for mitigating appearance bias in ReID models. We propose Colors See, Colors Ignore (CSCI), an RGB-only method that leverages color information directly from raw images or video frames. CSCI efficiently captures color-related appearance bias ('Color See') while disentangling it from identity-relevant ReID features ('Color Ignore'). To achieve this, we introduce S2A self-attention, a novel self-attention to prevent information leak between color and identity cues within the feature space. Our analysis shows a strong correspondence between learned color embeddings and clothing attributes, validating color as an effective proxy when explicit clothing labels are unavailable. We demonstrate the effectiveness of CSCI on both image and video ReID with extensive experiments on four CC-ReID datasets. We improve the baseline by Top-1 2.9% on LTCC and 5.0% on PRCC for image-based ReID, and 1.0% on CCVID and 2.5% on MeVID for video-based ReID without relying on additional supervision. Our results highlight the potential of color as a cost-effective solution for addressing appearance bias in CC-ReID. Github: https://github.com/ppriyank/ICCV-CSCI-Person-ReID. Priyank Pathak, Yogesh S. Rawat |
ICCV | 2 |
| 2025 | Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video GroundingabstractIn this work, we focus on Weakly Supervised Spatio-Temporal Video Grounding (WSTVG). It is a multimodal task aimed at localizing specific subjects spatio-temporally based on textual queries without bounding box supervision. Motivated by recent advancements in multi-modal foundation models for grounding tasks, we first explore the potential of state-of-the-art object detection models for WSTVG. Despite their robust zero-shot capabilities, our adaptation reveals significant limitations, including inconsistent temporal predictions, inadequate understanding of complex queries, and challenges in adapting to difficult scenarios. We propose CoSPaL (Contextual Self-Paced Learning), a novel approach which is designed to overcome these limitations. CoSPaL integrates three core components: (1) Tubelet Phrase Grounding (TPG), which introduces spatio-temporal prediction by linking textual queries to tubelets; (2) Contextual Referral Grounding (CRG), which improves comprehension of complex queries by extracting contextual information to refine object identification over time; and (3) Self-Paced Scene Understanding (SPS), a training paradigm that progressively increases task difficulty, enabling the model to adapt to complex scenarios by transitioning from coarse to fine-grained understanding. Akash Kumar 0016, Zsolt Kira, Yogesh S. Rawat |
ICLR | 3 |
| 2025 | Lr0.Fm: low-Resolution Zero-Shot Classification Benchmark for Foundation Models
Priyank Pathak, Shyam Marjit, Shruti Vyas, Yogesh S. Rawat |
ICLR | 4 |
| 2025 | MolVision: Molecular Property Prediction with Vision Language ModelsabstractMolecular property prediction is a fundamental task in computational chemistry with critical applications in drug discovery and materials science. While recent works have explored Large Language Models (LLMs) for this task, they primarily rely on textual molecular representations such as SMILES/SELFIES, which can be ambiguous and structurally uninformative. In this work, we introduce MolVision, a novel approach that leverages Vision-Language Models (VLMs) by integrating both molecular structure images and textual descriptions to enhance property prediction. We construct a benchmark spanning nine diverse datasets, covering both classification and regression tasks. Evaluating nine different VLMs in zero-shot, few-shot, and fine-tuned settings, we find that visual information improves prediction performance, particularly when combined with efficient fine-tuning strategies such as LoRA. Our results reveal that while visual information alone is insufficient, multimodal fusion significantly enhances generalization across molecular properties. Adaptation of vision encoder for molecular images in conjunction with LoRA further improves the performance. The code and data is available at : https://molvision.github.io/MolVision/. Deepan Adak, Yogesh S. Rawat, Shruti Vyas |
NeurIPS | 2 |
| 2025 | PULSE: Physiological Understanding with Liquid Signal ExtractionabstractThe non-contact estimation of vital signs, particularly heart rate, from video data is a promising method for remote health monitoring. 3D convolutional layers are widely used for this task due to their ability to capture both spatial and temporal features. However, traditional 3D convolutions, while effective in many cases, lack the capacity to adjust dy-namically to the temporal variability inherent in physiological signals such as remote photoplethysmography (rPPG), which are characterized by subtle frequency changes over time. To address this, we propose PULSE (Physiological Understanding with Liquid Signal Extraction), a frame-work that employs Liquid Time-Constant (LTC) models with 3D convolutional layers to enhance temporal sensitivity and improve the extraction of these fine-grained rPPG signals. In PULSE, traditional 3D-conv layers are deployed for ini-tial feature extraction, while LTC-based 3D-conv layers dy-namically adapt and guide the temporal processing, allowing the model to better track and interpret the subtle variations in heart rate signals under different conditions, such as motion artifacts and lighting changes. We evaluated the effectiveness of PULSE in an unsupervised training setting, demonstrating that our solution performs well even in the absence of labeled datasets a common challenge in rPPG signal extraction. Experimental evaluations on three public datasets confirm that PULSE achieves comparable or supe-rior results to existing methods, proving its robustness and efficacy for real-world, non-contact health monitoring applications. Shahzad Ahmad 0002, Sania Bano, Sachin Verma, Yogesh S. Rawat, Sukalpa Chanda, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
WACV | 4 |
| 2025 | Advancing automatic photovoltaic defect detection using semi-supervised semantic segmentation of electroluminescence images
Yogesh S. Rawat, Shruti Vyas |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Semi-supervised Active Learning for Video Action DetectionabstractIn this work, we focus on label efficient learning for video action detection. We develop a novel semi-supervised active learning approach which utilizes both labeled as well as un- labeled data along with informative sample selection for ac- tion detection. Video action detection requires spatio-temporal localization along with classification, which poses several challenges for both active learning (informative sample se- lection) as well as semi-supervised learning (pseudo label generation). First, we propose NoiseAug, a simple augmenta- tion strategy which effectively selects informative samples for video action detection. Next, we propose fft-attention, a novel technique based on high-pass filtering which enables effective utilization of pseudo label for SSL in video action detection by emphasizing on relevant activity region within a video. We evaluate the proposed approach on three different bench- mark datasets, UCF-101-24, JHMDB-21, and Youtube-VOS. First, we demonstrate its effectiveness on video action detec- tion where the proposed approach outperforms prior works in semi-supervised and weakly-supervised learning along with several baseline approaches in both UCF101-24 and JHMDB- 21. Next, we also show its effectiveness on Youtube-VOS for video object segmentation demonstrating its generalization capability for other dense prediction tasks in videos. Ayush Singh, Aayush Jung Rana, Akash Kumar 0016, Shruti Vyas, Yogesh S. Rawat |
AAAI | 5 |
| 2024 | Activity-Biometrics: Person Identification from Daily ActivitiesabstractIn this work, we study a novel problem which focuses on person identification while performing daily activities. Learning biometric features from RGB videos is challenging due to spatio-temporal complexity and presence of ap-pearance biases such as clothing color and background. We propose ABNet, a novel framework which leverages dis-entanglement of biometric and non-biometric features to perform effective person identification from daily activities. ABNet relies on a bias-less teacher to learn biometric features from RGB videos and explicitly disentangle non-biometric features with the help of biometric distortion. In addition, ABNet also exploits activity prior for biometrics which is enabled by joint biometric and activity learning. We perform comprehensive evaluation of the proposed approach across five different datasets which are derived from existing activity recognition benchmarks. Furthermore, we extensively compare ABNet with existing works in person identification and demonstrate its effectiveness for activity-based biometrics across all five datasets. The code and dataset can be accessed at: https: //github.com/ sacrcv/Activity-Biometrics/ Shehreen Azad, Yogesh S. Rawat |
CVPR | 2 |
| 2024 | AirSketch: Generative Motion to SketchabstractIllustration is a fundamental mode of human expression and communication. Certain types of motion that accompany speech can provide this illustrative mode of communication. While Augmented and Virtual Reality technologies (AR/VR) have introduced tools for producing drawings with hand motions (air drawing), they typically require costly hardware and additional digital markers, thereby limiting their accessibility and portability. Furthermore, air drawing demands considerable skill to achieve aesthetic results. To address these challenges, we introduce the concept of AirSketch, aimed at generating faithful and visually coherent sketches directly from hand motions, eliminating the need for complicated headsets or markers. We devise a simple augmentation-based self-supervised training procedure, enabling a controllable image diffusion model to learn to translate from highly noisy hand tracking images to clean, aesthetically pleasing sketches, while preserving the essential visual cues from the original tracking data. We present two air drawing datasets to study this problem. Our findings demonstrate that beyond producing photo-realistic images from precise spatial inputs, controllable image diffusion can effectively produce a refined, clear sketch from a noisy input. Our work serves as an initial step towards marker-less air drawing and reveals distinct applications of controllable diffusion models to AirSketch and AR/VR in general. Hui Xian Grace Lim, Xuanming Cui, Yogesh S. Rawat, Ser-Nam Lim |
NeurIPS | 3 |
| 2024 | Asynchronous Perception Machine for Efficient Test Time TrainingabstractIn this work, we propose Asynchronous Perception Machine (APM), a computationally-efficient architecture for test-time-training (TTT). APM can process patches of an image one at a time in any order asymmetrically and still encode semantic-awareness in the net. We demonstrate APM's ability to recognize out-of-distribution images without dataset-specific pre-training, augmentation or any-pretext task. APM offers competitive performance over existing TTT approaches. To perform TTT, APM just distills test sample's representation once. APM possesses a unique property: it can learn using just this single representation and starts predicting semantically-aware features.
APM demostrates potential applications beyond test-time-training: APM can scale up to a dataset of 2D images and yield semantic-clusterings in a single forward pass. APM also provides first empirical evidence towards validating GLOM's insight, i.e. input percept is a field. Therefore, APM helps us converge towards an implementation which can do both interpolation and perception on a shared-connectionist hardware. Our code is publicly available at https://rajatmodi62.github.io/apm_project_page/
--------
**It now appears that some of the ideas in GLOM could be made to work.**
https://www.technologyreview.com/2021/04/16/1021871/geoffrey-hinton-glom-godfather-ai-neural-networks/
GLOM = Geoff's Latest Original Model.
```
.-""""""-.
.' '.
/ O O \
| O |
\ '------' /
'. .'
'-....-'
Silent men in deep-contemplation.
Silent men emerges only sometimes.
Silent men love all.
Silent men practice slow science.
``` Rajat Modi, Yogesh S. Rawat |
NeurIPS | 2 |
| 2023 | Hybrid Active Learning via Deep Clustering for Video Action DetectionabstractIn this work, we focus on reducing the annotation cost for video action detection which requires costly frame-wise dense annotations. We study a novel hybrid active learning (AL) strategy which performs efficient labeling using both intra-sample and inter-sample selection. The intra-sample selection leads to labeling of fewer frames in a video as opposed to inter-sample selection which operates at video level. This hybrid strategy reduces the annotation cost from two different aspects leading to significant labeling cost reduction. The proposed approach utilize Clustering-Aware Uncertainty Scoring (CLAUS), a novel label acquisition strategy which relies on both informativeness and diversity for sample selection. We also propose a novel Spatio-Temporal Weighted (STeW) loss formulation, which helps in model training under limited annotations. The proposed approach is evaluated on UCF-101-24 and J-HMDB-21 datasets demonstrating its effectiveness in significantly reducing the annotation cost where it consistently outperforms other baselines. Project details available at https://tinyurl.com/hybridclaus Aayush Jung Rana, Yogesh S. Rawat |
CVPR | 2 |
| 2023 | A Large-Scale Robustness Analysis of Video Action Recognition ModelsabstractWe have seen a great progress in video action recognition in recent years. There are several models based on convolutional neural network (CNN) and some recent transformer based approaches which provide top performance on existing benchmarks. In this work, we perform a large-scale robustness analysis of these existing models for video action recognition. We focus on robustness against real-world distribution shift perturbations instead of adversarial perturbations. We propose four different benchmark datasets, HMDB51-P, UCF101-P, Kinetics400-P, and SSv2-P to perform this analysis. We study robustness of six state-of-the-art action recognition models against 90 different perturbations. The study reveals some interesting findings, 1) transformer based models are consistently more robust compared to CNN based models, 2) Pretraining improves robustness for Transformer based models more than CNN based models, and 3) All of the studied models are robust to temporal perturbations for all datasets but SSv2; suggesting the importance of temporal information for action recognition varies based on the dataset and activities. Next, we study the role of augmentations in model robustness and present a real-world dataset, UCF101-DS, which contains realistic distribution shifts, to further validate some of these findings. We believe this study will serve as a benchmark for future research in robust video action recognition11More details available at bit.1y/3TJLMUF.. Madeline Schiappa, Naman Biyani, Prudvi Kamtam, Shruti Vyas, Hamid Palangi, Vibhav Vineet, Yogesh S. Rawat |
CVPR | 7 |
| 2023 | Efficiently Robustify Pre-Trained ModelsabstractA recent trend in deep learning algorithms has been towards training large scale models, having high parameter count and trained on big dataset. However, robustness of such large scale models towards real-world settings is still a less-explored topic. In this work, we first benchmark the performance of these models under different perturbations and datasets thereby representing real-world shifts, and highlight their degrading performance under these shifts. We then discuss on how complete model fine-tuning based existing robustification schemes might not be a scalable option given very large scale networks and can also lead them to forget some of the desired characterstics. Finally, we propose a simple and cost-effective method to solve this problem, inspired by knowledge transfer literature. It involves robustifying smaller models, at a lower computation cost, and then use them as teachers to tune a fraction of these large scale networks, reducing the overall computational overhead. We evaluate our proposed method under various vision perturbations including ImageNet-C,R,S,A datasets and also for transfer learning, zero-shot evaluation setups on different datasets. Benchmark results show that our method is able to induce robustness to these large scale models efficiently, requiring significantly lower time and also preserves the transfer learning, zero-shot properties of the original model which none of the existing methods are able to achieve. Nishant Jain, Harkirat S. Behl, Yogesh S. Rawat, Vibhav Vineet |
ICCV | 3 |
| 2023 | Revealing the unseen: Benchmarking video action recognition under occlusionabstractIn this work, we study the effect of occlusion on video action recognition. Tofacilitate this study, we propose three benchmark datasets and experiment withseven different video action recognition models. These datasets include two synthetic benchmarks, UCF-101-O and K-400-O, which enabled understanding the effects of fundamental properties of occlusion via controlled experiments. We also propose a real-world occlusion dataset, UCF-101-Y-OCC, which helps in further validating the findings of this study. We find several interesting insights such as 1) transformers are more robust than CNN counterparts, 2) pretraining make modelsrobust against occlusions, and 3) augmentation helps, but does not generalize well to real-world occlusions. In addition, we propose a simple transformer based compositional model, termed as CTx-Net, which generalizes well under this distribution shift. We observe that CTx-Net outperforms models which are trained using occlusions as augmentation, performing significantly better under natural occlusions. We believe this benchmark will open up interesting future research in robust video action recognition Shresth Grover, Vibhav Vineet, Yogesh S. Rawat |
NeurIPS | 3 |
| 2023 | On Occlusions in Video Action Detection: Benchmark Datasets And Training RecipesabstractThis paper explores the impact of occlusions in video action detection. We facilitatethis study by introducing five new benchmark datasets namely O-UCF and O-JHMDB consisting of synthetically controlled static/dynamic occlusions, OVIS-UCF and OVIS-JHMDB consisting of occlusions with realistic motions and Real-OUCF for occlusions in realistic-world scenarios. We formally confirm an intuitiveexpectation: existing models suffer a lot as occlusion severity is increased andexhibit different behaviours when occluders are static vs when they are moving.We discover several intriguing phenomenon emerging in neural nets: 1) transformerscan naturally outperform CNN models which might have even used occlusion as aform of data augmentation during training 2) incorporating symbolic-componentslike capsules to such backbones allows them to bind to occluders never even seenduring training and 3) Islands of agreement (similar to the ones hypothesized inHinton et Al’s GLOM) can emerge in realistic images/videos without instance-levelsupervision, distillation or contrastive-based objectives(eg. video-textual training).Such emergent properties allow us to derive simple yet effective training recipeswhich lead to robust occlusion models inductively satisfying the first two stages ofthe binding mechanism (grouping/segregation). Models leveraging these recipesoutperform existing video action-detectors under occlusion by 32.3% on O-UCF,32.7% on O-JHMDB & 2.6% on Real-OUCF in terms of the vMAP metric. The code for this work has been released at https: //github.com/rajatmodi62/OccludedActionBenchmark. Rajat Modi, Vibhav Vineet, Yogesh S. Rawat |
NeurIPS | 3 |
| 2022 | End-to-End Semi-Supervised Learning for Video Action DetectionabstractIn this work, we focus on semi-supervised learning for video action detection which utilizes both labeled as well as unlabeled data. We propose a simple end-to-end consistency based approach which effectively utilizes the unlabeled data. Video action detection requires both, action class prediction as well as a spatio-temporal localization of actions. Therefore, we investigate two types of constraints, classification consistency, and spatio-temporal consistency. The presence of predominant background and static regions in a video makes it challenging to utilize spatio-temporal consistency for action detection. To address this, we propose two novel regularization constraints for spatio-temporal consistency; 1) temporal coherency, and 2) gradient smoothness. Both these aspects exploit the temporal continuity of action in videos and are found to be effective for utilizing unlabeled videos for action detection. We demonstrate the effectiveness of the proposed approach on two different action detection benchmark datasets, UCF101-24 and IHMDB-21. In addition, we also show the effectiveness of the proposed approach for video object segmentation on the Youtube-VOS which demonstrates its generalization capability The proposed approach achieves competitive performance by using merely 20% of annotations on UCF101-24 when compared with recent fully supervised methods. On UCF101-24, it improves the score by +8.9% and +11% at 0.5 f-mAP and v-mAP respectively, compared to supervised approach. The code and models will be made publicly available at: https://github.com/AKASH2907/End-to-End-Semi-Supervised-Learning-for-Video-Action-Detection. Akash Kumar 0016, Yogesh S. Rawat |
CVPR | 2 |
| 2022 | Don't Pour Cereal into Coffee: Differentiable Temporal Logic for Temporal Action SegmentationabstractWe propose Differentiable Temporal Logic (DTL), a model-agnostic framework that introduces temporal constraints to deep networks. DTL treats the outputs of a network as a truth assignment of a temporal logic formula, and computes a temporal logic loss reflecting the consistency between the output and the constraints. We propose a comprehensive set of constraints, which are implicit in data annotations, and incorporate them with deep networks via DTL. We evaluate the effectiveness of DTL on the temporal action segmentation task and observe improved performance and reduced logical errors in the output of different task models. Furthermore, we provide an extensive analysis to visualize the desirable effects of DTL. Ziwei Xu 0001, Yogesh S. Rawat, Yongkang Wong, Mohan Kankanhalli, Mubarak Shah |
NeurIPS | 2 |
| 2022 | Are all Frames Equal? Active Sparse Labeling for Video Action DetectionabstractVideo action detection requires annotations at every frame, which drastically increases the labeling cost. In this work, we focus on efficient labeling of videos for action detection to minimize this cost. We propose active sparse labeling (ASL), a novel active learning strategy for video action detection. Sparse labeling will reduce the annotation cost but poses two main challenges; 1) how to estimate the utility of annotating a single frame for action detection as detection is performed at video level?, and 2) how these sparse labels can be used for action detection which require annotations on all the frames? This work attempts to address these challenges within a simple active learning framework. For the first challenge, we propose a novel frame-level scoring mechanism aimed at selecting most informative frames in a video. Next, we introduce a novel loss formulation which enables training of action detection model with these sparsely selected frames. We evaluate the proposed approach on two different action detection benchmark datasets, UCF-101-24 and J-HMDB-21, and observed that active sparse labeling can be very effective in saving annotation costs. We demonstrate that the proposed approach performs better than random selection, outperforming all other baselines, with performance comparable to supervised approach using merely 10% annotations. Aayush Jung Rana, Yogesh S. Rawat |
NeurIPS | 2 |
| 2022 | Robustness Analysis of Video-Language Models Against Visual and Language PerturbationsabstractJoint visual and language modeling on large-scale datasets has recently shown good progress in multi-modal tasks when compared to single modal learning. However, robustness of these approaches against real-world perturbations has not been studied. In this work, we perform the first extensive robustness study of video-language models against various real-world perturbations. We focus on text-to-video retrieval and propose two large-scale benchmark datasets, MSRVTT-P and YouCook2-P, which utilize 90 different visual and 35 different text perturbations. The study reveals some interesting initial findings from the studied models: 1) models are more robust when text is perturbed versus when video is perturbed, 2) models that are pre-trained are more robust than those trained from scratch, 3) models attend more to scene and objects rather than motion and action. We hope this study will serve as a benchmark and guide future research in robust video-language learning. The benchmark introduced in this study along with the code and datasets is available at https://bit.ly/3CNOly4. Madeline Schiappa, Shruti Vyas, Hamid Palangi, Yogesh S. Rawat, Vibhav Vineet |
NeurIPS | 4 |
| 2022 | Pose-guided Generative Adversarial Net for Novel View Action SynthesisabstractWe focus on the problem of novel-view human action synthesis. Given an action video, the goal is to generate the same action from an unseen viewpoint. Naturally, novel view video synthesis is more challenging than image synthesis. It requires the synthesis of a sequence of realistic frames with temporal coherency. Besides, transferring different actions to a novel target view requires awareness of action category and viewpoint change simultaneously. To address these challenges we propose a novel framework named Pose-guided Action Separable Generative Adversarial Net (PAS-GAN), which utilizes pose to alleviate the difficulty of this task. First, we propose a recurrent pose-transformation module which transforms actions from the source view to the target view and generates novel view pose sequence in 2D coordinate space. Second, a well-transformed pose sequence enables us to separate the action and background in the target view. We employ a novel local-global spatial transformation module to effectively generate sequential video features in the target view using these action and background features. Finally, the generated video features are used to synthesize human action with the help of a 3D decoder. Moreover, to focus on dynamic action in the video, we propose a novel multi-scale action-separable loss which further improves the video quality. We conduct extensive experiments on two large-scale multi-view human action datasets, NTU-RGBD and PKU-MMD, demonstrating the effectiveness of PAS-GAN which outperforms existing approaches. The codes and models will be available on https://github.com/xhl-video/PAS-GAN. Xianhang Li, Junhao Zhang 0001, Kunchang Li 0002, Shruti Vyas, Yogesh S. Rawat |
WACV | 5 |
| 2021 | SSA2D: Single Shot Actor-Action Detection in Videos (Student Abstract)abstractWe propose a single-shot approach for actor-action detection in videos. The existing approaches use a two-step process, which rely on Region Proposal Network (RPN), where the action is estimated based on the detected proposals followed by post-processing such as non-maximal suppression. While effective in terms of performance, these methods pose limitations in scalability for dense video scenes with a high memory requirement for thousand of proposals, which leads to slow processing time. We propose SSA2D, a unified end-to-end deep network, which performs joint actor-action detection in a single-shot without the need of any proposals and post-processing, making it memory as well as time efficient. Aayush Jung Rana, Yogesh S. Rawat |
AAAI | 2 |
| 2021 | LARNet: Latent Action Representation for Human Action Synthesis
Naman Biyani, Aayush Jung Rana, Shruti Vyas, Yogesh S. Rawat |
BMVC | 4 |
| 2021 | Modeling Multi-Label Action Dependencies for Temporal Action LocalizationabstractReal-world videos contain many complex actions with inherent relationships between action classes. In this work, we propose an attention-based architecture that models these action relationships for the task of temporal action localization in untrimmed videos. As opposed to previous works that leverage video-level co-occurrence of actions, we distinguish the relationships between actions that occur at the same time-step and actions that occur at different time-steps (i.e. those which precede or follow each other). We define these distinct relationships as action dependencies. We propose to improve action localization performance by modeling these action dependencies in a novel attention-based Multi-Label Action Dependency (MLAD) layer. The MLAD layer consists of two branches: a Cooccurrence Dependency Branch and a Temporal Dependency Branch to model co-occurrence action dependencies and temporal action dependencies, respectively. We observe that existing metrics used for multi-label classification do not explicitly measure how well action dependencies are modeled, therefore, we propose novel metrics that consider both co-occurrence and temporal dependencies between action classes. Through empirical evaluation and extensive analysis, we show improved performance over state-of-the-art methods on multi-label action localization benchmarks (MultiTHUMOS and Charades) in terms of f-mAP and our proposed metric. Code is publicly available at https://github.com/ptirupat/MLAD. Praveen Tirupattur, Kevin Duarte, Yogesh S. Rawat, Mubarak Shah |
CVPR | 3 |
| 2021 | Novel View Video Prediction using a Dual RepresentationabstractWe address the problem of novel view video prediction; given a set of input video clips from a single/multiple views, our network is able to predict the video from a novel view. The proposed approach does not require any priors and is able to predict the video from wider angular distances, upto 45 degree, as compared to the recent studies predicting small variations in viewpoint. Moreover, our method relies only on RGB frames to learn a dual representation which is used to generate the video from a novel viewpoint. The dual representation encompasses a view-dependent and a global representation which incorporates complementary details to enable novel view video prediction. We demonstrate the effectiveness of our framework on two real world datasets: NTU- RGB+D and CMU Panoptic. A comparison with the State-of-the-art novel view video prediction methods shows an improvement of 26.1% in SSIM, 13.6% in PSNR, and 60% in FVD scores without using explicit priors from target views. Sarah Shiraz, Krishna Regmi, Shruti Vyas, Yogesh S. Rawat, Mubarak Shah |
ICIP | 4 |
| 2021 | Unsupervised Discriminative Embedding For Sub-Action Learning in Complex ActivitiesabstractThis paper proposes a novel approach for unsupervised sub-action learning in complex activities. The proposed method maps both visual and temporal representations to a latent space where the sub-actions are learnt discriminatively in an end-to-end fashion. To this end, we propose to learn sub-actions as latent concepts and a novel discriminative latent concept learning (DLCL) module aids in learning sub-actions. The proposed DLCL module lends on the idea of latent concepts to learn compact representations in the latent embedding space in an unsupervised way. The result is a set of latent vectors that can be interpreted as cluster centers in the embedding space. Our joint embedding learning with discriminative latent concept module is novel which eliminates the need for explicit clustering. We validate our approach on three benchmark datasets and show that the proposed combination of visual-temporal embedding and discriminative latent concepts allow to learn robust action representations in unsupervised setting. Sirnam Swetha, Hilde Kuehne, Yogesh S. Rawat, Mubarak Shah |
ICIP | 3 |
| 2021 | In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised Learning
Mamshad Nayeem Rizve, Kevin Duarte, Yogesh S. Rawat, Mubarak Shah |
ICLR | 3 |
| 2021 | Trustworthy AI'21: 1st International Workshop on Trustworthy AI for Multimedia ComputingabstractIn this workshop, we are addressing the trustworthy AI issues for Multimedia Computing. We aim to bring together researchers in the trustworthy aspects of Multimedia Computing and facilitate discussions in injecting trusts into multimedia to develop trustworthy AI techniques that are reliable and acceptable to multimedia researchers and practitioners. Our scope is at the conjunction of multimedia, computer vision and trustworthy AI, including Explainability, Robustness and Safety, Data Privacy, Accountability and Transparency, and Fairness. Teddy Furon, Jingen Liu, Yogesh S. Rawat, Wei Zhang 0031, Qi Zhao 0001 |
ACM Multimedia | 3 |
| 2021 | NoisyActions2M: A Multimedia Dataset for Video Understanding from Noisy LabelsabstractDeep learning has shown remarkable progress in a wide range of problems. However, efficient training of such models requires large-scale datasets, and getting annotations for such datasets can be challenging and costly. In this work, we explore user-generated freely available labels from web videos for video understanding. We create a benchmark dataset consisting of around 2 million videos with associated user-generated annotations and other meta information. We utilize the collected dataset for action classification and demonstrate its usefulness with existing small-scale annotated datasets, UCF101 and HMDB51. We study different loss functions and two pretraining strategies, simple and self-supervised learning. We also show how a network pretrained on the proposed dataset can help against video corruption and label noise in downstream datasets. We present this as a benchmark dataset in noisy learning for video understanding. The dataset, code, and trained models are publicly available here for future research. A longer version of our paper is also available here. Mohit Sharma 0004, Raj Patra, Harshal Desai, Shruti Vyas, Yogesh S. Rawat, Rajiv Ratn Shah |
MMAsia | 5 |
| 2021 | Reformulating Zero-shot Action Recognition for Multi-label ActionsabstractThe goal of zero-shot action recognition (ZSAR) is to classify action classes which were not previously seen during training. Traditionally, this is achieved by training a network to map, or regress, visual inputs to a semantic space where a nearest neighbor classifier is used to select the closest target class. We argue that this approach is sub-optimal due to the use of nearest neighbor on static semantic space and is ineffective when faced with multi-label videos - where two semantically distinct co-occurring action categories cannot be predicted with high confidence. To overcome these limitations, we propose a ZSAR framework which does not rely on nearest neighbor classification, but rather consists of a pairwise scoring function. Given a video and a set of action classes, our method predicts a set of confidence scores for each class independently. This allows for the prediction of several semantically distinct classes within one video input. Our evaluations show that our method not only achieves strong performance on three single-label action classification datasets (UCF-101, HMDB, and RareAct), but also outperforms previous ZSAR approaches on a challenging multi-label dataset (AVA) and a real-world surprise activity detection dataset (MEVA). Alec Kerrigan, Kevin Duarte, Yogesh S. Rawat, Mubarak Shah |
NeurIPS | 3 |
| 2021 | We don't Need Thousand Proposals: Single Shot Actor-Action Detection in VideosabstractWe propose SSA2D, a simple yet effective end-to-end deep network for actor-action detection in videos. The existing methods take a top-down approach based on region-proposals (RPN), where the action is estimated based on the detected proposals followed by post-processing such as non-maximal suppression. While effective in terms of performance, these methods pose limitations in scalability for dense video scenes with a high memory requirement for thousands of proposals. We propose to solve this problem from a different perspective where we don't need any proposals. SSA2D is a unified network, which performs pixel level joint actor-action detection in a single-shot, where every pixel of the detected actor is assigned an action label. SSA2D has two main advantages: 1) It is a fully convolutional network which does not require any proposals and post-processing making it memory as well as time efficient, 2) It is easily scalable to dense video scenes as its memory requirement is independent of the number of actors present in the scene. We evaluate the proposed method on the Actor-Action dataset (A2D) and Video Object Relation (VidOR) dataset, demonstrating its effectiveness in multiple actors and action detection in a video. SSA2D is 11x faster during inference with comparable (sometimes better) performance and fewer network parameters when compared with the prior works. Code available at https://github.com/aayushjr/ssa2d. Aayush Jung Rana, Yogesh S. Rawat |
WACV | 2 |
| 2021 | Adversarial Learning for Personalized Tag RecommendationabstractWe have recently seen great progress in image classification due to the success of deep convolutional neural networks and the availability of large-scale datasets. Most of the existing work focuses on single-label image classification. However, there are usually multiple tags associated with an image. The existing works on multi-label classification are mainly based on lab curated labels. Humans assign tags to their images differently, which is mainly based on their interests and personal tagging behavior. In this paper, we address the problem of personalized tag recommendation and propose an end-to-end deep network which can be trained on large-scale datasets. The user-preference is learned within the network in an unsupervised way where the network performs joint optimization for user-preference and visual encoding. A joint training of user-preference and visual encoding allows the network to efficiently integrate the visual preference with tagging behavior for a better user recommendation. In addition, we propose the use of adversarial learning, which enforces the network to predict tags resembling user-generated tags. We demonstrate the effectiveness of the proposed model on two different large-scale and publicly available datasets, YFCC100 M and NUS-WIDE. The proposed method achieves significantly better performance on both the datasets when compared to the baselines and other state-of-the-art methods. The code is publicly available at https://github.com/vyzuer/ALTReco. Erik Quintanilla, Yogesh S. Rawat, Andrey Sakryukin, Mubarak Shah, Mohan Kankanhalli |
IEEE Trans. Multim. | 2 |
| 2020 | Visual-Textual Capsule Routing for Text-Based Video SegmentationabstractJoint understanding of vision and natural language is a challenging problem with a wide range of applications in artificial intelligence. In this work, we focus on integration of video and text for the task of actor and action video segmentation from a sentence. We propose a capsule-based approach which performs pixel-level localization based on a natural language query describing the actor of interest. We encode both the video and textual input in the form of capsules, which provide a more effective representation in comparison with standard convolution based features. Our novel visual-textual routing mechanism allows for the fusion of video and text capsules to successfully localize the actor and action. The existing works on actor-action localization are mainly focused on localization in a single frame instead of the full video. Different from existing works, we propose to perform the localization on all frames of the video. To validate the potential of the proposed network for actor and action video localization, we extend an existing actor-action dataset (A2D) with annotations for all the frames. The experimental evaluation demonstrates the effectiveness of our capsule network for text selective actor and action localization in videos. The proposed method also improves upon the performance of the existing state-of-the art works on single frame-based localization. Bruce McIntosh, Kevin Duarte, Yogesh S. Rawat, Mubarak Shah |
CVPR | 3 |
| 2020 | A Recurrent Transformer Network for Novel View Action Synthesis
Kara Schatz, Erik Quintanilla, Shruti Vyas, Yogesh S. Rawat |
ECCV (27) | 4 |
| 2020 | Multi-view Action Recognition Using Cross-View Video Prediction
Shruti Vyas, Yogesh S. Rawat, Mubarak Shah |
ECCV (27) | 2 |
| 2020 | TinyVIRAT: Low-resolution Video Action RecognitionabstractThe existing research in action recognition is mostly focused on high-quality videos where the action is distinctly visible. In real-world surveillance environments, the actions in videos are captured at a wide range of resolutions. Most activities occur at a distance with a small resolution and recognizing such activities is a challenging problem. In this work, we focus on recognizing tiny actions in videos. We introduce a benchmark dataset, Tiny VIRAT, which contains natural low-resolution activities. The actions in Tiny VIRAT videos have multiple labels and they are extracted from surveillance videos which makes them realistic and more challenging. We propose a novel method for recognizing tiny actions in videos which utilizes a progressive generative approach to improve the quality of low-resolution actions. The proposed method also consists of a weakly trained attention mechanism which helps in focusing on the activity regions in the video. We perform extensive experiments to benchmark the proposed Tiny VIRAT dataset and observe that the proposed method significantly improves the action recognition performance over baselines. We also evaluate the proposed approach on synthetically resized action recognition datasets and achieve state-of-the-art results when compared with existing methods. The dataset and code are publicly available at https://github.com/UgurDemir/Tiny-VIRAT. Ugur Demir, Yogesh S. Rawat, Mubarak Shah |
ICPR | 2 |
| 2020 | Gabriella: An Online System for Real-Time Activity Detection in Untrimmed Security VideosabstractActivity detection in security videos is a difficult problem due to multiple factors such as large field of view, presence of multiple activities, varying scales and viewpoints, and its untrimmed nature. The existing research in activity detection is mainly focused on datasets, such as UCF-101, JHMDB, THUMOS, and AVA, which partially address these issues. The requirement of processing security videos in real-time makes this even more challenging. In this work, we propose Gabriella, a real-time online system to perform activity detection on untrimmed security videos. The proposed method consists of three stages: tubelet extraction, activity classification, and online tubelet merging. For tubelet extraction, we propose a localization network which takes a video clip as input and spatio-temporally detects potential foreground regions at multiple scales to generate action tubelets. We propose a novel Patch-Dice loss to handle large variations in actor size. Our online processing of videos at a clip level drastically reduces the computation time in detecting activities. The detected tubelets are assigned activity class scores by the classification network and merged together using our proposed Tubelet-Merge Action-Split (TMAS) algorithm to form the final action detections. The TMAS algorithm efficiently connects the tubelets in an online fashion to generate action detections which are robust against varying length activities. We perform our experiments on the VIRAT and MEVA (Multiview Extended Video with Activities) datasets and demonstrate the effectiveness of the proposed approach in terms of speed (~100 fps) and performance with state-of-the-art results. More details about this work are available on our project webpage11https://www.crcv.ucf.edu/research/projects/gabriella-an-online-system-for-real-time-activity-detection-in-untrimmed-security-videos. Mamshad Nayeem Rizve, Ugur Demir, Praveen Tirupattur, Aayush Jung Rana, Kevin Duarte, Ishan R. Dave, Yogesh S. Rawat, Mubarak Shah |
ICPR | 7 |
| 2020 | Photography and Exploration of Tourist Locations Based on Optimal Foraging TheoryabstractAnimals search for food in their environment with a decision strategy which keeps them fit. Optimal foraging theory models this foraging behavior to determine the optimal decision strategy followed by animals. This theory has been successfully applied for humans as they search for information and is termed as information foraging. When people visit a tourist location, they follow a similar strategy to move from one spot to another, and collect information by capturing photographs. This behavior has similarities with the foraging behavior of animals which has been widely studied by the researchers. In this paper, we propose to employ optimal foraging theory to help tourists explore a location and capture photographs in an optimal way. We determine a decision strategy for tourist which provides a list of interesting spots to visit in a tourist location along with corresponding stay time. Finally, we solve an optimization problem to find a path through these spots which can be followed by tourists. The experimental results on a public dataset demonstrate the effectiveness of the proposed method.1Code available at https://github.com/vyzuer/foraging_theory. Yogesh S. Rawat, Mubarak Shah, Mohan Kankanhalli |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | CapsuleVOS: Semi-Supervised Video Object Segmentation Using Capsule RoutingabstractIn this work we propose a capsule-based approach for semi-supervised video object segmentation. Current video object segmentation methods are frame-based and often require optical flow to capture temporal consistency across frames which can be difficult to compute. To this end, we propose a video based capsule network, CapsuleVOS, which can segment several frames at once conditioned on a reference frame and segmentation mask. This conditioning is performed through a novel routing algorithm for attention-based efficient capsule selection. We address two challenging issues in video object segmentation: 1) segmentation of small objects and 2) occlusion of objects across time. The issue of segmenting small objects is addressed with a zooming module which allows the network to process small spatial regions of the video. Apart from this, the framework utilizes a novel memory module based on recurrent networks which helps in tracking objects when they move out of frame or are occluded. The network is trained end-to-end and we demonstrate its effectiveness on two benchmark video object segmentation datasets; it outperforms current offline approaches on the Youtube-VOS dataset while having a run-time that is almost twice as fast as competing methods. The code is publicly available at https://github.com/KevinDuarte/CapsuleVOS. Kevin Duarte, Yogesh S. Rawat, Mubarak Shah |
ICCV | 2 |
| 2018 | ThoughtViz: Visualizing Human Thoughts Using Generative Adversarial NetworkabstractStudying human brain signals has always gathered great attention from the scientific community. In Brain Computer Interface (BCI) research, for example, changes of brain signals in relation to specific tasks (e.g., thinking something) are detected and used to control machines. While extracting spatio-temporal cues from brain signals for classifying state of human mind is an explored path, decoding and visualizing brain states is new and futuristic. Following this latter direction, in this paper, we propose an approach that is able not only to read the mind, but also to decode and visualize human thoughts. More specifically, we analyze brain activity, recorded by an ElectroEncephaloGram (EEG), of a subject while thinking about a digit, character or an object and synthesize visually the thought item. To accomplish this, we leverage the recent progress of adversarial learning by devising a conditional Generative Adversarial Network (GAN), which takes, as input, encoded EEG signals and generates corresponding images. In addition, since collecting large EEG signals in not trivial, our GAN model allows for learning distributions with limited training data. Performance analysis carried out on three different datasets -- brain signals of multiple subjects thinking digits, characters, and objects -- show that our approach is able to effectively generate images from thoughts of a person. They also demonstrate that EEG signals encode explicitly cues from thoughts which can be effectively used for generating semantically relevant visualizations. Praveen Tirupattur, Yogesh S. Rawat, Concetto Spampinato, Mubarak Shah |
ACM Multimedia | 2 |
| 2018 | VideoCapsuleNet: A Simplified Network for Action DetectionabstractThe recent advances in Deep Convolutional Neural Networks (DCNNs) have shown extremely good results for video human action classification, however, action detection is still a challenging problem. The current action detection approaches follow a complex pipeline which involves multiple tasks such as tube proposals, optical flow, and tube classification. In this work, we present a more elegant solution for action detection based on the recently developed capsule network. We propose a 3D capsule network for videos, called VideoCapsuleNet: a unified network for action detection which can jointly perform pixel-wise action segmentation along with action classification. The proposed network is a generalization of capsule network from 2D to 3D, which takes a sequence of video frames as input. The 3D generalization drastically increases the number of capsules in the network, making capsule routing computationally expensive. We introduce capsule-pooling in the convolutional capsule layer to address this issue and make the voting algorithm tractable. The routing-by-agreement in the network inherently models the action representations and various action characteristics are captured by the predicted capsules. This inspired us to utilize the capsules for action localization and the class-specific capsules predicted by the network are used to determine a pixel-wise localization of actions. The localization is further improved by parameterized skip connections with the convolutional capsule layers and the network is trained end-to-end with a classification as well as localization loss. The proposed network achieves state-of-the-art performance on multiple action detection datasets including UCF-Sports, J-HMDB, and UCF-101 (24 classes) with an impressive ~20% improvement on UCF-101 and ~15% improvement on J-HMDB in terms of v-mAP scores. Kevin Duarte, Yogesh S. Rawat, Mubarak Shah |
NeurIPS | 2 |
| 2018 | A Spring-Electric Graph Model for Socialized Group PhotographyabstractVisual balance is considered as one of the important factors in defining the aesthetic quality of visual arts. In this paper, we propose a novel method to obtain visual balance in a layout with dynamic visual elements. We use the idea of a spring-electric graph model and augment it with the concept of color energy from the literature of visual arts. We also present an interesting application of the proposed model in photography assistance. We focus on group photography and utilize social media images along with the proposed spring-electric model for providing a recommendation to the user. The proposed method can provide real-time feedback to the user regarding the arrangement of people, their position, and relative size on the image frame. We conducted qualitative experiments along with user studies to evaluate the proposed method. Experimental results and user studies show the effectiveness of the proposed model in obtaining visual balance and group photography recommendation. Yogesh S. Rawat, Mingli Song, Mohan Kankanhalli |
IEEE Trans. Multim. | 1 |
| 2017 | ClickSmart: A Context-Aware Viewpoint Recommendation System for Mobile PhotographyabstractIn this paper, we propose ClickSmart, a viewpoint recommendation system that can assist a user in capturing high-quality photographs at well-known tourist locations. ClickSmart can provide real-time viewpoint recommendation based on the preview on the user's camera, current time, and user's geolocation. It makes use of publicly available geotagged images along with the associated metadata for learning a recommendation model. We define view-cells, macroblocks in geospace, and propose the concepts of popularity, quality, and uniqueness of view-cells from the viewpoint perspective. Viewpoint recommendation is generated at the granularity of a view-cell and is based on its popularity, quality, and uniqueness, which are estimated using social media cues associated with images. We further observe that contextual information such as time and weather conditions play an important role in photography, and therefore augment the recommendation system with the associated context. ClickSmart also takes into account the presence of people in the view for making the recommendation. It can provide two kinds of recommendations, quality based and uniqueness based. Although both were found effective in the experimental evaluation, our user study showed that uniqueness-based recommendation was preferred more by skilled photographers compared with amateurs. Yogesh S. Rawat, Mohan Kankanhalli |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | ConTagNet: Exploiting User Context for Image Tag RecommendationabstractIn recent years, deep convolutional neural networks have shown great success in single-label image classification. However, images usually have multiple labels associated with them which may correspond to different objects or actions present in the image. In addition, a user assigns tags to a photo not merely based on the visual content but also the context in which the photo has been captured. Inspired by this, we propose a deep neural network which can predict multiple tags for an image based on the content as well as the context in which the image is captured. The proposed model can be trained end-to-end and solves a multi-label classification problem. We evaluate the model on a dataset of 1,965,232 images which is drawn from the YFCC100M dataset provided by the organizers of Yahoo-Flickr Grand Challenge. We observe a significant improvement in the prediction accuracy after integrating user-context and the proposed model performs very well in the Grand Challenge. Yogesh S. Rawat, Mohan Kankanhalli |
ACM Multimedia | 1 |
| 2015 | Real-Time Assistance in Multimedia Capture Using Social MediaabstractIn the last decade, we have seen significant improvement in the ease and cost of capturing multimedia content. However, the aesthetic quality of the content captured by an amateur user still needs substantial improvement. This doctoral research aims at providing real-time assistance to amateur users so that they can capture high quality photographs and home videos. Our approach is focused on learning the art of photography and videography from multimedia content shared on social media. We have proposed a context-based photography learning method which can assist a user in capturing high quality photographs. The photography learning is augmented with contextual information such as time, geo-location, environmental conditions and type of image, which have an impact on photography. The proposed method can provide real-time feedback to the user regarding scene composition, camera parameters, camera movement and viewpoint. We have presented some preliminary results and also described the planned future work. Yogesh S. Rawat |
ACM Multimedia | 1 |
| 2015 | Context-Aware Photography Learning for Smart Mobile DevicesabstractIn this work we have developed a photography model based on machine learning which can assist a user in capturing high quality photographs. As scene composition and camera parameters play a vital role in aesthetics of a captured image, the proposed method addresses the problem of learning photographic composition and camera parameters. Further, we observe that context is an important factor from a photography perspective, we therefore augment the learning with associated contextual information. The proposed method utilizes publicly available photographs along with social media cues and associated metainformation in photography learning. We define context features based on factors such as time, geolocation, environmental conditions and type of image, which have an impact on photography. We also propose the idea of computing the photographic composition basis, eigenrules and baserules , to support our composition learning. The proposed system can be used to provide feedback to the user regarding scene composition and camera parameters while the scene is being captured. It can also recommend position in the frame where people should stand for better composition. Moreover, it also provides camera motion guidance for pan, tilt and zoom to the user for improving scene composition. Yogesh S. Rawat, Mohan Kankanhalli |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2014 | Context-Based Photography Learning using Crowdsourced Images and Social MediaabstractThis paper presents a photography model based on machine learning which utilizes crowd-sourced images along with social media cues. As scene composition and camera parameters play a vital role in aesthetics of a captured image, the proposed system addresses the problem of learning photographic composition and camera parameters. Further, we observe that context is an important factor from a photography perspective, we therefore augment the learning with associated contextual information. We define context features based on factors such as time, geo-location, environmental conditions and type of image, which have an impact on photography. The meta information available with crowd-sourced images is utilized for context identification and social media cues are used for photo quality evaluation. We also propose the idea of computing the photographic composition basis, eigenrules and baserules, to support our composition learning method. The trained photography model can provide assistance to the user in determining image composition and camera parameters. Yogesh S. Rawat, Mohan Kankanhalli |
ACM Multimedia | 1 |