EDBT 2026 Demo / reviewers in the wild / expert
Shreyank N. Gowda
dblp:191/1593 · also Shreyank Narayana Gowda
· DBLP profile ↗
20ranked-venue papers
11as first author
18since 2021 · last 2026
0000-0002-4975-0705ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 11 first-author · 16 since 2021Artificial intelligence and machine learning · 16 · 10 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adversarial Robustness in Zero-Shot Learning: An Empirical Study on Class and Concept-Level VulnerabilitiesabstractZero-shot Learning (ZSL) aims to enable image classifiers to recognize images from unseen classes that were not included during training. Unlike traditional supervised classification, ZSL typically relies on learning a mapping from visual features to predefined, human-understandable class concepts. While ZSL models promise to improve generalization and interpretability, their robustness under systematic input perturbations remain unclear. In this study, we present an empirical analysis about the robustness of existing ZSL methods at both class-level and concept-level. Specifically, we successfully disrupted their class prediction by the well-known non-target class attack (clsA). However, in the Generalized Zero-shot Learning (GZSL) setting, we observe that the success of clsA is only at the original best-calibrated point. After the attack, the optimal best-calibration point shifts, and ZSL models maintain relatively strong performance at other calibration points, indicating that clsA results in a spurious attack success in the GZSL. To address this, we propose the Class-Bias Enhanced Attack (CBEA), which completely eliminates GZSL accuracy across all calibrated points by enhancing the gap between seen and unseen class probabilities. Next, at concept-level attack, we introduce two novel attack modes: Class-Preserving Concept Attack (CPconA) and Non-Class-Preserving Concept Attack (NCPconA). Our extensive experiments evaluate three typical ZSL models across various architectures from the past three years and reveal that ZSL models are vulnerable not only to the traditional class attack but also to concept-based attacks. These attacks allow malicious actors to easily manipulate class predictions by erasing or introducing concepts. Our findings highlight a significant performance gap between existing approaches, emphasizing the need for improved adversarial robustness in current ZSL models. Our codes are available at https://github.com/FouriYe/AttackZSL_TIP26. Zihan Ye, Shreyank N. Gowda, Yuping Yan, Ling Shao 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Principles of Visual Tokens for Efficient Video UnderstandingabstractVideo understanding has made huge strides in recent years, relying largely on the power of transformers. As this architecture is notoriously expensive and video data is highly redundant, research into improving efficiency has become particularly relevant. Some creative solutions include token selection and merging. While most methods succeed in reducing the cost of the model and maintaining accuracy, an interesting pattern arises: most methods do not outperform the baseline of randomly discarding tokens. In this paper we take a closer look at this phenomenon and observe 5 principles of the nature of visual tokens. For example, we observe that the value of tokens follows a clear Pareto-distribution where most tokens have remarkably low value, and just a few carry most of the perceptual information. We build on these and further insights to propose a lightweight video model, LITE, that can select a small number of tokens effectively, outperforming state-of-the-art and existing baselines across datasets (Kinetics-400 and Something-Something-V2) in the challenging trade-off of computation (GFLOPs) vs accuracy. Experiments also show that LITE generalizes across datasets and even other tasks without the need for retraining. Xinyue Hao 0001, Gen Li 0008, Shreyank N. Gowda, Robert B. Fisher, Jonathan Huang, Anurag Arnab, Laura Sevilla-Lara |
ICCV | 3 |
| 2025 | ZeroDiff: Solidified Visual-semantic Correlation in Zero-Shot LearningabstractZero-shot Learning (ZSL) aims to enable classifiers to identify unseen classes. This is typically achieved by generating visual features for unseen classes based on learned visual-semantic correlations from seen classes. However, most current generative approaches heavily rely on having a sufficient number of samples from seen classes. Our study reveals that a scarcity of seen class samples results in a marked decrease in performance across many generative ZSL techniques. We argue, quantify, and empirically demonstrate that this decline is largely attributable to spurious visual-semantic correlations. To address this issue, we introduce ZeroDiff, an innovative generative framework for ZSL that incorporates diffusion mechanisms and contrastive representations to enhance visual-semantic correlations. ZeroDiff comprises three key components: (1) Diffusion augmentation, which naturally transforms limited data into an expanded set of noised data to mitigate generative model overfitting; (2) Supervised-contrastive (SC)-based representations that dynamically characterize each limited sample to support visual feature generation; and (3) Multiple feature discriminators employing a Wasserstein-distance-based mutual learning approach, evaluating generated features from various perspectives, including pre-defined semantics, SC-based representations, and the diffusion process. Extensive experiments on three popular ZSL benchmarks demonstrate that ZeroDiff not only achieves significant improvements over existing ZSL methods but also maintains robust performance even with scarce training data. Our codes are available at https://github.com/FouriYe/ZeroDiff_ICLR25. Zihan Ye, Shreyank N. Gowda, Shiming Chen 0002, Xiaowei Huang 0001, Fahad Shahbaz Khan, Yaochu Jin, Kaizhu Huang, Xiao-Bo Jin |
ICLR | 2 |
| 2025 | FATE: A Prompt-Tuning-Based Semi-Supervised Learning Framework for Extremely Limited Labeled DataabstractSemi-supervised learning (SSL) has achieved significant progress by leveraging both labeled data and unlabeled data. Existing SSL methods overlook a common real-world scenario when labeled data is extremely scarce, potentially as limited as a single labeled sample in the dataset. General SSL approaches struggle to train effectively from scratch under such constraints, while methods utilizing pre-trained models often fail to find an optimal balance between leveraging limited labeled data and abundant unlabeled data. To address this challenge, we propose Firstly Adapt, Then catEgorize (FATE), a novel SSL framework tailored for scenarios with extremely limited labeled data. At its core, the two-stage prompt tuning paradigm FATE exploits unlabeled data to compensate for scarce supervision signals, then transfers to downstream tasks. Concretely, FATE first adapts a pre-trained model to the feature distribution of downstream data using volumes of unlabeled samples in an unsupervised manner. It then applies an SSL method specifically designed for pre-trained models to complete the final classification task. FATE is designed to be compatible with both vision and vision-language pre-trained models. Extensive experiments demonstrate that FATE effectively mitigates challenges arising from the scarcity of labeled samples in SSL, achieving an average performance improvement of 33.74% across seven benchmarks compared to state-of-the-art SSL methods. Code is available at https://github.com/ganchi-huanggua/FATE.git. Hezhao Liu, Yang Lu 0009, Mengke Li 0001, Yiqun Zhang 0006, Shreyank N. Gowda, Chen Gong 0002, Hanzi Wang |
ACM Multimedia | 5 |
| 2025 | Progressive Data Dropout: An Embarrassingly Simple Approach to Train FasterabstractThe success of the machine learning field has reliably depended on training on large datasets. While effective, this trend comes at an extraordinary cost. This is due to two deeply intertwined factors: the size of models and the size of datasets. While promising research efforts focus on reducing the size of models, the other half of the equation remains fairly mysterious. Indeed, it is surprising that the standard approach to training remains to iterate over and over, uniformly sampling the training dataset. In this paper we explore a series of alternative training paradigms that leverage insights from hard-data-mining and dropout, simple enough to implement and use that can become the new training standard. The proposed Progressive Data Dropout reduces the number of effective epochs to as little as 12.4\% of the baseline. This savings actually do not come at any cost for accuracy. Surprisingly, the proposed method improves accuracy by up to 4.82\%. Our approach requires no changes to model architecture or optimizer, and can be applied across standard training pipelines, thus posing an excellent opportunity for wide adoption. Code can be found here: \url{https://github.com/bazyagami/LearningWithRevision}. Shriram M. S, Xinyue Hao 0001, Shihao Hou, Yang Lu 0009, Laura Sevilla-Lara, Anurag Arnab, Shreyank N. Gowda |
NeurIPS | 7 |
| 2025 | Performance is not All You Need: Sustainability Considerations for Algorithms
Chong Zhang 0006, Shreyank N. Gowda, Yushi Li, Xiao-Bo Jin |
PRCV (12) | 4 |
| 2024 | Continual Learning Improves Zero-Shot Action Recognition
Shreyank N. Gowda, Davide Moltisanti, Laura Sevilla-Lara |
ACCV (3) | 1 |
| 2024 | Telling Stories for Common Sense Zero-Shot Action Recognition
Shreyank N. Gowda, Laura Sevilla-Lara |
ACCV (3) | 1 |
| 2024 | Bridging the Projection Gap: Overcoming Projection Bias Through Parameterized Distance Learning
Chong Zhang 0006, Mingyu Jin, Qinkai Yu, Haochen Xue, Shreyank N. Gowda, Xiao-Bo Jin |
ACCV (8) | 5 |
| 2024 | Optimizing Factorized Encoder Models: Time and Memory Reduction for Scalable and Efficient Action Recognition
Shreyank N. Gowda, Anurag Arnab, Jonathan Huang |
ECCV (10) | 1 |
| 2024 | CC-SAM: SAM with Cross-Feature Attention and Context for Ultrasound Image Segmentation
Shreyank N. Gowda, David A. Clifton |
ECCV (45) | 1 |
| 2024 | FE-Adapter: Adapting Image-Based Emotion Classifiers to VideosabstractUtilizing large pre-trained models for specific tasks has yielded impressive results. However, fully fine-tuning these increasingly large models is becoming prohibitively resource-intensive. This has led to a focus on more parameter-efficient transfer learning, primarily within the same modality. But this approach has limitations, particularly in video understanding where suitable pre-trained models are less common. Addressing this, our study introduces a novel cross-modality transfer learning approach from images to videos, which we call parameter-efficient image-to-video transfer learning. We present the Facial-Emotion Adapter (FE-Adapter), designed for efficient fine-tuning in video tasks. This adapter allows pre-trained image models, which traditionally lack temporal processing capabilities, to analyze dynamic video content efficiently. Notably, it uses about 15 times fewer parameters than previous methods, while improving accuracy. Our experiments in video emotion recognition demonstrate that the FE-Adapter can match or even surpass existing fine-tuning and video emotion models in both performance and efficiency. This breakthrough highlights the potential for cross-modality approaches in enhancing the capabilities of AI models, particularly in fields like video emotion analysis where the demand for efficiency and accuracy is constantly rising. Shreyank N. Gowda, Boyan Gao, David A. Clifton |
FG | 1 |
| 2024 | Masks and Manuscripts: Advancing Medical Pre-training with End-to-End Masking and Narrative Structuring
Shreyank N. Gowda, David A. Clifton |
MICCAI (11) | 1 |
| 2022 | A Closer Look at Temporal Ordering in the Segmentation of Instructional Videos
Anil Batra, Shreyank N. Gowda, Frank Keller, Laura Sevilla-Lara |
BMVC | 2 |
| 2022 | Capturing Temporal Information in a Single Frame: Channel Sampling Strategies for Action Recognition
Kiyoon Kim, Shreyank N. Gowda, Oisin Mac Aodha, Laura Sevilla-Lara |
BMVC | 2 |
| 2022 | Learn2Augment: Learning to Composite Videos for Data Augmentation in Action Recognition
Shreyank N. Gowda, Marcus Rohrbach, Frank Keller, Laura Sevilla-Lara |
ECCV (31) | 1 |
| 2022 | CLASTER: Clustering with Reinforcement Learning for Zero-Shot Action Recognition
Shreyank N. Gowda, Laura Sevilla-Lara, Frank Keller, Marcus Rohrbach |
ECCV (20) | 1 |
| 2021 | SMART Frame Selection for Action RecognitionabstractVideo classification is computationally expensive. In this paper, we address theproblem of frame selection to reduce the computational cost of video classification.Recent work has successfully leveraged frame selection for long, untrimmed videos,where much of the content is not relevant, and easy to discard. In this work, however,we focus on the more standard short, trimmed video classification problem. Weargue that good frame selection can not only reduce the computational cost of videoclassification but also increase the accuracy by getting rid of frames that are hard toclassify. In contrast to previous work, we propose a method that instead of selectingframes by considering one at a time, considers them jointly. This results in a moreefficient selection, where “good" frames are more effectively distributed over thevideo, like snapshots that tell a story. We call the proposed frame selection SMARTand we test it in combination with different backbone architectures and on multiplebenchmarks (Kinetics [5], Something-something [14], UCF101 [31]). We showthat the SMART frame selection consistently improves the accuracy compared toother frame selection strategies while reducing the computational cost by a factorof 4 to 10 times. Additionally, we show that when the primary goal is recognitionperformance, our selection strategy can improve over recent state-of-the-art modelsand frame selection strategies on various benchmarks (UCF101, HMDB51 [21],FCVID [17], and ActivityNet [4]). Shreyank N. Gowda, Marcus Rohrbach, Laura Sevilla-Lara |
AAAI | 1 |
| 2020 | ALBA: Reinforcement Learning for Video Object Segmentation
Shreyank N. Gowda, Panagiotis Eustratiadis, Timothy M. Hospedales, Laura Sevilla-Lara |
BMVC | 1 |
| 2018 | ColorNet: Investigating the Importance of Color Spaces for Image Classification
Shreyank N. Gowda, Chun Yuan 0003 |
ACCV (4) | 1 |