VLDB 2026 Research / reviewers in the wild / expert
Muhammad Uzair Khattak
dblp:324/2256
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Vision and language · 37% Transfer learning and domain adaptation · 22% Efficient and distributed learning · 11% |
Topics — the 19 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
vision-language model adaptation |
1.7 | 3 | 2025 | Learning to Prompt with Text Only Supervision for Vision-Language Models · AAAI 2025 MaPLe: Multi-modal Prompt Learning · CVPR 2023 Self-regulating Prompts: Foundational Model Adaptation without Forgetting · ICCV 2023 |
Machine learning › Transfer learning and domain adaptation
zero-shot transfer |
1.7 | 3 | 2025 | Learning to Prompt with Text Only Supervision for Vision-Language Models · AAAI 2025 Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization · NeurIPS 2023 Bridging the Gap between Object and Image-level Representations for Open-Vocabulary Detection · NeurIPS 2022 |
Computer vision › Vision and language › vision-language model
prompt learning |
1.5 | 2 | 2025 | Learning to Prompt with Text Only Supervision for Vision-Language Models · AAAI 2025 Self-regulating Prompts: Foundational Model Adaptation without Forgetting · ICCV 2023 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
1.3 | 2 | 2023 | Self-regulating Prompts: Foundational Model Adaptation without Forgetting · ICCV 2023 MaPLe: Multi-modal Prompt Learning · CVPR 2023 |
Computer vision › Video understanding and tracking
action recognition |
0.7 | 1 | 2023 | Video-FocalNets: Spatio-Temporal Focal Modulation for Video Action Recognition · ICCV 2023 |
Machine learning › Trustworthy machine learning › robustness
distribution shift |
0.7 | 1 | 2023 | Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization · NeurIPS 2023 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › visual domain adaptation
image-to-video adaptation |
0.7 | 1 | 2023 | Fine-tuned CLIP Models are Efficient Video Learners · CVPR 2023 |
Computer vision › Vision and language
multimodal prompt learning |
0.7 | 1 | 2023 | MaPLe: Multi-modal Prompt Learning · CVPR 2023 |
Natural language and speech › Language models and text generation
prompt tuning |
0.7 | 1 | 2023 | MaPLe: Multi-modal Prompt Learning · CVPR 2023 |
Machine learning › Transfer learning and domain adaptation › test-time adaptation
test-time prompt tuning |
0.7 | 1 | 2023 | Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization · NeurIPS 2023 |
Computer vision › Video understanding and tracking
video classification |
0.7 | 1 | 2023 | Fine-tuned CLIP Models are Efficient Video Learners · CVPR 2023 |
Computer vision › Vision and language
vision-language model |
0.7 | 1 | 2023 | Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization · NeurIPS 2023 |
Computer vision › Image recognition and object detection › object detection
open-vocabulary object detection |
0.6 | 1 | 2022 | Bridging the Gap between Object and Image-level Representations for Open-Vocabulary Detection · NeurIPS 2022 |
Computer vision › Vision and language › cross-modal alignment › visual-semantic alignment
vision-language representation alignment |
0.6 | 1 | 2022 | Bridging the Gap between Object and Image-level Representations for Open-Vocabulary Detection · NeurIPS 2022 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
LLM distillation |
0.3 | 1 | 2025 | Learning to Prompt with Text Only Supervision for Vision-Language Models · AAAI 2025 |
Machine learning › Transfer learning and domain adaptation
generalization to unseen classes |
0.2 | 1 | 2023 | MaPLe: Multi-modal Prompt Learning · CVPR 2023 |
Computer vision › Image recognition and object detection
image classification |
0.2 | 1 | 2023 | MaPLe: Multi-modal Prompt Learning · CVPR 2023 |
Machine learning › Deep learning architectures and training
transformer |
0.2 | 1 | 2023 | Video-FocalNets: Spatio-Temporal Focal Modulation for Video Action Recognition · ICCV 2023 |
Computer vision › Vision and language
vision-language pretraining |
0.2 | 1 | 2023 | Fine-tuned CLIP Models are Efficient Video Learners · CVPR 2023 |
Methods — techniques the papers use, named apart from their topics
prompt learning · 1.3prompt ensembling · 0.9knowledge distillation · 0.9self-regularization · 0.7self-ensemble · 0.7mutual agreement maximization · 0.7feature pooling · 0.7convolution · 0.7contrastive vision-language pretraining · 0.7CLIP fine-tuning · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning to Prompt with Text Only Supervision for Vision-Language ModelsabstractFoundational vision-language models like CLIP are emerging as a promising paradigm in vision due to their excellent generalization. However, adapting these models for downstream tasks while maintaining their generalization remains challenging. In literature, one branch of methods adapts CLIP by learning prompts using images. While effective, these methods often rely on image-label data, which is not always practical, and struggle to generalize to new datasets due to overfitting on few-shot source data. Another approach explores training-free methods by generating class captions from large language models (LLMs) and performing prompt ensembling, but these methods often produce static, class-specific prompts that cannot be transferred to new classes and incur additional costs by generating LLM descriptions for each class separately. In this work, we aim to combine the strengths of both approaches by learning prompts using only text data derived from LLMs. As supervised training of prompts in the image-free setup is non-trivial, we develop a language-only efficient training approach that enables prompts to distill rich contextual knowledge from LLM data. Furthermore, by mapping the LLM contextual text data within the learned prompts, our approach enables zero-shot transfer of prompts to new classes and datasets, potentially reducing the LLM prompt engineering cost. To the best of our knowledge, this is the first work that learns generalized and transferable prompts for image tasks using only text data. We perform evaluations on 4 benchmarks, where ProText improves over ensembling methods while being competitive with those using labeled images. Muhammad Uzair Khattak, Muhammad Ferjad Naeem, Muzammal Naseer, Luc Van Gool, Federico Tombari |
AAAI | 1 |
| 2023 | MaPLe: Multi-modal Prompt LearningabstractPre-trained vision-language (V-L) models such as CLIP have shown excellent generalization ability to downstream tasks. However, they are sensitive to the choice of input text prompts and require careful selection of prompt templates to perform well. Inspired by the Natural Language Processing (NLP) literature, recent CLIP adaptation approaches learn prompts as the textual inputs to fine-tune CLIP for downstream tasks. We note that using prompting to adapt representations in a single branch of CLIP (language or vision) is sub-optimal since it does not allow the flexibility to dynamically adjust both representation spaces on a downstream task. In this work, we propose Multi-modal Prompt Learning (MaPLe) for both vision and language branches to improve alignment between the vision and language representations. Our design promotes strong coupling between the vision-language prompts to ensure mutual synergy and discourages learning independent uni-modal solutions. Further, we learn separate prompts across different early stages to progressively model the stage-wise feature relationships to allow rich context learning. We evaluate the effectiveness of our approach on three representative tasks of generalization to novel classes, new target datasets and unseen domain shifts. Compared with the state-of-the-art method Co-CoOp, MaPLe exhibits favorable performance and achieves an absolute gain of 3.45% on novel classes and 2.72% on overall harmonic-mean, averaged over 11 diverse image recognition datasets. Our code and pre-trained models are available at https://github.com/muzairkhattak/multimodal-prompt-learning. Muhammad Uzair Khattak, Hanoona Abdul Rasheed, Muhammad Maaz 0001, Salman Khan 0001, Fahad Shahbaz Khan |
CVPR | 1 |
| 2023 | Fine-tuned CLIP Models are Efficient Video LearnersabstractLarge-scale multi-modal training with image-text pairs imparts strong generalization to CLIP model. Since training on a similar scale for videos is infeasible, recent approaches focus on the effective transfer of image-based CLIP to the video domain. In this pursuit, new parametric modules are added to learn temporal information and inter-frame relationships which require meticulous design efforts. Furthermore, when the resulting models are learned on videos, they tend to overfit on the given task distribution and lack in generalization aspect. This begs the following question: How to effectively transfer image-level CLIP representations to videos? In this work, we show that a simple Video Fine-tuned CLIP (ViFi-CLIP) baseline is generally sufficient to bridge the domain gap from images to videos. Our qualitative analysis illustrates that the frame-level processing from CLIP image-encoder followed by feature pooling and similarity matching with corresponding text embeddings helps in implicitly modeling the temporal cues within ViFi-CLIP. Such fine-tuning helps the model to focus on scene dynamics, moving objects and inter-object relationships. For low-data regimes where full fine-tuning is not viable, we propose a ‘bridge and prompt’ approach that first uses fine-tuning to bridge the domain gap and then learns prompts on language and vision side to adapt CLIP representations. We extensively evaluate this simple yet strong baseline on zero-shot, base-to-novel generalization, few-shot and fully supervised settings across five video benchmarks. Our code and pre-trained models are available at https://github.com/muzairkhattak/ViFi-CLIP. Hanoona Abdul Rasheed, Muhammad Uzair Khattak, Muhammad Maaz 0001, Salman Khan 0001, Fahad Shahbaz Khan |
CVPR | 2 |
| 2023 | Self-regulating Prompts: Foundational Model Adaptation without ForgettingabstractPrompt learning has emerged as an efficient alternative for fine-tuning foundational models, such as CLIP, for various downstream tasks. Conventionally trained using the task-specific objective, i.e., cross-entropy loss, prompts tend to overfit downstream data distributions and find it challenging to capture task-agnostic general features from the frozen CLIP. This leads to the loss of the model’s original generalization capability. To address this issue, our work introduces a self-regularization framework for prompting called PromptSRC (Prompting with Self-regulating Constraints). PromptSRC guides the prompts to optimize for both task-specific and task-agnostic general representations using a three-pronged approach by: (a) regulating prompted representations via mutual agreement maximization with the frozen model, (b) regulating with self-ensemble of prompts over the training trajectory to encode their complementary strengths, and (c) regulating with textual diversity to mitigate sample diversity imbalance with the visual branch. To the best of our knowledge, this is the first regularization framework for prompt learning that avoids overfitting by jointly attending to pre-trained model features, the training trajectory during prompting, and the textual diversity. PromptSRC explicitly steers the prompts to learn a representation space that maximizes performance on downstream tasks without compromising CLIP generalization. We perform extensive experiments on 4 benchmarks where PromptSRC overall performs favorably well compared to the existing methods. Our code and pre-trained models are publicly available at: https://github.com/muzairkhattak/PromptSRC. Muhammad Uzair Khattak, Syed Talal Wasim, Muzammal Naseer, Salman Khan 0001, Ming-Hsuan Yang 0001, Fahad Shahbaz Khan |
ICCV | 1 |
| 2023 | Video-FocalNets: Spatio-Temporal Focal Modulation for Video Action RecognitionabstractRecent video recognition models utilize Transformer models for long-range spatio-temporal context modeling. Video transformer designs are based on self-attention that can model global context at a high computational cost. In comparison, convolutional designs for videos offer an efficient alternative but lack long-range dependency modeling. Towards achieving the best of both designs, this work proposes Video-FocalNet, an effective and efficient architecture for video recognition that models both local and global contexts. Video-FocalNet is based on a spatiotemporal focal modulation architecture that reverses the interaction and aggregation steps of self-attention for better efficiency. Further, the aggregation step and the interaction step are both implemented using efficient convolution and element-wise multiplication operations that are computationally less expensive than their self-attention counterparts on video representations. We extensively explore the design space of focal modulation-based spatiotemporal context modeling and demonstrate our parallel spatial and temporal encoding design to be the optimal choice. Video-FocalNets perform favorably well against the state-of-the-art transformer-based models for video recognition on five large-scale datasets (Kinetics-400, Kinetics-600, SS-v2, Diving-48, and ActivityNet-1.3) at a lower computational cost. Our code/models are released at https://github.com/TalalWasim/Video-FocalNets. Syed Talal Wasim, Muhammad Uzair Khattak, Muzammal Naseer, Salman Khan 0001, Mubarak Shah, Fahad Shahbaz Khan |
ICCV | 2 |
| 2023 | Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot GeneralizationabstractThe promising zero-shot generalization of vision-language models such as CLIP has led to their adoption using prompt learning for numerous downstream tasks. Previous works have shown test-time prompt tuning using entropy minimization to adapt text prompts for unseen domains. While effective, this overlooks the key cause for performance degradation to unseen domains -- distribution shift. In this work, we explicitly handle this problem by aligning the out-of-distribution (OOD) test sample statistics to those of the source data using prompt tuning. We use a single test sample to adapt multi-modal prompts at test time by minimizing the feature distribution shift to bridge the gap in the test domain. Evaluating against the domain generalization benchmark, our method improves zero-shot top-1 accuracy beyond existing prompt-learning techniques, with a 3.08% improvement over the baseline MaPLe. In cross-dataset generalization with unseen categories across 10 datasets, our method improves consistently across all datasets compared to the existing state-of-the-art. Our source code and models are available at [https://jameelhassan.github.io/promptalign](https://jameelhassan.github.io/promptalign) Jameel Abdul Samadh, Hanan Gani, Noor Hussein, Muhammad Uzair Khattak, Muzammal Naseer, Fahad Shahbaz Khan, Salman Khan 0001 |
NeurIPS | 4 |
| 2022 | Bridging the Gap between Object and Image-level Representations for Open-Vocabulary DetectionabstractExisting open-vocabulary object detectors typically enlarge their vocabulary sizes by leveraging different forms of weak supervision. This helps generalize to novel objects at inference. Two popular forms of weak-supervision used in open-vocabulary detection (OVD) include pretrained CLIP model and image-level supervision. We note that both these modes of supervision are not optimally aligned for the detection task: CLIP is trained with image-text pairs and lacks precise localization of objects while the image-level supervision has been used with heuristics that do not accurately specify local object regions. In this work, we propose to address this problem by performing object-centric alignment of the language embeddings from the CLIP model. Furthermore, we visually ground the objects with only image-level supervision using a pseudo-labeling process that provides high-quality object proposals and helps expand the vocabulary during training. We establish a bridge between the above two object-alignment strategies via a novel weight transfer function that aggregates their complimentary strengths. In essence, the proposed model seeks to minimize the gap between object and image-centric representations in the OVD setting. On the COCO benchmark, our proposed approach achieves 36.6 AP50 on novel classes, an absolute 8.2 gain over the previous best performance. For LVIS, we surpass the state-of-the-art ViLD model by 5.0 mask AP for rare categories and 3.4 overall. Code: https://github.com/hanoonaR/object-centric-ovd. Hanoona Abdul Rasheed, Muhammad Maaz 0001, Muhammad Uzair Khattak, Salman Khan 0001, Fahad Shahbaz Khan |
NeurIPS | 3 |