EDBT 2026 Demo / reviewers in the wild / expert
Sara Atito Ali Ahmed
dblp:248/3915 · also Sara Atito 0001, Sara Atito Ali, Sara Atito Aly
· DBLP profile ↗
31ranked-venue papers
6as first author
31since 2021 · last 2026
0000-0002-7576-5791ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 3 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 19 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DAT3: Dual-Teacher Topology Adversarial Training for Defending Against Adversarial Attacks
Huaxing Feng, He-Feng Yin, Sara Atito Ali Ahmed, Muhammad Awais 0001 |
ICPR (11) | 5 |
| 2026 | Channel-Aware Probing for Multi-channel Imaging
Umar Marikkar, Syed Sameed Husain, Muhammad Awais 0001, Sara Atito Ali Ahmed |
ICPR (7) | 4 |
| 2026 | Gaze-Guided Multimodal LLMs for Social Scene Understanding
Shayan Nasiriboukani, Muhammad Awais 0001, Sara Atito Ali Ahmed |
ICPR (11) | 3 |
| 2026 | CoZSR-VAD: Contextual Zero-Shot Reasoning for Video Anomaly Detection
Mohd Ubaid Wani, Sara Atito Ali Ahmed, Srinivasa Rao Nandam, Josef Kittler, Muhammad Awais 0001 |
ICPR (12) | 2 |
| 2026 | Probabilistically Aligned View-Unaligned Clustering With Adaptive Template SelectionabstractIn most existing multi-view modeling scenarios, cross-view correspondence (CVC) between instances of the same target from different views, like paired image-text data, is a crucial prerequisite for effortlessly deriving a consistent representation. Nevertheless, this premise is frequently compromised in certain applications, where each view is organized and transmitted independently, resulting in the view-unaligned problem (VuP). Restoring CVC of unaligned multi-view data is a challenging and highly demanding task that has received limited attention from the research community. To tackle this practical challenge, we propose to integrate the permutation derivation procedure into the bipartite graph paradigm for view-unaligned clustering, termed Probabilistically Aligned View-unaligned Clustering with Adaptive Template Selection (PAVuC-ATS). Specifically, we learn consistent anchors and view-specific graphs by the bipartite graph, and derive permutations applied to the unaligned graphs by reformulating the alignment between two latent representations as a 2-step transition of a Markov chain with adaptive template selection, thereby achieving the probabilistic alignment. The convergence of the resultant optimization problem is validated both experimentally and theoretically. Extensive experiments on six benchmark datasets demonstrate the superiority of the proposed PAVuC-ATS over the baseline methods. Wenhua Dong, Xiaojun Wu 0001, Zhenhua Feng 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | DeepChest: Dynamic Gradient-Free Task Weighting for Effective Multi-Task Learning in Chest X-Ray Classification
Youssef Mohamed, Noran Mohamed, Khaled Abouhashad, Sara Atito Ali Ahmed, Shoaib Jameel, Muhammad Imran Razzak, Ahmed B. Zaky |
IEEE Big Data | 5 |
| 2025 | One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image FusionabstractAdvanced image fusion methods mostly prioritise high-level missions, where task interaction struggles with semantic gaps, requiring complex bridging mechanisms. In contrast, we propose to leverage low-level vision tasks from digital photography fusion, allowing for effective feature interaction through pixel-level supervision. This new paradigm provides strong guidance for unsupervised multimodal fusion without relying on abstract semantics, enhancing task-shared feature learning for broader applicability. Owning to the hybrid image features and enhanced universal representations, the proposed GIFNet supports diverse fusion tasks, achieving high performance across both seen and unseen scenarios with a single model. Uniquely, experimental results reveal that our framework also supports single-modality enhancement, offering superior flexibility for practical applications. Our code will be available at https://github.com/AWCXV/GIFNet. Chunyang Cheng, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Zhangyong Tang, Hui Li 0037, Zeyang Zhang 0002, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler |
CVPR | 8 |
| 2025 | Text Augmented Correlation Transformer For Few-shot Classification & SegmentationabstractFoundation models like CLIP and ALIGN have transformed few-shot and zero-shot vision applications by fusing visual and textual data, yet the integrative few-shot classification and segmentation (FS-CS) task primarily leverages visual cues, overlooking the potential of textual support. In FS-CS scenarios, ambiguous object boundaries and overlapping classes often hinder model performance, as limited visual data struggles to fully capture high-level semantics. To bridge this gap, we present a novel multi-modal FS-CS framework that integrates textual cues into support data, facilitating enhanced semantic disambiguation and fine-grained segmentation. Our approach first investigates the unique contributions of exclusive text-based support, using only class labels to achieve FS-CS. This strategy alone achieves performance competitive with vision-only methods on FS-CS tasks, underscoring the power of textual cues in few-shot learning. Building on this, we introduce a dualmodal prediction mechanism that synthesizes insights from both textual and visual support sets, yielding robust multimodal predictions. This integration significantly elevates FS-CS performance, with classification and segmentation improvements of +3.7/6.6% (1-way 1-shot) and +8.0/6.5% (2-way 1-shot) on COCO-20i, and +2.2/3.8% (1-way 1shot) and +4.3/4.0% (2-way 1-shot) on Pascal-5i. Additionally, in weakly supervised FS-CS settings, our method surpasses visual-only benchmarks using textual support exclusively, further enhanced by our dual-modal predictions. By rethinking the role of text in FS-CS, our work establishes new benchmarks for multi-modal few-shot learning and demonstrates the efficacy of textual cues for improving model generalization and segmentation accuracy. Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
CVPR | 2 |
| 2025 | Enhanced Weakly Supervised Few-shot Classification & SegmentationabstractThe emergence of vision-language foundation models has enabled the integration of textual information into vision-based applications. However, in few-shot classification and segmentation (FS-CS), this potential remains underutilised. Commonly, self-supervised vision models have been employed, particularly in weakly-supervised scenarios, to generate pseudo-segmentation masks, as ground truth masks are typically unavailable and only target classification is provided. Despite their success, such models find it difficult to capture accurate semantics when compared to vision-language models. To address this limitation, we propose a novel FS-CS approach that leverages the rich semantic alignment of vision-language models to generate more precise pseudo ground-truth masks. While current vision-language models excel in global visual-text alignment, they struggle with finer, patch-level alignment, which is crucial for detailed segmentation tasks. To overcome this, we introduce a method that enhances patch-level alignment without requiring additional training. In addition, existing FS-CS frameworks typically lacks multi-scale information, limiting their ability to capture fine and coarse features simultaneously. To overcome this, we incorporate a module based on atrous convolutions to inject multi-scale information into the feature maps. Together, these contributions - text enhanced pseudo-mask generation and improved multi-scale feature representation - significantly boost the performance of our model in weakly-supervised settings, surpassing state-of-the-art methods and demonstrating the importance of integrating multi-modal information for robust FS-CS solutions. Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
ICASSP | 2 |
| 2025 | SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic SoundscapesabstractSelf-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed in a frozen state, under the assumption that the self-supervised pre-training has sufficiently equipped them to handle real-world audio. However, a critical question remains: how well do these models actually perform in real-world conditions, where audio is typically polyphonic and complex, involving multiple overlapping sound sources? Current audio self-supervised learning (SSL) methods are often benchmarked on datasets predominantly featuring monophonic audio, such as environmental sounds, and speech. As a result, the ability of SSL models to generalize to polyphonic audio, a common characteristic in natural scenarios, remains underexplored. This limitation raises concerns about the practical robustness of SSL models in more realistic audio settings. To address this gap, we introduce Self-Supervised Learning from Audio Mixtures (SSLAM), a novel direction in audio SSL research, designed to improve the model’s ability to learn from polyphonic data while maintaining strong performance on monophonic data. We thoroughly evaluate SSLAM on standard audio SSL benchmark datasets which are predominantly monophonic and conduct a comprehensive comparative analysis against state-of-the-art (SOTA) methods using a range of high-quality, publicly available polyphonic datasets. SSLAM not only improves model performance on polyphonic audio, but also maintains or exceeds performance on standard audio SSL benchmarks. Notably, it achieves up to a 3.9% improvement on the AudioSet-2M(AS-2M), reaching a mean average precision (mAP) of 50.2. For polyphonic datasets, SSLAM sets new SOTA in both linear evaluation and fine-tuning regimes with performance improvements of up to 9.1%(mAP). These results demonstrate SSLAM's effectiveness in both polyphonic and monophonic soundscapes, significantly enhancing the performance of audio SSL models. Code and pre-trained models are available at https://github.com/ta012/SSLAM. Tony Alex, Sara Atito Ali Ahmed, Armin Mustafa, Muhammad Awais 0001, Philip J. B. Jackson |
ICLR | 2 |
| 2025 | CG-SSL: Concept-Guided Self-Supervised LearningabstractHumans understand visual scenes by first capturing a global impression and then refining this understanding into distinct, object-like components. Inspired by this process, we introduce \textbf{C}oncept-\textbf{G}uided \textbf{S}elf-\textbf{S}upervised \textbf{L}earning (CG-SSL), a novel framework that brings structure and interpretability to representation learning through a curriculum of three training phases: (1) global scene encoding, (2) discovery of visual concepts via tokenised cross-attention, and (3) alignment of these concepts across views.
Unlike traditional SSL methods, which simply enforce similarity between multiple augmented views of the same image, CG-SSL accounts for the fact that these views may highlight different parts of an object or scene. To address this, our method establishes explicit correspondences between views and aligns the representations of meaningful image regions. At its core, CG-SSL augments standard SSL with a lightweight decoder that learns and refines concept tokens via cross-attention with patch features. The concept tokens are trained using masked concept distillation and a feature-space reconstruction objective. A final alignment stage enforces view consistency by geometrically matching concept regions under heavy augmentation, enabling more compact, robust, and disentangled representations of scene regions.
Across multiple backbone sizes, CG-SSL achieves state-of-the-art results on image segmentation benchmarks using $k$-NN and linear probes, substantially outperforming prior methods and approaching, or even surpassing, the performance of leading SSL models trained on over $100\times$ more data. Code and pretrained models will be released. Sara Atito Ali Ahmed, Josef Kittler, Muhammad Imran Razzak, Muhammad Awais 0001 |
NeurIPS | 1 |
| 2025 | Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM CaptioningabstractDespite significant advances in inference-time search for vision–language models (VLMs), existing approaches remain both computationally expensive and prone to unpenalized, low-confidence generations which often lead to persistent hallucinations. We introduce \textbf{Value-guided Inference with Margin-based Reward (ViMaR)}, a two-stage inference framework that improves both efficiency and output fidelity by combining a temporal-difference value model with a margin-aware reward adjustment. In the first stage, we perform a single pass to identify the highest-value caption among diverse candidates. In the second stage, we selectively refine only those segments that were overlooked or exhibit weak visual grounding, thereby eliminating frequently rewarded evaluations. A calibrated margin-based penalty discourages low-confidence continuations while preserving descriptive richness. Extensive experiments across multiple VLM architectures demonstrate that ViMaR generates captions that are significantly more reliable, factually accurate, detailed, and explanatory, while achieving over 4$\times$ speedup compared to existing value-guided methods. Specifically, we show that ViMaR trained solely on LLaVA Mistral-7B \textit{generalizes effectively to guide decoding in stronger unseen models}. To further validate this, we adapt ViMaR to steer generation in both LLaVA-OneVision-Qwen2-7B and Qwen2.5-VL-3B, leading to consistent improvements in caption quality and demonstrating robust cross-model guidance. This cross-model generalization highlights ViMaR's flexibility and modularity, positioning it as a scalable and transferable inference-time decoding strategy. Furthermore, when ViMaR-generated captions are used for self-training, the underlying models achieve substantial gains across a broad suite of visual comprehension benchmarks, underscoring the potential of fast, accurate, and self-improving VLM pipelines.
Code: https://github.com/ankan8145/ViMaR Ankan Deria, Adinath Madhavrao Dukre, Sara Atito Ali Ahmed, Sudipta Roy 0002, Muhammad Awais 0001, Muhammad Haris Khan, Muhammad Imran Razzak |
NeurIPS | 4 |
| 2025 | Investigating Self-Supervised Methods for Label-Efficient LearningabstractAbstract Vision transformers combined with self-supervised learning have enabled the development of models which scale across large datasets for several downstream tasks, including classification, segmentation, and detection. However, the potential of these models for low-shot learning across several downstream tasks remains largely under explored. In this work, we conduct a systematic examination of different self-supervised pretext tasks, namely contrastive learning, clustering, and masked image modelling, to assess their low-shot capabilities by comparing different pretrained models. In addition, we explore the impact of various collapse avoidance techniques, such as centring, ME-MAX, and sinkhorn, on these downstream tasks. Based on our detailed analysis, we introduce a framework that combines mask image modelling and clustering as pretext tasks. This framework demonstrates superior performance across all examined low-shot downstream tasks, including multi-class classification, multi-label classification and semantic segmentation. Furthermore, when testing the model on large-scale datasets, we show performance gains in various tasks. Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Correction: Investigating Self-Supervised Methods for Label-Efficient Learning
Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Which images can be effectively learnt from self-supervised learning?
Michalis Lazarou, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler |
Pattern Recognit. Lett. | 2 |
| 2024 | Pseudo Labelling for Enhanced Masked Auto Encoders
Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
BMVC | 2 |
| 2024 | C2C: Component-to-Composition Learning for Zero-Shot Compositional Action Recognition
Rongchang Li 0001, Zhenhua Feng 0001, Tianyang Xu 0001, Linze Li 0002, Xiaojun Wu 0001, Muhammad Awais 0001, Sara Atito Ali Ahmed, Josef Kittler |
ECCV (38) | 7 |
| 2024 | Improved Image Captioning Via Knowledge Graph-Augmented ModelsabstractMultimodal foundation models, pre-trained on large-scale data, effectively capture vast amounts of factual and commonsense knowledge. However, these models store all their knowledge within their parameters, requiring increasingly larger models and training data to capture more knowledge. To address this limitation and achieve a more scalable and modular integration of knowledge, we propose a novel knowledge graph-augmented multimodal model. This approach enables a base multimodal model to access pertinent information from an external knowledge graph. Our methodology leverages existing general domain knowledge to facilitate vision-language pre-training using paired images and text descriptions. We conduct comprehensive evaluations demonstrating that our model outperforms state-of-the-art models and yields comparable results to much larger models trained on more extensive datasets. Notably, our model reached a 145 Cider score on MS COCO Captions using only 2.9 million samples, outperforming a 1.4B parameter model by 1.7% despite having 11 times fewer parameters. Sergio Sánchez Santiesteban, Sara Atito Ali Ahmed, Muhammad Awais 0001, Yi-Zhe Song, Josef Kittler |
ICASSP | 2 |
| 2024 | SS-CXR: Self-Supervised Pretraining Using Chest X-Rays Towards A Domain Specific Foundation ModelabstractChest X-rays (CXRs) are widely used imaging modality for the diagnosis and prognosis of lung disease. There is a large body of work where machine learning algorithms are developed for specific tasks. However, the traditional diagnostic tool design methods based on supervised learning are burdened by the need to provide training data annotation, which should be of good quality for better clinical outcomes. Here, we propose an alternative solution, a new self-supervised paradigm, where a general representation from CXRs is learned using a group-masked self-supervised framework. The pre-trained model is then fine-tuned for domain-specific tasks such as covid-19, pneumonia detection, and general health screening. We show that the same pre-training can be used for the lung segmentation task. Our proposed paradigm shows robust performance in multiple downstream tasks which demonstrates the success of the pre-training. Moreover, the performance of the pre-trained models on data with significant drift during test time proves the learning of a better generic representation. The methods are further validated by covid-19 detection in a unique small-scale pediatric data set. The performance gain ($\sim 25 \%$) is significant when compared to a supervised transformer-based method. This adds credence to the strength and reliability of our proposed framework and pre-training strategy. Syed Muhammad Anwar, Abhijeet Parida, Sara Atito Ali Ahmed, Muhammad Awais 0001, Gustavo Nino, Josef Kittler, Marius George Linguraru |
ICIP | 3 |
| 2024 | Investigating Self-Supervised Methods for Label-Efficient LearningabstractVision transformers combined with self-supervised learning have enabled the development of models which scale across large datasets for several downstream tasks like classification, segmentation and detection. The low-shot learning capability of these models, across several low-shot downstream tasks, has been largely under explored. We perform a system level study of different self supervised pretext tasks, namely contrastive learning, clustering, and masked image modelling for their low-shot capabilities by comparing the pretrained models. In addition we also study the effects of collapse avoidance methods, namely centring, ME-MAX, sinkhorn, on these downstream tasks. Based on our detailed analysis, we introduce a framework involving both mask image modelling and clustering as pretext tasks, which performs better across all low-shot downstream tasks, including multi-class classification, multi-label classification and semantic segmentation. Furthermore, when testing the model on full scale datasets, we show performance gains in multi-class classification, multi-label classification and semantic segmentation. Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001 |
ICIP | 2 |
| 2024 | Masked Momentum Contrastive Learning for Semantic Understanding by ObservationabstractLarge language models (LLMs) have shown excellent performance in zero-shot learning using natural language prompts. However, in the domain of computer vision (CV), the paradigm of pretraining followed by finetuning remains dominant. The aim of this study is to reduce this gap by utilizing the capability of Self-Supervised Learning (SSL) in semantic understanding for zero-shot segmentation, without relying on human-provided labels or vision-language supervision. We introduce a novel evaluation framework that employs visual prompts, including a threshold and a query patch. This framework evaluates the ability of SSL models to derive concepts from observational data. Through this evaluation, we identify the strengths and limitations of SSL models in understanding semantics. Building on the insights from various SSL methods, we further propose the MMC approach to enhance the representations for objects, which integrates Masked image modeling, Momentum-based self-distillation, and global Contrastive learning. MMC achieves a better balance between the inter-object discriminability and the intra-object compactness of learned features. Our experiments on COCO, DAVIS-2017, PASCAL VOC, and ADE20K demonstrate outstanding performance of MMC’s representations. Jiantao Wu, Shentong Mo, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Syed Sameed Husain, Muhammad Awais 0001 |
ICIP | 3 |
| 2024 | View-shuffled clustering via the modified Hungarian algorithm
Wenhua Dong, Xiaojun Wu 0001, Tianyang Xu 0001, Zhenhua Feng 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler |
Neural Networks | 5 |
| 2024 | Self-supervised learning for RGB-D object tracking
Xuefeng Zhu 0003, Tianyang Xu 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Xiaojun Wu 0001, Zhenhua Feng 0001, Josef Kittler |
Pattern Recognit. | 3 |
| 2024 | ASiT: Local-Global Audio Spectrogram Vision Transformer for Event ClassificationabstractTransformers, which were originally developed for natural language processing, have recently generated significant interest in the computer vision and audio communities due to their flexibility in learning long-range relationships. Constrained by the data hungry nature of transformers and the limited amount of labelled data, most transformer-based models for audio tasks are finetuned from ImageNet pretrained models, despite the huge gap between the domain of natural images and audio. This has motivated the research in self-supervised pretraining of audio transformers, which reduces the dependency on large amounts of labeled data and focuses on extracting concise representations of audio spectrograms. In this paper, we proposeLocal-GlobalAudioSpectrogram vIsionTransformer, namely ASiT, a novel self-supervised learning framework that captures local and global contextual information by employing group masked model learning and self-distillation. We evaluate our pretrained models on both audio and speech classification tasks, including audio event classification, keyword spotting, and speaker identification. We further conduct comprehensive ablation studies, including evaluations of different pretraining strategies. The proposed ASiT framework significantly boosts the performance on all tasks and sets a new state-of-the-art performance in five audio and speech classification tasks, outperforming recent methods, including the approaches that use additional datasets for pretraining. Sara Atito Ali Ahmed, Muhammad Awais 0001, Wenwu Wang 0001, Mark D. Plumbley, Josef Kittler |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2024 | One-pass View-unaligned ClusteringabstractGiven a set of multi-view instances, the prevailing assumption in most existing clustering approaches is that they are complete and exhibit cross-view alignment. However, this assumption is often unrealistic. In such scenarios, it could be satisfied at the cost of data pre-processing, but this would be complex and inconsistent with practical applications. Therefore, developing more effective solutions for the View-unaligned Problem (VuP) is highly desirable. Several pioneering works have tackled the partially VuP, yet handling fully VuP remains a challenge due to the reliance on partially pre-aligned instances. In this paper, we propose One-pass View-unaligned Clustering (OpVuC) that simultaneously aligns and clusters instances in a unified framework. Specifically, we alig shuffled instances with a selected template using an innovative global-local alignment scheme based on the notion of geometric invariance and separate the fully aligned instances using a relaxed$k$-means algorithm. The proposed OpVuC method can handle VuP at any alignment level without requiring any pre-aligned instances. Extensive experiments conducted on several benchmark datasets demonstrate the effectiveness and merits of the proposed OpVuC method. Wenhua Dong, Xiaojun Wu 0001, Zhenhua Feng 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler |
IEEE Trans. Multim. | 4 |
| 2023 | Group Masked Model Learning for General Audio RepresentationabstractVision transformers have recently generated significant interest in the computer vision and audio communities due to their flexibility in learning long-range relationships. However, transformers are known to be data hungry which require orders of magnitude more data [1] to train. This has motivated the research in self-supervised pretraining of audio transformers, which reduces the dependency on large amounts of labeled data and focuses on extracting concise representation of the audio spectrograms. In this paper, we propose Audio-GMML, a self-supervised transformer for general audio representations that is based on Group Masked Model Learning (GMML) and a patch aggregation strategy to improve the performance of learned representations and enforce global structure of the given audio. We evaluate our pretrained models on several downstream tasks, setting a new state-of-the-art performance on five audio and speech classification tasks. The code and pretrained weights will be made publicly available for the scientific community. Sara Atito Ali Ahmed, Muhammad Awais 0001, Tony Alex, Josef Kittler |
ICIP | 1 |
| 2023 | GMML is All You NeedabstractVision transformers (ViTs) have generated significant interest in the computer vision community because of their flexibility in exploiting contextual information, whether it is sharply confined local, or long range global. However, they are known to be data hungry and therefore often pretrained on large-scale datasets, e.g. JFT-300M or ImageNet. An ideal learning method would perform best regardless of the size of the dataset, a property lacked by current learning methods, with merely a few existing works studying ViTs with limited data. We propose Group Masked Model Learning (GMML), a self-supervised learning (SSL) method that is able to train ViTs and achieve state-of-the-art (SOTA) performance when pre-trained with limited data. The GMML uses the information conveyed by all concepts in the image. This is achieved by manipulating randomly groups of connected tokens, successively covering different meaningful parts of the image content, and then recovering the hidden information from the visible part of the concept. Unlike most of the existing SSL approaches, GMML does not require momentum encoder, nor relies on careful implementation details such as large batches and gradient stopping. Pretraining, finetuning, and evaluation codes are available under: https://github.com/GMML. Sara Atito Ali Ahmed, Muhammad Awais 0001, Srinivasa Rao Nandam, Josef Kittler |
ICIP | 1 |
| 2023 | LT-ViT: A Vision Transformer for Multi-Label Chest X-Ray ClassificationabstractVision Transformers (ViTs) are widely adopted in medical imaging tasks, and some existing efforts have been directed towards vision-language training for Chest X-rays (CXRs). However, we envision that there still exists a potential for improvement in vision-only training for CXRs using ViTs, by aggregating information from multiple scales, which has been proven beneficial for non-transformer networks. Hence, we have developed LT-ViT, a transformer that utilizes combined attention between image tokens and randomly initialized auxiliary tokens that represent labels. Our experiments demonstrate that LT-ViT (1) surpasses the state-of-the-art performance using pure ViTs on two publicly available CXR datasets, (2) is generalizable to other pre-training methods and therefore is agnostic to model initialization, and (3) enables model interpretability without grad-cam and its variants. Umar Marikkar, Sara Atito Ali Ahmed, Muhammad Awais 0001, Adam Mahdi |
ICIP | 2 |
| 2022 | Relative attributes classification via transformers and rank SVM lossabstractWe propose a new model for learning to rank two images with respect to their relative strength of expression for a given attribute. We address this problem – called relative attribute learning — using a vision transformer backbone. The embedded representations of the two images to be compared are extracted and used for comparison with a ranking head, in an end-to-end fashion. The results demonstrate the strength of vision transformers and their suitability for relative attributes classification. Our proposed approach outperforms the state-of-the-art by a large margin, achieving 90.40% and 98.14% mean accuracy over the attributes of LFW-10 and Pubfig datasets. Sara Atito Ali Ahmed, Berrin A. Yanikoglu |
ICMV | 1 |
| 2022 | Face attribute classification with evidential deep learningabstractWe address the problem of uncertainty quantification in the domain of face attribute classification, using Evidential Deep Learning (EDL) framework. The proposed EDL approach leverages the strength of Convolution Neural Networks (CNN), with the objective of representing the uncertainty in the output predictions. Predominantly, the softmax/sigmoid activation functions are applied to map the output logits of the CNN to target class probabilities in multi-class classification problems. By replacing the standard softmax/sigmoid output of a CNN with the parameters of the evidential distribution, EDL learns to represent the uncertainty in its predictions. The proposed approach is evaluated on CelebA and LFWA datasets. The quantitative and qualitative analysis demonstrate the suitability and strength of EDL to estimate the uncertainty in the output predictions without hindering the accuracy of CNN-based models. Arin Zeyneloglu, Sara Atito Ali Ahmed, Berrin A. Yanikoglu |
ICMV | 2 |
| 2022 | Comparison and ensemble of 2D and 3D approaches for COVID-19 detection in CT images
Sara Atito Ali Ahmed, Mehmet Can Yavuz, Mehmet Umut Sen, Fatih Gulsen, Onur Tutar, Bora Korkmazer, Cesur Samanci, Sabri Sirolu, Rauf Hamid, Ali Ergun Eryurekli, Toghrul Mammadov, Berrin A. Yanikoglu |
Neurocomputing | 1 |