Md. Zarif Hossain

dblp:336/6776 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 SLADE: Shielding against Dual Exploits in Large Vision-Language Models
abstract
Large Vision-Language Models (LVLMs) have emerged as transformative tools in multimodal tasks, seamlessly integrating pretrained vision encoders to align visual and textual modalities. Prior works have highlighted the susceptibility of LVLMs to dual exploits (gradient-based and optimization-based jailbreak attacks), which leverage the expanded attack surface introduced by the image modality. Despite advancements in enhancing robustness, existing methods fall short in their ability to defend against dual exploits while preserving fine-grained semantic details and overall semantic coherence under intense adversarial perturbations. To bridge this gap, we introduce SheiLding against Dual Exploits (SLADE), a novel unsupervised adversarial fine-tuning scheme that enhances the resilience of CLIP-based vision encoders. SLADE’s dual-level contrastive learning approach balances the granular and the holistic, capturing fine-grained image details without losing sight of high-level semantic coherence. Extensive experiments demonstrate that SLADE-equipped LVLMs set a new benchmark for robustness against dual exploits while preserving fine-grained semantic details of perturbed images. Notably, SLADE achieves these results without compromising the core functionalities of LVLMs, such as instruction following, or requiring the computational overhead (e.g., large batch sizes, momentum encoders) commonly associated with traditional contrastive learning methods.
Md. Zarif Hossain, Ahmed Imteaj
CVPR1
2025 Breaking and Securing Vision-Language Models: An Adversarial Robustness Study
abstract
Vision-Language Models (VLMs) have achieved remarkable performance in various downstream tasks, such as image captioning and visual question answering (VQA), by learning joint representations from paired image-text data. However, the robustness of VLMs against adversarial attacks is a critical concern as these models become widely adopted in real-world applications. This study provides a comprehensive overview of the current state of research on the adversarial robustness of VLMs. We introduce a novel taxonomy that categorizes the diverse landscape of adversarial attacks on VLMs, considering factors such as the modality of the attack, the level of model access, and the attack objective. We explore various types of attacks, including jailbreak, backdoor, data poisoning, gradient-based, and patch-based attacks, and discuss their underlying principles and assumptions. Furthermore, we present a comparative analysis of existing defense strategies and their strengths and limitations. By identifying open challenges and future research directions, we aim to stimulate further exploration of this crucial area and contribute to the development of robust and trustworthy VLMs for real-world applications.
Md. Zarif Hossain, Abdur Rahman Bin Shahid, Ahmed Imteaj
ICMLA1
2024 Securing Vision-Language Models with a Robust Encoder Against Jailbreak and Adversarial Attacks
abstract
Large Vision-Language Models (LVLMs), trained on multimodal big datasets, have significantly advanced AI by excelling in vision-language tasks. However, these models remain vulnerable to adversarial attacks, particularly jailbreak attacks, which bypass safety protocols and cause the model to generate misleading or harmful responses. This vulnerability stems from both the inherent susceptibilities of LLMs and the expanded attack surface introduced by the visual modality. We propose SimCLIP+, a novel defense mechanism that adversarially fine-tunes the CLIP vision encoder by leveraging a Siamese architecture. This approach maximizes cosine similarity between perturbed and clean samples, facilitating resilience against adversarial manipulations. Sim-CLIP+ offers a plug-and-play solution, allowing seamless integration into existing LVLM architectures as a robust vision encoder. Unlike previous defenses, our method requires no structural modifications to the LVLM and incurs minimal computational overhead. Sim-CLIP+ demonstrates effectiveness against both gradient-based adversarial attacks and various jailbreak techniques. We evaluate Sim-CLIP+ against three distinct jailbreak attack strategies and perform clean evaluations using standard downstream datasets, including COCO for image captioning and OKVQA for visual question answering. Extensive experiments demonstrate that Sim-CLIP+ maintains high clean accuracy while substantially improving robustness against both gradient-based adversarial attacks and jailbreak techniques.
Md. Zarif Hossain, Ahmed Imteaj
IEEE Big Data1
2024 FedAVO: Improving Communication Efficiency in Federated Learning with African Vultures Optimizer
abstract
Federated Learning (FL) has recently experienced tremendous popularity due to its emphasis on user data privacy. However, the distributed computations of FL can result in constrained communication and drawn-out learning processes, necessitating the client-server communication cost optimization. The ratio of chosen clients and the quantity of local training passes are two hyperparameters that have a significant impact on the performance of FL. Due to different training preferences across various applications, it can be difficult for FL practitioners to manually select such hyperparameters. In this paper, we introduce FedAVO, a novel FL algorithm that enhances communication effectiveness by selecting the best hyperparameters leveraging the African Vulture Optimizer (AVO). Our research demonstrates that the communication costs associated with FL operations can be substantially reduced by adopting AVO for FL hyperparameter adjustment. Through extensive evaluations of FedAVO on benchmark datasets, we identify the optimal hyperparameters that are appropriately fitted for the benchmark datasets, eventually increasing global model accuracy by 6% in comparison to the state-of-the-art FL algorithms (such as FedAvg, FedProx, FedPSO). The code, data, and experiments have been made publicly available on our GitHub repository11https://github.com/speedlab-git/FedAVO.
Md. Zarif Hossain, Ahmed Imteaj, Abdur Rahman Bin Shahid
COMPSAC1
2024 Enhancing Road Safety Through Cost-Effective, Real-Time Monitoring of Driver Awareness with Resource-Constrained IoT Devices
abstract
The prevalence of road and highway accidents, largely attributed to driver distraction, highlights the critical need for an intelligent system that can assess driver alertness and provide timely alerts. Current solutions in the market are often characterized by their high costs and complex installation processes, which significantly limit their accessibility and practical application on a broader scale. In response to this challenge, our research introduces a cost-effective, real-time framework designed to monitor driver alertness utilizing the Raspberry Pi, a device known for its limited processing capabilities. This inherent limitation prompted us to implement several optimizations, which are elaborated upon within our study, to equip the Raspberry Pi with the ability to make real-time decisions effectively. Our proposed approach features a novel algorithm that integrates Haar Cascade and facial landmark detection techniques, enabling the rapid and precise identification of facial features, thereby surpassing the accuracy of existing leading methods. This system meticulously tracks facial points to evaluate driver attentiveness through indicators such as drowsiness, yawning, and unusual facial movements. It utilizes metrics including the Eye Aspect Ratio (EAR), Lips Movement Ratio (LMR), and Face Position Difference (FPD) to initiate driver alerts, thereby contributing to the prevention of potential accidents. Furthermore, our system is enhanced with a GSM module, facilitating emergency notifications to the vehicle owner in critical situations. Extensive testing of our framework, involving participants of diverse sizes, skin colors, and ages, has demonstrated its efficacy in sustaining driver awareness with minimal processing delays, even when deployed on devices with limited computational resources. This affirms the potential of our proposed solution to serve as a viable and scalable option for enhancing road safety through improved driver alertness monitoring.
Ahmed Imteaj, Tanveer Rahman, Saika Zaman, Md. Zarif Hossain, Abdur Rahman Bin Shahid
COMPSAC4
2024 WatchOverGPT: A Framework for Real-Time Crime Detection and Response Using Wearable Camera and Large Language Model
abstract
In the era of Large Language Models (LLMs), the application of advanced AI technologies to data captured by wearables devices, combined with the fusion of contextual data, presents a revolutionary approach to enhancing real-time public safety, individual security, and emergency response. In this paper, we introduce WatchOverGPT, a novel framework that leverages this integration to promptly identify and respond to potential life-threatening criminal activities and safety concerns. WatchOverGPT combines the capabilities of wearable cameras, smartphones' location data, and LLM-based advanced con-versational AI communication through Generative Pre-trained Transformer (GPT). The core of this framework involves a wearable camera connected to the user's smartphone, which continuously captures and analyzes the environment for signs of distress or criminal behaviors, including human actions and the presence of weapons, coupled with location and other information from the smartphone by which GPT-based application provides an autonomous decision-making process. This paper explores the framework's design, implementation, and potential impact of LLM applications on public safety. The proposed framework aims to bridge the gap between safety threats and emergency response teams in the fight against crime through real-time data processing and AI -driven autonomous communication, enhancing the security of individuals in various settings,
Abdur Rahman Bin Shahid, Syed Mhamudul Hasan, Malithi Wanniarachchi Kankanamge, Md. Zarif Hossain, Ahmed Imteaj
COMPSAC4
2024 Towards Communication-Efficient Federated Learning Through Particle Swarm Optimization and Knowledge Distillation
abstract
The widespread popularity of Federated Learning (FL) has led researchers to delve into its various facets, primarily focusing on personalization, fair resource allocation, privacy, and global optimization, with less attention puts towards the crucial aspect of ensuring efficient and cost-optimized communication between the FL server and its agents. A major challenge in achieving successful model training and inference on distributed edge devices lies in optimizing communication costs amid resource constraints, such as limited bandwidth, and selecting efficient agents. In resource-limited FL scenarios, where agents often rely on unstable networks, the transmission of large model weights can substantially degrade model accuracy and increase communication latency between the FL server and agents. Addressing this challenge, we propose a novel strategy that integrates a knowledge distillation technique with a Particle Swarm Optimization (PSO)-based FL method. This approach focuses on transmitting model scores instead of weights, significantly reducing communication overhead and enhancing model accuracy in unstable environments. Our method, with potential applications in smart city services and industrial IoT, marks a significant step forward in reducing network communication costs and mitigating accuracy loss, thereby optimizing the communication efficiency between the FL server and its agents.
Saika Zaman, Sajedul Talukder, Md. Zarif Hossain, Sai Puppala, Ahmed Imteaj
COMPSAC3
2024 TriplePlay: Enhancing Federated Learning with CLIP for Non-IID Data and Resource Efficiency
abstract
The recent advancement of pretrained models shows great potential as well as challenges for privacy-preserving distributed machine learning technique called Federated Learning (FL). With the growing demands of foundation models, it is now an urgent need to explore the potential of such foundation models in a distributed setting. In this paper, In this paper, we delve into the complexities of leveraging foundation models, like CLIP into FL frameworks to preserve data privacy, and efficiently training distributed network clients across heterogeneous data landscapes. We specifically aim to address the issues related to non-IID data distributions, skewed class representation of FL clients' local dataset, communication overhead and high resource consumption due to large, complex model training in an FL setting. To address these, we propose TriplePlay, a framework that tailors CLIP foundation model as an adapter to strengthen FL model's performance and adaptability across heterogeneous data distributions among the clients. Besides, we address the long-tail distribution problem in an FL environment to maintain fairness and optimize the computational resource demands of the FL clients through quantization and low-rank adaptation techniques. A comprehensive simulations results with two distinct datasets and different FL settings demonstrate that TriplePlay efficiently reduces GPU usage and accelerates the convergence time that ultimately reduces the communication cost.
Ahmed Imteaj, Md. Zarif Hossain, Saika Zaman, Abdur Rahman Bin Shahid
ICMLA2
2023 Assessing Wearable Human Activity Recognition Systems Against Data Poisoning Attacks in Differentially-Private Federated Learning
abstract
Differentially-Private Federated Learning (DPFL) is an emerging privacy-preserving distributed machine learning paradigm that allows for the automatic recognition of human activities using wearable sensors without compromising users’ sensitive data. However, this decentralized approach makes the system vulnerable to poisoning attacks, where malicious agents can inject contaminated data during local model training. This paper presents the results of our research on designing, developing, and evaluating a holistic model for data poisoning attacks in DPFL-based human activity recognition (HAR) systems. Specifically, we focus on label-flipping poisoning attacks, where the label of a sensor reading is maliciously changed during data collection. To investigate the impact of such attacks, we develop a simulator that explores key design issues, such as the correlation between the level of differential privacy, the level of poisoning, the number of communication rounds, and the number of agents in the system. Our findings shed light on the effectiveness of label contamination attacks in DPFL-based HAR systems and can inform the development of more robust and secure models.
Abdur Rahman Bin Shahid, Ahmed Imteaj, Shahriar Badsha, Md. Zarif Hossain
SMARTCOMP4