Abdul Basit 0013

dblp:28/1807-13 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0005-9890-5565ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 MetaCipher: A Time-Persistent and Universal Multi-Agent Framework for Cipher-Based Jailbreak Attacks for LLMs
abstract
Large language models (LLMs) face persistent vulnerability to jailbreak attacks despite their increasing capabilities. While developers deploy alignment finetuning and safety guardrails, researchers consistently devise novel attacks that circumvent these defenses. This dynamic mirrors a strategic game of continual evolution. However, two challenges hinder jailbreak development: the high cost of querying top-tier LLMs and the short lifespan of effective attacks due to frequent safety updates. These factors limit cost-efficiency and impact. To address this, we propose MetaCipher, a low-cost, multi-agent jailbreak framework that generalizes across LLMs with varying safety measures. Using reinforcement learning, MetaCipher is modular and adaptive, supporting extensibility to future strategies. Within as few as 10 queries, MetaCipher achieves state-of-the-art attack success rates on recent malicious prompt benchmarks, outperforming prior jailbreak methods. We conduct a large-scale empirical evaluation across diverse victim models, demonstrating its robustness and adaptability.
Boyuan Chen 0004, Minghao Shao, Abdul Basit 0013, Siddharth Garg, Muhammad Shafique 0001
AAAI3
2026 PatchBlock: A Lightweight Defense Against Adversarial Patches for Embedded EdgeAI Devices
abstract
Adversarial attacks pose a significant challenge to the reliable deployment of machine learning models in EdgeAI applications, such as autonomous driving and surveillance, which rely on resource-constrained devices for real-time inference. Among these, patch-based adversarial attacks, where small malicious patches (e.g., stickers) are applied to objects, can deceive neural networks into making incorrect predictions with potentially severe consequences. In this paper, we present PatchBlock, a lightweight framework designed to detect and neutralize adversarial patches in images. Leveraging outlier detection and dimensionality reduction, PatchBlock identifies regions affected by adversarial noise and suppresses their impact. It operates as a pre-processing module at the sensor level, efficiently running on CPUs in parallel with GPU inference, thus preserving system throughput while avoiding additional GPU overhead. The framework follows a three-stage pipeline: splitting the input into chunks (Chunking), detecting anomalous regions via a redesigned isolation forest with targeted cuts for faster convergence (Separating), and applying dimensionality reduction on the identified outliers (Mitigating). PatchBlock is both model- and patch-agnostic, can be retrofitted to existing pipelines, and integrates seamlessly between sensor inputs and downstream models. Evaluations across multiple neural architectures, benchmark datasets, attack types, and diverse edge devices demonstrate that PatchBlock consistently improves robustness, recovering up to 77% of model accuracy under strong patch attacks such as the Google Adversarial Patch, while maintaining high portability and minimal clean accuracy loss. Additionally, PatchBlock outperforms the state-of-the-art defenses in efficiency, in terms of computation time and energy consumption per sample, making it suitable for EdgeAI applications.
Nandish Chattopadhyay, Abdul Basit 0013, Amira Guesmi, Muhammad Abdullah Hanif, Bassem Ouni, Muhammad Shafique 0001
DATE2
2026 Focus Session: Hardware and Software Techniques for Accelerating Multimodal Foundation Models
abstract
This work presents a multi-layered methodology for efficiently accelerating multimodal foundation models (MFMs). It combines hardware and software co-design of transformer blocks with an optimization pipeline that reduces computational and memory requirements. During model development, it employs performance enhancements through fine-tuning for domain-specific adaptation. Our methodology further incorporates hardware and software techniques for optimizing MFMs. Specifically, it employs MFM compression using hierarchy-aware mixed-precision quantization and structural pruning for transformer blocks and MLP channels. It also optimizes operations through speculative decoding, model cascading that routes queries through a small-to-large cascade and uses lightweight self-tests to determine when to escalate to larger models, as well as co-optimization of sequence length, visual resolution & stride, and graph-level operator fusion. To efficiently execute the model, the processing dataflow is optimized based on the underlying hardware architecture together with memory-efficient attention to meet on-chip bandwidth and latency budgets. To support this, a specialized hardware accelerator for the transformer workloads is employed, which can be developed through expert design or an LLM-aided design approach. We demonstrate the effectiveness of the proposed methodology on medical-MFMs and on code generation tasks, and conclude with extensions toward energy-efficient spiking-MFMs.
Muhammad Shafique 0001, Abdul Basit 0013, Muhammad Abdullah Hanif, Alberto Marchisio, Rachmad Vidya Wicaksana Putra, Minghao Shao
DATE2
2025 CognitiveArm: Enabling Real-Time EEG-Controlled Prosthetic Arm Using Embodied Machine Learning
abstract
Efficient control of prosthetic limbs via non-invasive brain-computer interfaces (BCIs) requires advanced EEG processing capabilities-including pre-filtering, feature extraction, and action pre-diction-all performed in real-time on edge AI hardware. Achieving this level of real-time processing on resource-constrained edge devices presents significant challenges in balancing model complexity, computational efficiency, and latency. We present CognitiveArm, an EEG-driven, brain-controlled prosthetic system implemented on edge AI hardware, achieving real-time performance without compromising accuracy. The system integrates BrainFlow-an open-source library for EEG data acquisition and streaming-and optimized deep learning (DL) models for precise brain signal classification. By leveraging evolutionary search, we identify Pareto-optimal DL model configurations through hyper-parameter tuning, optimizer analysis, and window selection, analyzed individually and in ensemble configurations. We further apply model compression techniques like pruning and quantization to optimize these models for embedded deployment, balancing computational efficiency and accuracy. We collected EEG dataset and designed an annotation pipeline, enabling precise labeling of brain signals corresponding to specific intended actions, which forms the foundation for training our optimized deep learning (DL) models. Our CognitiveArm system also supports voice commands for seamless mode switching, enabling control of the prosthetic arm’s 3 degrees of freedom (DoF). Running independently on embedded hardware, CognitiveArm ensures low latency and facilitates real-time interaction. We developed a full-scale prototype of the CognitiveArm, interfaced with the OpenBCI UltraCortex Mark IV EEG headset. Our evaluations demonstrate a significant improvement in accuracy, reaching up to $\mathbf{9 0 \%}$, for classifying three core actions (left, right, and stay idle). The integration of voice commands allows for multiplexed, variable movement, enabling multi-action control for various everyday tasks (e.g., handshake, cup picking). This enhances CognitiveArm’s real-world performance for prosthetic control, demonstrating its potential as a practical solution for individuals requiring advanced prosthetic limb control.
Abdul Basit 0013, Maha Nawaz, Saim Rehman, Muhammad Shafique 0001
DAC1
2025 BRAVE: Brain-Controlled Prosthetic Arm with Voice Integration and Embodied Learning for Enhanced Mobility
abstract
Non-invasive brain-computer interfaces (BCIs) have the potential to enable intuitive control of prosthetic limbs for individuals with upper limb amputations. However, existing EEG-based control systems face challenges related to signal noise, classification accuracy, and real-time adaptability. In this work, we present BRAVE, a hybrid EEG and voice-controlled prosthetic system that integrates ensemble learning-based EEG classification with a human-in-the-loop (HITL) correction framework for enhanced responsiveness. Unlike traditional electromyography (EMG)-based prosthetic control, BRAVE aims to interpret EEG-driven motor intent, enabling movement control without reliance on residual muscle activity. To improve classification robustness, BRAVE combines Long Short-Term Memory (LSTM), Convolutional Neural Network (CNN), and Random Forest models in an ensemble framework, achieving a classification accuracy of 96% across test subjects. EEG signals are preprocessed using a bandpass filter (0.5–45 Hz), Independent Component Analysis (ICA) for artifact removal, and Common Spatial Pattern (CSP) feature extraction to minimize contamination from electromyographic (EMG) and electrooculographic (EOG) signals. Additionally, BRAVE incorporates automatic speech recognition (ASR) to facilitate intuitive mode switching between different degrees of freedom (DOF) in the prosthetic arm. The system operates in real time, with a response latency of 150 ms, leveraging Lab Streaming Layer (LSL) networking for synchronized data acquisition. The system is evaluated on an in-house fabricated prosthetic arm and on multiple participants highlighting the generalizability across users. The system is optimized for low-power embedded deployment, ensuring practical real-world application beyond high-performance computing environments. Our results indicate that BRAVE offers a promising step towards robust, real-time, non-invasive prosthetic control.
Abdul Basit 0013, Maha Nawaz, Muhammad Shafique 0001
IJCNN1
2025 ROVER: Autonomous Open-Vocabulary Object Searching in Unexplored Environments Using VLM-Driven Scene Understanding
abstract
Autonomous robots operating in unstructured and unknown environments require efficient object searching capabilities to enable real-world applications such as service robotics, warehouse automation, and disaster response. Traditional object detection models are inherently limited to predefined categories, making them unsuitable for open-ended object search tasks where the target object is specified dynamically. In this work, we propose ROVER, an end-to-end open-vocabulary object searching framework which we implement and test on the AgileX LIMO UGV in real-world environments. ROVER integrates frontier-based exploration, RTABMap SLAM for real-time localization and dynamic navigation, and language-driven object detection in unexplored environments. The system autonomously explores, dynamically detects objects specified via natural language input, and navigates precisely to the detected object’s location using depth-based coordinate estimation. Unlike prior works, our approach is fully deployable on resource-constrained mobile robots and optimized for real-time execution on embedded hardware. Our ported model on the Orin Nano edge platform runs at around a Hz. Extensive real-world experiments demonstrate ROVER achieves 100% precision/recall for small object detection (screws/kettles), 91.3% chair detection precision in cluttered environments, and an average 95.5% F1-score, outperforming YOLO-World baselines by 13.1 percentage points using Grounding DINO in real-demonstration settings. Unlike prior works limited to simulations, our framework demonstrates full operational capability on physical robots with open-vocabulary robotic perception, proving language-driven object navigation in dynamic and unexplored settings.
Abdul Basit 0013, Niraj Pudasaini, Muhammad Shafique 0001
IJCNN1
2024 RoboMed: On-Premise Medical Assistance Leveraging Large Language Models in Robotics
abstract
Large language models (LLMs) are revolutionizing numerous domains with their remarkable natural language processing (NLP) capabilities, attracting significant interest and widespread adoption. However, deploying LLMs in resource-constrained environments, such as edge computing and robotics systems without server infrastructure, while also aiming to minimize latency, presents significant challenges. Another challenge lies in delivering medical assistance to remote areas with limited healthcare facilities and infrastructure. To address this, we introduce RoboMed, an on-premise healthcare robot that utilizes compact versions of large language models (tiny-LLMs) integrated with LangChain as its backbone. Moreover, it incorporates automatic speech recognition (ASR) models for user interface, enabling efficient, edge-based preliminary medical diagnostics and support. RoboMed employs model optimizations to achieve minimal memory footprint and reduced latency during inference on embedded edge devices. The training process optimization involves low-rank adaptation (LoRA), which reduces the model's complexity without significantly impacting its performance. For fine-tuning, the LLM is trained on a diverse medical dataset compiled from online health forums, clinical case studies, and a distilled medicine corpus. This fine-tuning process utilizes reinforcement learning from human feedback (RLHF) to further enhance its domain-specific capabilities. The system is deployed on Nvidia Jetson development board and achieves 78% accuracy in medical consultations and scores 56 in USMLE benchmark, enabling an resource-efficient healthcare assistance robot that alleviates privacy concerns due to edge-based deployment, thereby empowering the community.
Abdul Basit 0013, Khizar Hussain, Muhammad Abdullah Hanif, Muhammad Shafique 0001
ICARCV1
2024 MindArm: Mechanized Intelligent Non-Invasive Neuro-Driven Prosthetic Arm System
abstract
Currently, individuals with arm mobility impairments (referred to as ‘'patients') face limited technological solutions due to two key challenges: (1) non-invasive prosthetic devices are often prohibitively expensive and costly to maintain, and (2) invasive solutions require high-risk, costly brain surgery, which can pose a health risk. Therefore, current technological solutions are not accessible for all patients with different financial backgrounds. Toward this, we propose a low-cost technological solution called MindArm, an affordable, non-invasive neuro-driven prosthetic arm system. MindArm employs a deep neural network (DNN) to translate brain signals, captured by low-cost surface electroencephalogram (EEG) electrodes, into prosthetic arm movements. Utilizing an Open Brain Computer Interface and UDP networking for signal processing, the system seamlessly controls arm motion. In the compute module, we run a trained DNN model to interpret filtered micro-voltage brain signals, and then translate them into a prosthetic arm action via serial communication seamlessly. Experimental results from a fully functional prototype show high accuracy across three actions, with 91% for idle/stationary, 85% for handshake, and 84% for cup pickup. The system costs approximately $500-550, including $400 for the EEG headset and $100-150 for motors, 3D printing, and assembly, offering an affordable alternative for mind-controlled prosthetic devices.
Maha Nawaz, Abdul Basit 0013, Muhammad Shafique 0001
ICARCV2
2024 tinyDigiClones: A Multi-Modal LLM-Based Framework for Edge-optimized Personalized Avatars
abstract
Conversational AI has made significant strides, however the integration of multi-modal interactions, particularly on edge devices, presents a novel frontier to be explored primarily due to the computational constraints. This paper proposes the tinyDigiClones framework, which enables communication with a personalized AI assistant leveraging optimized large language models (LLMs) for natural language processing (NLP), and deep-learning models for automatic speech recognition (ASR) and realistic voice synthesis. This paper explores various options for different AI models employed in our framework, with a primary focus on ensuring efficient deployment on edge devices while maintaining high accuracy. To replicate users’ voice fonts and learn the unique vocal characteristics, the Text-to-Speech (TTS) models are trained using a custom dataset of audio-text pairs. It is generated automatically by the ASR module which segments extended sentences into shorter, transcript-matched audio files. Moreover, deploying state-of-the-art LLMs on resource-constrained devices presents a significant challenge, particularly in maintaining minimal latency, given their extensive parameter counts. Towards this, we explore several lightweight LLMs and employ optimization techniques aimed at reducing computational costs. The integration of these models is personified through a digital avatar mirroring the user’s facial and voice likeness, offering an immersive experience. Deployment on the edge alleviates the server latency and enhances privacy enabling real-time interaction capabilities of AI chatbots, ideal for interactive digital avatars.
Abdul Basit 0013, Muhammad Shafique 0001
IJCNN1