VLDB 2026 Research / reviewers in the wild / expert
Muhammad Sohaib Ayub
dblp:134/5998
· DBLP profile ↗
12ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0001-9206-1545ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MaskedVerbalizer: Automatic Verbalizer Construction for Few-Shot Text Classification in Low-Resource Right-to-Left Languages
Faizad Ullah, Furqan Sikandar, Areeba Waqar, Faizan Ali, Muhammad Sohaib Ayub, Mubashar Mushtaq, Asim Karim |
LREC | 5 |
| 2025 | Clustering-Based Balance Phenotyping in Older Adults from Treadmill Training Interventions
Shafiq Alam, Imran Khan Niazi, Hina Shafi, Waqar Ahmed Awan, Imran Amjad, Muhammad Sohaib Ayub, Mufti Mahmud |
IEEE Big Data | 6 |
| 2025 | Evaluating Explainable AI Implementation and User Agency Across Major Social Media PlatformsabstractAI-driven recommendation algorithms increasingly shape user experience on social media, raising concerns about transparency, accountability, and user agency. This paper presents a comparative analysis of Explainable AI (XAI) implementations on Facebook, Instagram, TikTok, and Twitter/X. Using a structured framework, we assess two key dimensions: Explanation Adequacy, defined by clarity, specificity, relevance, and verifiability, and Explanation Actionability, defined by proximity, granularity, and reversibility of controls. Our evaluation combines feature audits, cross-platform comparisons, and rubric-based scoring. Results show Facebook provides the strongest balance of adequacy (4/5) and actionability (4/5), Twitter/X offers limited adequacy (2/5) but moderate actionability (3/5), TikTok demonstrates strong actionability (4/5) but generic explanations, and Instagram performs moderately ($3 / 5$on both dimensions). We further extend the analysis by linking adequacy and actionability to perceived usefulness, trust, satisfaction, and algorithmic scepticism. Findings highlight tensions between algorithmic sophistication, user comprehension, and engagement optimization, while also revealing regulatory implications under GDPR and the DSA. This work contributes a standardized evaluation framework for XAI in social computing, empirical evidence of platform disparities, and practical design guidelines for enhancing transparency and user empowerment in recommender systems. Shafiq Alam, Aditya Pawade, Muhammad Sohaib Ayub, Saeed Ur Rehman 0001, Asma Ayub |
IEEE Big Data | 3 |
| 2025 | Abstractive Summarization for Urdu Video Description GenerationabstractAutomatic summarization condenses content while retaining key ideas and details. Urdu, with over 230 million speakers globally, is one of the most widely spoken languages. The rise of Urdu content on social media platforms has driven the need for tools that enhance accessibility and engagement. The growing popularity of social media has increased the number of Urdu instructional videos. Well-written video descriptions can boost viewer engagement and improve search engine optimization; however, many lack these. Therefore, an automatic description generation system for Urdu videos is needed, which can be achieved by abstractive summarization of video transcripts. However, such public datasets are not available in Urdu. To address this problem, we investigate the usability of high-resource language datasets for Urdu abstractive text summarization. We created the first Urdu video transcription dataset Urdu How2 and evaluated its quality using intrinsic evaluation. We leverage transfer learning, a technique where knowledge from pretrained models (like mT5) is adapted to new tasks, to develop the uT5 model for generating Urdu text summaries. We further trained the model to improve its Urdu text generation capability. The machine-generated summaries are evaluated using ROUGE scores, human evaluation scores, and adversarial evaluation, providing a reliable assessment of the quality of generated descriptions and the robustness of the model against noisy text data. The human evaluation shows the proposed method generates accurate and coherent summaries compared to the translated ground truth. To the best of our knowledge, this is the first attempt to utilize a cross-lingual dataset for Urdu abstractive text summarization and video description generation. This research enhances Urdu content accessibility and lays the groundwork for advancing multilingual content generation and multimodal analysis in other low-resource languages. Ali Faheem, Faizad Ullah, Muhammad Sohaib Ayub, Asim Karim |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2024 | UrduMASD: A Multimodal Abstractive Summarization Dataset for UrduabstractIn this era of multimedia dominance, the surge of multimodal content on social media has transformed our methods of communication and information exchange. With the widespread use of multimedia content, the ability to effectively summarize this multimodal content is crucial for enhancing consumption, searchability, and retrieval. The scarcity of such training datasets has been a barrier to research in this area, especially for low-resource languages like Urdu. To address this gap, this paper introduces “UrduMASD”, a video-based Urdu multimodal abstractive text summarization dataset. The dataset contains 15,374 collections of videos, audio, titles, transcripts, and corresponding text summaries. To ensure the quality of the dataset, intrinsic evaluation metrics such as Abstractivity, Compression, Redundancy, and Semantic coherence have been employed. It was observed that our dataset surpasses existing datasets on numerous key quality metrics. Additionally, we present baseline results achieved using both text-based and state-of-the-art multimodal summarization models. On adding visual information, an improvement of 2.6% was observed in the ROUGE scores, highlighting the efficacy of utilizing multimodal inputs for summarization. To the best of our knowledge, this is the first dataset in Urdu that provides video-based multimodal data for abstractive text summarization, making it a valuable resource for advancing research in this field. Ali Faheem, Faizad Ullah, Muhammad Sohaib Ayub, Asim Karim |
LREC/COLING | 3 |
| 2024 | Detecting Cybercrimes in Accordance with Pakistani Law: Dataset and Evaluation Using PLMsabstractCybercrime is a serious and growing threat affecting millions of people worldwide. Detecting cybercrimes from text messages is challenging, as it requires understanding the linguistic and cultural nuances of different languages and regions. Roman Urdu is a widely used language in Pakistan and other South Asian countries, however, it lacks sufficient resources and tools for natural language processing and cybercrime detection. To address this problem, we make three main contributions in this paper. (1) We create and release CRU, a benchmark dataset for text-based cybercrime detection in Roman Urdu, which covers a number of cybercrimes as defined by the Prevention of Electronic Crimes Act (PECA) of Pakistan. This dataset is annotated by experts following a standardized procedure based on Pakistan’s legal framework. (2) We perform experiments on four pre-trained language models (PLMs) for cybercrime text classification in Roman Urdu. Our results show that xlm-roberta-base is the best model for this task, achieving the highest performance on all metrics. (3) We explore the utility of prompt engineering techniques, namely prefix and cloze prompts, for enhancing the performance of PLMs for low-resource languages such as Roman Urdu. We analyze the impact of different prompt shapes and k-shot settings on the performance of xlm-roberta-base and bert-base-multilingual-cased. We find that prefix prompts are more effective than cloze prompts for Roman Urdu classification tasks, as they provide more contextually relevant completions for the models. Our work provides useful insights and resources for future research on cybercrime detection and text classification in low-resource languages. Faizad Ullah, Ali Faheem, Ubaid Azam, Muhammad Sohaib Ayub, Faisal Kamiran, Asim Karim |
LREC/COLING | 4 |
| 2024 | Enhanced Facial Emotion Detection Models Utilizing Geometry-Based Features for Superior Human-Computer Interaction
Shafiq Alam, Muhammad Sohaib Ayub, Rohan Sathasivam, Muhammad Asad Khan |
ICONIP (10) | 2 |
| 2023 | Towards Developing an Automated Chatbot for Predicting Legal Case Outcomes: A Deep Learning Approach
Shafiq Alam, Rohit Pande, Muhammad Sohaib Ayub, Muhammad Asad Khan |
ACIIDS (1) | 3 |
| 2023 | Sequence-Based Nanobody-Antigen Binding Prediction
Usama Sardar, Sarwan Ali, Muhammad Sohaib Ayub, Khurram Bashir, Murray Patterson |
ISBRA | 3 |
| 2017 | Experience Report: Verifying MPI Java Programs Using Software Model CheckingabstractParallel and distributed computing have enabled development of much more scalable software. However, developing concurrent software requires the programmer to be aware of nondeterminism, data races, and deadlocks. MPI (message passing interface) is a popular standard for writing message-oriented distributed applications. Some messages in MPI systems can be processed by one of the many machines and in many possible orders. This non-determinism can affect the result of an MPI application. The alternate results may or may not be correct. To verify MPI applications, we need to check all these possible orderings and use an application specific oracle to decide if these orderings give correct output. MPJ Express is an open source Java implementation of the MPI standard. Model checking of MPI Java programs is a challenging task due to their parallel nature. We developed a Java based model of MPJ Express, where processes are modeled as threads, and which can run unmodified MPI Java programs on a single system. This model enabled us to adapt the Java PathFinder explicit state software model checker (JPF) using a custom listener to verify our model running real MPI Java programs. The evaluation of our approach shows that model checking reveals incorrect system behavior that results in very intricate message orderings. Muhammad Sohaib Ayub, Waqas ur Rehman, Junaid Haroon Siddiqui |
ISSRE | 1 |
| 2016 | Verification of MPI Java programs using software model checkingabstractDevelopment of concurrent software requires the programmer to be aware of non-determinism, data races, and deadlocks. MPI (message passing interface) is a popular standard for writing message oriented distributed applications. Some messages in MPI systems can be processed by one of the many machines and in many possible orders. This non-determinism can affect the result of an MPI application. The alternate results may or may not be correct. To verify MPI applications, we need to check all these possible orderings and use an application specific oracle to decide if these orderings give correct output. MPJ Express is an open source Java implementation of the MPI standard. We developed a Java based model of MPJ Express, where processes are modeled as threads, and which can run unmodified MPI Java programs on a single system. This enabled us to adapt the Java PathFinder explicit state software model checker (JPF) using a custom listener to verify our model running real MPI Java programs. We evaluated our approach using small examples where model checking revealed message orders that would result in incorrect system behavior. Waqas ur Rehman, Muhammad Sohaib Ayub, Junaid Haroon Siddiqui |
PPoPP | 2 |
| 2012 | Towards Efficient Support for Parallel I/O in Java HPCabstractModern HPC applications put forward significant I/O requirements. To deal with them, MPI provides the MPI-IO API for parallel file access. ROMIO library implements MPI-IO and provides efficient support for parallel I/O in C and Fortran based applications. On the other hand, Java based MPI-like libraries such as MPJ Express and F-MPJ have emerged but they lack parallel I/O support. Little research has been done to provide Java based ROMIO-like libraries due to the non-availability of MPI-IO-like API for the Java language. In this paper, we take the first step towards the development of parallel I/O API in Java by evaluating the newly introduced Java NIO API versus the legacy Java I/O API. We propose two simple approaches for performing parallel file I/O using NIO and evaluate them on two different computational platforms. The implementation of proposed approaches exploits the view buffers concept of NIO API to perform efficient array based file I/O operations from multiple processes. We report encouraging speedups and suggest that design of a parallel I/O API in Java should be based on the NIO API. Ammar Ahmad Awan, Muhammad Sohaib Ayub, Aamir Shafi, Sungyoung Lee 0001 |
PDCAT | 2 |