VLDB 2026 Research / reviewers in the wild / expert
Md Fahim
dblp:311/0390
· DBLP profile ↗
15ranked-venue papers
6as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BanHADEX: Towards Explainable HAte Speech Detection in Bangla Using Human Annotated EXplanationabstractFaisal Hossain Raquib, Akm Moshiur Rahman Mazumder, Md Fahim, Md Tahmid Hasan Fuad, Md Farhan Ishmam, Faria Sultana, M Ashraful Amin, Amin Ahsan Ali, Akmmahbubur Rahman. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Faisal Hossain Raquib, Akm Moshiur Rahman Mazumder, Md Fahim, Md. Tahmid Hasan Fuad, Md Farhan Ishmam, Faria Sultana, M. Ashraful Amin, Amin Ahsan Ali, Akmmahbubur Rahman |
ACL (1) | 3 |
| 2026 | Embodied Referring Expression Comprehension in Human-Robot InteractionabstractAs robots enter human workspaces, there is a crucial need for them to comprehend embodied human instructions, enabling intuitive and fluent human-robot interaction (HRI). However, accurate comprehension is challenging due to a lack of large-scale datasets that capture natural embodied interactions in diverse HRI settings. Existing datasets suffer from perspective bias, single-view data collection, inadequate coverage of nonverbal gestures, and a predominant focus on indoor environments. To address these issues, we present the Refer360 dataset, a large-scale dataset of embodied verbal and nonverbal interactions collected across diverse viewpoints in both indoor and outdoor settings. Additionally, we introduce MuRes, a multimodal guided residual module designed to improve embodied referring expression comprehension. MuRes acts as an information bottleneck, extracting salient modality-specific signals and reinforcing them into pre-trained representations to form complementary features for downstream tasks. We conduct extensive experiments on four datasets, including our Refer360 dataset, and demonstrate that current multimodal models fail to capture embodied interactions comprehensively; however, augmenting them with MuRes consistently improves performance. These findings establish Refer360 as a valuable benchmark and exhibit the potential of guided residual learning to advance embodied referring expression comprehension in robots operating within human environments. Md. Mofijul Islam, Alexi Gladstone, Sujan Sarker, Ganesh Nanduru, Md Fahim, Keyan Du, Aman Chadha, Tariq Iqbal |
HRI | 5 |
| 2026 | PRiSM: Partial Ranking via Inter-layer Semantic Measurement for Efficient Fine-tuning of Language Models
Aldrin Kabya Biswas, Md Fahim, M. Ashraful Amin, Amin Ahsan Ali, A. K. M. Mahbubur Rahman |
LREC | 2 |
| 2026 | R-MMA: Enhancing Vision-Language Models with Recurrent Adapters for Few-Shot and Cross-Domain GeneralizationabstractPre-trained vision-language models (VLMs) such as CLIP exhibit strong generalization but struggle with few-shot adaptation due to the trade-off between gaining task-specific knowledge and preserving general performance. While multimodal adapters add trainable modules that improve alignment and excel in few-shot generality, they greatly increase the trainable parameter count while relying heavily on the prior layer’s frozen representation. Addressing these limitations, we introduce Recurrent Multi-Modal Adapter (R-MMA), a lightweight and efficient adapter that uses self-attention to compute a unified latent representation with a single set of shared adapter weights. Our attention-based alignment harmonizes the adapter outputs with the frozen encoder features before fusing the modalities, ensuring better preservation of pre-trained representations and cross-modal consistency. Our experiments show that R-MMA achieves state-of-the-art performance on most datasets for base-to-novel generalization, cross-dataset evaluation, and domain generalization, under few-shot settings. Our approach also achieves one of the highest forms of parameter efficiency with only a few trainable weight matrices for the whole network, regardless of its depth. Our code is available at: https://github.com/farhanishmam/R-MMA. Md Fahim, Md Farhan Ishmam, Mir Sazzat Hossain, M. Ashraful Amin, Amin Ahsan Ali, A. K. M. Mahbubur Rahman |
WACV | 1 |
| 2026 | BanglaProtha: Evaluating Vision Language Models in Underrepresented Long-tail Cultural ContextsabstractThe advanced multimodal processing of current vision language models (VLMs) has prompted rigorous benchmarking across multicultural settings, revealing a clear inclination toward Western culture. While the bias likely stems from the predominance of Western-centric images in the VLM pretraining data, the resulting long-tail distribution problem is only exacerbated in underrepresented cultural settings, such as Bengali. Our work explores this problem through an aspect-based evaluation of several classes of VLMs on the rich Bengali culture. Our BanglaProtha dataset is a VQA dataset, containing images that encapsulate Bengali cultural elements, questions in native Bengali, and semantically similar multiple-choice answer options. Our experiments provide behavioral insights into VLMs across prompting & fine-tuning strategies, cultural aspects, model size, and augmentation methods. Our work serves as a diagnostic tool for addressing and mitigating inequalities in multicultural and multilingual settings, thereby bringing efforts to democratize AI systems. Our code and data are available at https://github.com/farhanishmam/BanglaProtha. Md Fahim, Md Sakib Ul Rahman Sourove, Akm Moshiur Rahman, Md Farhan Ishmam, Md Tasmim Rahman Adib, Fariha Tanjim Shifat, Fabiha Haider, Farhad Alam Bhuiyan |
WACV | 1 |
| 2026 | Training-free layer selection for partial fine-tuning of language modelsabstractThe growing scale of pre-trained language models poses a challenge in fine-tuning for downstream tasks, especially in resource-constrained settings. Recent studies highlight that not all layers in Transformer-based language models contribute equally to downstream task performance, giving rise to various partial fine-tuning strategies. We propose a training-free approach for layer-wise partial fine-tuning that leverages the cosine similarity between representative tokens across layers to identify inter-layer relationships. Our method comprises two stages: (i) scoring layers based on their relevance to the task via a single forward pass, and (ii) fine-tuning a subset of layers, either highest-scoring, lowest-scoring, or block-wise, while keeping others frozen. We conduct experiments on 16 diverse NLP datasets, including single-sentence and sentence-pair classification tasks, as well as generation tasks. Our method achieves competitive performance compared to full fine-tuning, with an average training speedup of 1.5 and a reduction of trainable parameters by 75%, and outperforms all comparative baselines in 14 out of 16 evaluated datasets. Additionally, our approach does not cause any notable drop in performance when the domain is changed for the evaluation tasks, demonstrating a robust cross-domain performance. • Efficient training-free layer selection uses token cosine similarity. • Inter-layer token relationships identify optimal layers for fine-tuning. • Reduces trainable parameters by 75% with a 1.5x training speedup. • Our paper yields better accuracy on 16 diverse datasets compared to the SOTA works. • Layer selection strategy preserves robust cross-domain generalization. Aldrin Kabya Biswas, Md Fahim, Md. Tahmid Hasan Fuad, Akm Moshiur Rahman Mazumder, Amin Ahsan Ali, A. K. M. Mahbubur Rahman |
Inf. Sci. | 2 |
| 2025 | BANMIME : Misogyny Detection with Metaphor Explanation on Bangla MemesabstractMd Ayon Mia, Akm Moshiur Rahman Mazumder, Khadiza Sultana Sayma, Md Fahim, Md Tahmid Hasan Fuad, Muhammad Ibrahim Khan, Akmmahbubur Rahman. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Md Ayon Mia, Akm Moshiur Rahman Mazumder, Khadiza Sultana Sayma, Md Fahim, Md. Tahmid Hasan Fuad, Muhammad Ibrahim Khan, A. K. M. Mahbubur Rahman |
EMNLP | 4 |
| 2025 | ChitroJera: A Regionally Relevant Visual Question Answering Dataset for Bangla
Deeparghya Dutta Barua, Md Sakib Ul Rahman Sourove, Md Fahim, Fabiha Haider, Fariha Tanjim Shifat, Md Tasmim Rahman Adib, Anam Borhan Uddin, Md Farhan Ishmam, Md. Farhad Alam |
ECML/PKDD (4) | 3 |
| 2024 | Improving the Performance of Transformer-Based Models Over Classical Baselines in Multiple Transliterated LanguagesabstractSocial media users express their feelings, experiences, ideas, and stories with little or no regard for the conventions of traditional grammar. Online discourse, by its very nature, is rife with transliterated text along with code-mixing and code-switching. Transliteration is heavily featured due to the ease of inputting romanized text with standard keyboards over native scripts. Due to its ubiquity, it is a critical area of study to ensure NLP models perform well in real-world scenarios. In this paper, we analyze the performance of various language models, Tiny Large Language models, TF-IDF and Bag-of-Words feature extraction-based classical ML models, as well as zero-shot classification with ChatGPT on romanized/transliterated social media text. We chose the tasks of sentiment analysis and offensive language identification and we carried out experiments for three different languages, namely Bangla, Hindi, and Arabic, for six datasets. To our surprise, we discovered across multiple datasets that the non-neural methods perform very competitively with fine-tuned transformer-based mono/multilingual language models, tiny large language models, and ChatGPT for classification tasks in transliterated text. These classical models train in seconds using only a fraction of the computing power, and thus the carbon footprint, required by language models. We demonstrate TF-IDF and BoW-based classifiers achieve performance within around 3% of fine-tuned LMs and could thus be considered as a strong baseline for transliterated text-based NLP tasks. Additionally, we investigated various mitigation strategies such as translation and augmentation via the use of ChatGPT, as well as Masked Language Modelling to dataset-specific pretraining for language models. Depending on the dataset and language, employing those mitigation techniques yields a 2-3% further improvement in accuracy and macro-F1 above baseline. Fahim Ahmed, Md Fahim, M. Ashraful Amin, Amin Ahsan Ali, A. K. M. Mahbubur Rahman |
ECAI | 2 |
| 2024 | TinyLLM Efficacy in Low-Resource Language: An Experiment on Bangla Text Classification Task
Farhan Noor Dehan, Md Fahim, A. K. M. Mahbubur Rahman, M. Ashraful Amin, Amin Ahsan Ali |
ICPR (19) | 2 |
| 2024 | How Good are LM and LLMs in Bangla Newspaper Article Summarization?
Faria Sultana, Md. Tahmid Hasan Fuad, Md Fahim, Rahat Rizvi Rahman, Meheraj Hossain, M. Ashraful Amin, A. K. M. Mahbubur Rahman, Amin Ahsan Ali |
ICPR (20) | 3 |
| 2023 | EDAL: Entropy based Dynamic Attention Loss for HateSpeech Classification
Md Fahim, Amin Ahsan Ali, M. Ashraful Amin, A. K. M. Mahbubur Rahman |
PACLIC | 1 |
| 2021 | Identifying Social Media Content Supporting Proud BoysabstractWhile most conversations on social media are inspiring and uplifting, a fraction of the users do engage in sharing content that supports extremist, radical philosophies and organizations. Such rhetoric, if left unchecked, can propagate virally on these platforms, ultimately escalating to turbulence and violence in the physical, offline spaces. Identifying such content from the volumes of social media feeds is therefore necessary to prevent the damage that it may cause, yet it is infeasible to undertake manually, calling for an automated approach. This paper demonstrates the potential of machine learning to separate social media feeds that are sympathetic to radical extremism using the example of Proud Boys, a contemporary right-wing group. From the tweets collected after Proud Boys protests in August 2020; linguistic, social and auxiliary features are extracted. Significance tests are used to select a subset of these features that contribute towards separating between extremist and normal content. Several machine learning models are trained based on a combination of these features. These models can identify tweets that support Proud Boys with excellent performance metrics. Artificial Neural Networks offer the best accuracy, precision, recall and F1-score. Feature importance, assessed using the Random Forest model indicates that users rely both on the expressive power of the language and other metadata features such as punctuations, mentions and URLs to voice and spread hateful ideology on online platforms. Md Fahim, Swapna S. Gokhale |
IEEE BigData | 1 |
| 2021 | Identifying Social Media Content Supporting Proud BoysabstractWhile most conversations on social media are inspiring and uplifting, a fraction of the users do engage in sharing content that supports extremist, radical philosophies and organizations. Such rhetoric, if left unchecked, can propagate virally on these platforms, ultimately escalating to turbulence and violence in the physical, offline spaces. Identifying such content from the volumes of social media feeds is therefore necessary to prevent the damage that it may cause, yet it is infeasible to undertake manually, calling for an automated approach. This paper demonstrates the potential of machine learning to separate social media feeds that are sympathetic to radical extremism using the example of Proud Boys, a contemporary right-wing group. From the tweets collected after Proud Boys protests in August 2020; linguistic, social and auxiliary features are extracted. Significance tests are used to select a subset of these features that contribute towards separating between extremist and normal content. Several machine learning models are trained based on a combination of these features. These models can identify tweets that support Proud Boys with excellent performance metrics. Artificial Neural Networks offer the best accuracy, precision, recall and F1-score. Feature importance, assessed using the Random Forest model indicates that users rely both on the expressive power of the language and other metadata features such as punctuations, mentions and URLs to voice and spread hateful ideology on online platforms. Md Fahim, Swapna S. Gokhale |
IEEE BigData | 1 |
| 2021 | Detecting Offensive Content on Twitter During Proud Boys RiotsabstractHateful and offensive speech on online social media platforms has seen a rise in the recent years. Often used to convey humor through sarcasm or to emphasize a point, offensive speech may also be employed to insult, deride and mock alternate points of view. In turbulent and chaotic circumstances, insults and mockery can lead to violence and unrest, and hence, such speech must be identified and tagged to limit its damage. This paper presents an application of machine learning to detect hateful and offensive content from Twitter feeds shared after the protests by Proud Boys, an extremist, ideological and violent hate group. A comprehensive coding guide, consolidating definitions of what constitutes offensive content based on the potential to trigger and incite people is developed and used to label the tweets. Linguistic, auxiliary and social features extracted from these labeled tweets were used to train machine learning classifiers, which detect offensive content with an accuracy of about 92%. An analysis of the importance scores reveals that offensiveness is pre-dominantly a function of words and their combinations, rather than meta features such as punctuations and quotes. This observation can form the foundation of pre-trained classifiers that can be deployed to automatically detect offensive speech in new and unforeseen circumstances. Md Fahim, Swapna S. Gokhale |
ICMLA | 1 |