Md Fahim

dblp:311/0390 · DBLP profile ↗
← Back
4ranked-venue papers in the field
2as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2 (2 first)Data Mining & Knowledge Discovery · 1Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2026 Training-free layer selection for partial fine-tuning of language models
abstract
The growing scale of pre-trained language models poses a challenge in fine-tuning for downstream tasks, especially in resource-constrained settings. Recent studies highlight that not all layers in Transformer-based language models contribute equally to downstream task performance, giving rise to various partial fine-tuning strategies. We propose a training-free approach for layer-wise partial fine-tuning that leverages the cosine similarity between representative tokens across layers to identify inter-layer relationships. Our method comprises two stages: (i) scoring layers based on their relevance to the task via a single forward pass, and (ii) fine-tuning a subset of layers, either highest-scoring, lowest-scoring, or block-wise, while keeping others frozen. We conduct experiments on 16 diverse NLP datasets, including single-sentence and sentence-pair classification tasks, as well as generation tasks. Our method achieves competitive performance compared to full fine-tuning, with an average training speedup of 1.5 and a reduction of trainable parameters by 75%, and outperforms all comparative baselines in 14 out of 16 evaluated datasets. Additionally, our approach does not cause any notable drop in performance when the domain is changed for the evaluation tasks, demonstrating a robust cross-domain performance. • Efficient training-free layer selection uses token cosine similarity. • Inter-layer token relationships identify optimal layers for fine-tuning. • Reduces trainable parameters by 75% with a 1.5x training speedup. • Our paper yields better accuracy on 16 diverse datasets compared to the SOTA works. • Layer selection strategy preserves robust cross-domain generalization.
Aldrin Kabya Biswas, Md Fahim, Md. Tahmid Hasan Fuad, Akm Moshiur Rahman Mazumder, Amin Ahsan Ali, A. K. M. Mahbubur Rahman
Inf. Sci.2
2025 ChitroJera: A Regionally Relevant Visual Question Answering Dataset for Bangla
Deeparghya Dutta Barua, Md Sakib Ul Rahman Sourove, Md Fahim, Fabiha Haider, Fariha Tanjim Shifat, Md Tasmim Rahman Adib, Anam Borhan Uddin, Md Farhan Ishmam, Md. Farhad Alam
ECML/PKDD (4)3
2021 Identifying Social Media Content Supporting Proud Boys
abstract
While most conversations on social media are inspiring and uplifting, a fraction of the users do engage in sharing content that supports extremist, radical philosophies and organizations. Such rhetoric, if left unchecked, can propagate virally on these platforms, ultimately escalating to turbulence and violence in the physical, offline spaces. Identifying such content from the volumes of social media feeds is therefore necessary to prevent the damage that it may cause, yet it is infeasible to undertake manually, calling for an automated approach. This paper demonstrates the potential of machine learning to separate social media feeds that are sympathetic to radical extremism using the example of Proud Boys, a contemporary right-wing group. From the tweets collected after Proud Boys protests in August 2020; linguistic, social and auxiliary features are extracted. Significance tests are used to select a subset of these features that contribute towards separating between extremist and normal content. Several machine learning models are trained based on a combination of these features. These models can identify tweets that support Proud Boys with excellent performance metrics. Artificial Neural Networks offer the best accuracy, precision, recall and F1-score. Feature importance, assessed using the Random Forest model indicates that users rely both on the expressive power of the language and other metadata features such as punctuations, mentions and URLs to voice and spread hateful ideology on online platforms.
Md Fahim, Swapna S. Gokhale
IEEE BigData1
2021 Identifying Social Media Content Supporting Proud Boys
abstract
While most conversations on social media are inspiring and uplifting, a fraction of the users do engage in sharing content that supports extremist, radical philosophies and organizations. Such rhetoric, if left unchecked, can propagate virally on these platforms, ultimately escalating to turbulence and violence in the physical, offline spaces. Identifying such content from the volumes of social media feeds is therefore necessary to prevent the damage that it may cause, yet it is infeasible to undertake manually, calling for an automated approach. This paper demonstrates the potential of machine learning to separate social media feeds that are sympathetic to radical extremism using the example of Proud Boys, a contemporary right-wing group. From the tweets collected after Proud Boys protests in August 2020; linguistic, social and auxiliary features are extracted. Significance tests are used to select a subset of these features that contribute towards separating between extremist and normal content. Several machine learning models are trained based on a combination of these features. These models can identify tweets that support Proud Boys with excellent performance metrics. Artificial Neural Networks offer the best accuracy, precision, recall and F1-score. Feature importance, assessed using the Random Forest model indicates that users rely both on the expressive power of the language and other metadata features such as punctuations, mentions and URLs to voice and spread hateful ideology on online platforms.
Md Fahim, Swapna S. Gokhale
IEEE BigData1