VLDB 2026 Research / reviewers in the wild / expert
Shafin Rahman
dblp:95/10398
· DBLP profile ↗
7ranked-venue papers in the field
0as first author
7since 2021 · last 2026
0000-0001-7169-0318ORCID · conflict
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 5Data Mining & Knowledge Discovery · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Figures as Evidence: Multi-image Scientific Generation
Jawad Ibn Ahad, Mritunjoy Chakraborty, Fuad Rahman 0001, Sifat Momen, Shafin Rahman, Nabeel Mohammed |
ICDAR (3) | 5 |
| 2025 | LAET: A Layer-Wise Adaptive Ensemble Tuning Framework for Pretrained Language Models
Jawad Ibn Ahad, Muhammad Rafsan Kabir, Robin Krambroeckers, Sifat Momen, Nabeel Mohammed, Shafin Rahman |
IEEE Big Data | 6 |
| 2025 | Teacher-Guided One-Shot Pruning via Context-Aware Knowledge Distillation
Md. Samiul Alim, Sharjil Khan, Amrijit Biswas, Fuad Rahman 0001, Shafin Rahman, Nabeel Mohammed |
IEEE Big Data | 5 |
| 2025 | Dynamic Temperature Scheduler for Knowledge Distillation
Sibgat Ul Islam, Jawad Ibn Ahad, Fuad Rahman 0001, Mohammad Ruhul Amin, Nabeel Mohammed, Shafin Rahman |
IEEE Big Data | 6 |
| 2025 | HadaSmileNet: Hadamard Fusion of Handcrafted and Deep-Learning Features for Enhancing Facial Emotion Recognition of Genuine SmilesabstractThe distinction between genuine and posed emotions represents a fundamental pattern recognition challenge with significant implications for data mining applications in social sciences, healthcare, and human-computer interaction. While recent multitask learning frameworks have shown promise in combining deep learning architectures with handcrafted D-Marker features for smile facial emotion recognition, these approaches exhibit computational inefficiencies due to auxiliary task supervision and complex loss balancing requirements. This paper introduces HadaSmileNet, a novel feature fusion framework that directly integrates transformer-based representations with physiologically-grounded D-Markers through parameter-free multiplicative interactions. Through systematic evaluation of 15 fusion strategies, we demonstrate that Hadamard multi-plicative fusion achieves optimal performance by enabling direct feature interactions while maintaining computational efficiency. The proposed approach establishes new state-of-the-art results for deep learning methods across four benchmark datasets: UvA-NEMO (88.7%, +0.8%), MMI (99.7%), SPOS (98.5%, +0.7%), and BBC (100%, +5.0%). Comprehensive computational analysis reveals 26% parameter reduction and simplified training compared to multitask alternatives, while feature visualization demonstrates enhanced discriminative power through direct domain knowledge integration. The framework's efficiency and effectiveness make it particularly suitable for practical deployment in multimedia data mining applications that require realtime affective computing capabilities. Mohammad Junayed Hasan, Nabeel Mohammed, Shafin Rahman, Philipp Koehn |
ICDM | 3 |
| 2024 | Empowering Meta-Analysis: Leveraging Large Language Models for Scientific SynthesisabstractThis study investigates the automation of metaanalysis in scientific documents using large language models (LLMs). Meta-analysis is a robust statistical method that synthesizes the findings of multiple studies (support articles) to provide a comprehensive understanding. We know that a metaarticle provides a structured analysis of several articles. However, conducting meta-analysis by hand is labor-intensive, time-consuming, and susceptible to human error, highlighting the need for automated pipelines to streamline the process. Our research introduces a novel approach that fine-tunes the LLM on extensive scientific datasets to address challenges in big data handling and structured data extraction. We automate and optimize the meta-analysis process by integrating Retrieval Augmented Generation (RAG). Tailored through prompt engineering and a new loss metric, Inverse Cosine Distance (ICD), designed for fine-tuning on large contextual datasets, LLMs efficiently generate structured meta-analysis content. Human evaluation then assesses relevance and provides information on model performance in key metrics. This research demonstrates that fine-tuned models outperform non-fine-tuned models, with fine-tuned LLMs generating 87.6% relevant meta-analysis abstracts. The relevance of the context, based on human evaluation, shows a reduction in irrelevancy from 4.56% to 1.9%. These experiments were conducted in a low-resource environment, highlighting the study’s contribution to enhancing the efficiency and reliability of meta-analysis automation. Jawad Ibn Ahad, Rafeed Mohammad Sultan, Abraham Kaikobad, Fuad Rahman 0001, Mohammad Ruhul Amin, Nabeel Mohammed, Shafin Rahman |
IEEE Big Data | 7 |
| 2024 | BanglaDialecto: An End-to-End AI-Powered Regional Speech StandardizationabstractThis study focuses on recognizing Bangladeshi dialects and converting diverse Bengali accents into standardized formal Bengali speech. Dialects, often referred to as regional languages, are distinctive variations of a language spoken in a particular location and are identified by their phonetics, pronunciations, and lexicon. Subtle changes in pronunciation and intonation are also influenced by geographic location, educational attainment, and socioeconomic status. Dialect standardization is needed to ensure effective communication, educational consistency, access to technology, economic opportunities, and the preservation of linguistic resources while respecting cultural diversity. Being the fifth most spoken language with around 55 distinct dialects spoken by 160 million people, addressing Bangla dialects is crucial for developing inclusive communication tools. However, limited research exists due to a lack of comprehensive datasets and the challenges of handling diverse dialects. With the advancement in multilingual Large Language Models (mLLMs), emerging possibilities have been created to address the challenges of dialectal Automated Speech Recognition (ASR) and Machine Translation (MT). This study presents an end-to-end pipeline for converting dialectal Noakhali speech to standard Bangla speech. This investigation includes constructing a large-scale diverse dataset with dialectal speech signals that tailored the fine-tuning process in ASR and LLM for transcribing the dialect speech to dialect text and translating the dialect text to standard Bangla text. Our experiments demonstrated that fine-tuning the Whisper ASR model achieved a CER of 0.8% and WER of 1.5%, while the BanglaT5 model attained a BLEU score of 41.6% for dialect-to-standard text translation. We completed our end-to-end pipeline for dialect standardization by utilizing AlignTTS, a text-to-speech (TTS) model. With potential applications across different dialects, this research lays the groundwork for future research into Bangla dialect standardization. Md. Nazmus Sadat Samin, Jawad Ibn Ahad, Tanjila Ahmed Medha, Fuad Rahman 0001, Mohammad Ruhul Amin, Nabeel Mohammed, Shafin Rahman |
IEEE Big Data | 7 |