Fuad Rahman 0001

dblp:11/478 · also Ahmad Fuad Rezaur Rahman · DBLP profile ↗
← Back
21ranked-venue papers in the field
5as first author
11since 2021 · last 2026
0000-0002-8670-7124ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 14 (2 first)Other / Interdisciplinary · 7 (3 first)
YearPublicationVenuePosition
2026 Figures as Evidence: Multi-image Scientific Generation
Jawad Ibn Ahad, Mritunjoy Chakraborty, Fuad Rahman 0001, Sifat Momen, Shafin Rahman, Nabeel Mohammed
ICDAR (3)3
2025 Teacher-Guided One-Shot Pruning via Context-Aware Knowledge Distillation
Md. Samiul Alim, Sharjil Khan, Amrijit Biswas, Fuad Rahman 0001, Shafin Rahman, Nabeel Mohammed
IEEE Big Data4
2025 Beyond Characters: Position-Aware Metric and a Lightweight LLM for Low-Resource Bangla Punctuation Restoration
Md Mehedi Hasan, S. M. Jishanul Islam, AKM Shahariar Azad Rabby, Fuad Rahman 0001
IEEE Big Data4
2025 Dynamic Temperature Scheduler for Knowledge Distillation
Sibgat Ul Islam, Jawad Ibn Ahad, Fuad Rahman 0001, Mohammad Ruhul Amin, Nabeel Mohammed, Shafin Rahman
IEEE Big Data3
2024 Empowering Meta-Analysis: Leveraging Large Language Models for Scientific Synthesis
abstract
This study investigates the automation of metaanalysis in scientific documents using large language models (LLMs). Meta-analysis is a robust statistical method that synthesizes the findings of multiple studies (support articles) to provide a comprehensive understanding. We know that a metaarticle provides a structured analysis of several articles. However, conducting meta-analysis by hand is labor-intensive, time-consuming, and susceptible to human error, highlighting the need for automated pipelines to streamline the process. Our research introduces a novel approach that fine-tunes the LLM on extensive scientific datasets to address challenges in big data handling and structured data extraction. We automate and optimize the meta-analysis process by integrating Retrieval Augmented Generation (RAG). Tailored through prompt engineering and a new loss metric, Inverse Cosine Distance (ICD), designed for fine-tuning on large contextual datasets, LLMs efficiently generate structured meta-analysis content. Human evaluation then assesses relevance and provides information on model performance in key metrics. This research demonstrates that fine-tuned models outperform non-fine-tuned models, with fine-tuned LLMs generating 87.6% relevant meta-analysis abstracts. The relevance of the context, based on human evaluation, shows a reduction in irrelevancy from 4.56% to 1.9%. These experiments were conducted in a low-resource environment, highlighting the study’s contribution to enhancing the efficiency and reliability of meta-analysis automation.
Jawad Ibn Ahad, Rafeed Mohammad Sultan, Abraham Kaikobad, Fuad Rahman 0001, Mohammad Ruhul Amin, Nabeel Mohammed, Shafin Rahman
IEEE Big Data4
2024 One Model Multiple Insights: CNNs for Medical Image Classification
abstract
Breast cancer diagnosis from ultrasound imaging poses a significant challenge due to the difficulty of simultaneously distinguishing between normal, benign, and malignant tissues within a unified framework. Existing approaches often prioritize binary classification, overlooking the clinical necessity for comprehensive multi-class differentiation. In this work, we propose an end-to-end convolutional neural network (CNN) architecture that directly addresses this gap. Leveraging a modified ResNet-18 backbone with domain-specific adaptations, our model integrates feature extraction and multi-class decision-making into a single efficient pipeline. Trained on the BUSI dataset, the proposed framework achieves 92.31% accuracy in multi-class classification, outperforming state-of-the-art methods by a significant margin. Through extensive ablation studies, we demonstrate the robustness and scalability of our approach, highlighting its clinical relevance for early detection and effective management of breast cancer. This work sets a new benchmark for ultrasound-based breast cancer diagnostics, offering a reliable and interpretable framework for real-world deployment.
AKM Shahariar Azad Rabby, Pratim Saha, Sheikh Abujar, Fuad Rahman 0001
IEEE Big Data4
2024 BanglaDialecto: An End-to-End AI-Powered Regional Speech Standardization
abstract
This study focuses on recognizing Bangladeshi dialects and converting diverse Bengali accents into standardized formal Bengali speech. Dialects, often referred to as regional languages, are distinctive variations of a language spoken in a particular location and are identified by their phonetics, pronunciations, and lexicon. Subtle changes in pronunciation and intonation are also influenced by geographic location, educational attainment, and socioeconomic status. Dialect standardization is needed to ensure effective communication, educational consistency, access to technology, economic opportunities, and the preservation of linguistic resources while respecting cultural diversity. Being the fifth most spoken language with around 55 distinct dialects spoken by 160 million people, addressing Bangla dialects is crucial for developing inclusive communication tools. However, limited research exists due to a lack of comprehensive datasets and the challenges of handling diverse dialects. With the advancement in multilingual Large Language Models (mLLMs), emerging possibilities have been created to address the challenges of dialectal Automated Speech Recognition (ASR) and Machine Translation (MT). This study presents an end-to-end pipeline for converting dialectal Noakhali speech to standard Bangla speech. This investigation includes constructing a large-scale diverse dataset with dialectal speech signals that tailored the fine-tuning process in ASR and LLM for transcribing the dialect speech to dialect text and translating the dialect text to standard Bangla text. Our experiments demonstrated that fine-tuning the Whisper ASR model achieved a CER of 0.8% and WER of 1.5%, while the BanglaT5 model attained a BLEU score of 41.6% for dialect-to-standard text translation. We completed our end-to-end pipeline for dialect standardization by utilizing AlignTTS, a text-to-speech (TTS) model. With potential applications across different dialects, this research lays the groundwork for future research into Bangla dialect standardization.
Md. Nazmus Sadat Samin, Jawad Ibn Ahad, Tanjila Ahmed Medha, Fuad Rahman 0001, Mohammad Ruhul Amin, Nabeel Mohammed, Shafin Rahman
IEEE Big Data4
2023 Versatile Bengali OCR: Document Analysis Technique for Varied Document Styles and Content
abstract
In our research paper, we introduce a distinctive Bengali OCR system that boasts impressive capabilities. This system excels in reconstructing document layouts while maintaining the integrity of structure, alignment, and even images. It integrates advanced image and signature detection for precise extraction. Specifically, tailored models for word segmentation accommodate various document types, such as computer-compose, letterpress, typewritten, and handwritten documents. Notably, the system handles static and dynamic handwritten inputs, recognizing diverse writing styles. Additionally, it achieves remarkable recognition of compound characters in the Bengali language. The comprehensive data collection contributes to a diverse corpus, and sophisticated technical components enhance character and word recognition. Other notable features include image, logo, signature recognition, table recognition, perspective correction, layout reconstruction, and a queuing module for efficient and scalable processing. The system showcases exceptional performance in the efficient and accurate extraction and analysis of text.
AKM Shahariar Azad Rabby, Hasmot Ali, Md. Majedul Islam, Fuad Rahman 0001
IEEE Big Data4
2021 Towards building a Bangla text recognition solution with a Multi-Headed CNN architecture
abstract
Bangla is among the ten most popular languages in the world by the number of speakers. The task of Bangla recognition is quite challenging than other languages because of the existence of graphemes of multiple single characters, and diacritics of vowels and consonants. The purpose of this study is to develop an innovative large-scale Bangla OCR solution based on character-level recognition. Two types of documents were used to test our method: handwritten and printed. In addition, our method was applied to the handwritten documents as well as three subdomains of the printed domain: computer-composed, letterpress, and typewritten documents using our proposed attentionbased multi-headed CNN architecture. Extensive testing shows that our method provides state-of-the-art performance on both handwritten and printed texts.
Md. Majedul Islam, Avishek Das, Ibna Kowsar, AKM Shahariar Azad Rabby, Nazmul Hasan, Fuad Rahman 0001
IEEE BigData6
2021 Modeling Influenza with a Forest Deep Neural Network Utilizing a Virtualized Clinical Semantic Network
abstract
CoViD-19 pandemic has shown that we have deep gaps in understanding this extremely infectious virus—not only both from a clinical diagnosis and treatment perspective—but also from a forecasting point of view, so that we are better prepared for the next onset of a similar pandemic, which, at this point, seems almost inevitable. In this paper, we present a novel approach towards modeling influenza, a closely related disease to CoViD-19, marrying clinical understanding with artificial intelligence, exploiting the Forest Deep Neural Network (fDNN) with accuracy rates in the 90% range.
Fuad Rahman 0001, Abrar Rahman, AKM Shahariar Azad Rabby, Md Jamiur Rahman Rifat, Mridul Banik, Md. Majedul Islam, Nor Azriah Aziz, Rick Meyer, John Kriak, Sidney Goldblatt
IEEE BigData1
2021 A Large Multi-target Dataset of Common Bengali Handwritten Graphemes
Samiul Alam, Tahsin Reasat, Asif Shahriyar Sushmit, Sadi Mohammad Siddiquee, Fuad Rahman 0001, Mahady Hasan, Ahmed Imtiaz Humayun
ICDAR (4)5
2020 A Novel Deep Learning Character-Level Solution to Detect Language and Printing Style from a Bilingual Scanned Document
abstract
Bangla is one of the world's most widely-spoken languages, but few languages (or "script") automation solutions have been reported for it. To build an OCR system, it is very important to detect the language and type of printing style to run specific character recognition and segmentation modules. This paper presents a novel solution to automatically detect the language (Bangla vs English in terms of the script), and printing style (printed vs handwritten) from any given bilingual scanned document using multiple deep learning models.
AKM Shahariar Azad Rabby, Md. Majedul Islam, Nazmul Hasan, Jebun Nahar, Fuad Rahman 0001
IEEE BigData5
2020 A Machine Learning Based Modeling of the Cytokine Storm as it Relates to COVID-19 Using a Virtual Clinical Semantic Network (vCSN)
abstract
This paper presents a targeted, machine learning based solution to model the phenomenon known as the `cytokine storm,' which is suspected to play a major role in explaining the highly variable severity of COVID-19 among patients. It describes how a Natural Language Processing (NLP) approach, augmented by biomedical knowledge databases, can extract pre-existing conditions and relevant clinical markers from Electronic Health Records (EHRs). These extracted variables can be modeled to demonstrate correlation with the severity of infection outcomes, the building blocks of a comprehensive risk assessment and stratification strategy to predict which patients have higher or lower risks in terms of the disease severity and likelihood of hospitalization, exclusively from insights taken from the natural language data. The model has been applied to a cohort of patients from a large database of real, anonymized patients and has displayed demonstrable results.
Abrar Rahman, John Kriak, Rick Meyer, Sidney Goldblatt, Fuad Rahman 0001
IEEE BigData5
2019 Sankhya: An Unbiased Benchmark for Bangla Handwritten Digits Recognition
abstract
The rise of artificial intelligence technology along with machine and deep learning are opening up almost limitless possibilities. In recent years, application-based researchers in machine learning and deep learning have started developing solutions for many practical problems. Handwriting recognition is one such area of interest. Bangla, being the seventh most spoken language in the world, is not an exception. However, unlike English, there have not been concerted formal attempts in building a benchmark in comparing the different approaches reported in the literature, mainly because of the lack of openly and freely available datasets and diversity of the approaches without formal comparative studies. In this research paper, we seek to rectify this gap. We have focused on benchmarking five robust algorithms: K-Nearest Neighbor (KNN), Support Vector Machine (SVM), Random Forest (RF), Multi-layer Perceptron (MLP), Convolutional Neural Network (CNN), on all publicly available Bangla handwriting digits datasets, including Ekush, NumtaDB, CMARTdb, and BDRW. NumtaDB itself is a collection of five handwriting datasets. We have worked on fine-tuning these algorithms by finding the best possible hyper-parameters of these algorithms. It is our hope that Sankhya will work as a beginning point of an open and verifiable benchmarking process that we plan to repeat every two years for now on and set a standard for testing and validating newer and novel algorithms that will be reported in this area in the future. In addition, we have extensively compared our research with other states of the art research and our versions of these algorithms are now outperforming every reported result on these datasets.All the datasets we used are open-sourced. In addition, we are making the.csv version of these datasets available in public GitHub. Of all the models we tested, the Sankhya CNN model performed the best for all these datasets, which we fine-tuned specifically for Bangla character recognition. We are making this CNN model available in public GitHub.
Fuad Rahman 0001, AKM Shahariar Azad Rabby
IEEE BigData2
2019 Smart EHR - A Big-Data Approach to Automated Collection and Processing of Multi-Modal Health Signals in a Doctor-patient Encounter
abstract
This work focuses on creating a smart Electronic Health Record (EHR) platform to collect and analyze doctor-patient interactions. It is an often cited fact that doctors are spending less and less time with patients, but two recent surveys have quantified what these numbers really are and the results are disturbing. In one survey [1], it was found that during the office day, physicians spent 27.0% of their total time on direct clinical face time with patients and 49.2% of their time on EHR and deskwork. In another survey [2], results “suggest that the physicians logged an average of 3.08 hours on office visits and 3.17 hours on desktop medicine each day,... Over time, log records from physicians showed a decline in the time allocated to face-to-face visits, accompanied by an increase in time allocated to desktop medicine It is precisely this problem that the University of Arizona's “Wired Room” project seeks to address.
Abrar Rahman, Ari Mitra, Fuad Rahman 0001, Marvin J. Slepian
IEEE BigData3
2016 A novel big-data processing framwork for healthcare applications: Big-data-healthcare-in-a-box
abstract
Herein we present a novel big-data framework for healthcare applications. Healthcare data is well suited for bigdata processing and analytics because of the variety, veracity and volume of these types of data. In recent times, many areas within healthcare have been identified that can directly benefit from such treatment. However, setting up these types of architecture is not trivial. We present a novel approach of building a big-data framework that can be adapted to various healthcare applications with relative use, making this a one-stop “Big-Data-Healthcare-in-a-Box”.
Fuad Rahman 0001, Marvin J. Slepian, Ari Mitra
IEEE BigData1
2003 Web Page Summarization for Handheld Devices: A Natural Language Approach
abstract
Summarization of web pages is a very interesting topic from both academic and commercial point of view. Academically, it is challenging to create a summary of a document (e.g. a web page) that is highly structured and has multi-media components in it. From the commercial point of view, it is advantageous to summarize web pages to be viewed in small display devices such as PDAs and cell phones. Summarization not only makes web browsing and navigation easier, but it makes browsing faster as complete web pages need not be downloaded before viewing. In this paper, a novel combination of natural language and non-natural language based summarization techniques have been used to automatically generate an intelligent re-authored display of web pages in real time. 1.
Hassan Alam, Rachmat Hartono, Fuad Rahman 0001, Yuliya Tarnikova, Che Wilcox
ICDAR4
2003 Structured and Unstructured Document Summarization: Design of a Commercial Summarizer using Lexical Chains
abstract
The process of summarizing documents is becoming increasingly important in the light of recent advances in document creation/distribution technology, and the resulting influx of large numbers of documents in every day life. This paper presents a document summarizer that combines document analysis, structural decomposition, XML representation and lexical chain analysis. The proposed summarizer is compared to three commercially available summarizers and it is shown that it produces either comparable or better summaries overall. 1.
Hassan Alam, Mikako Nakamura, Fuad Rahman 0001, Yuliya Tarnikova, Che Wilcox
ICDAR4
2002 Multiple Classifier Combination for Character Recognition: Revisiting the Majority Voting System and Its Variations
Fuad Rahman 0001, Hassan Alam, Michael C. Fairhurst
Document Analysis Systems1
2001 Automatic Summarization of Web Content to Smaller Display Devices
abstract
Web documents usually have complicated layouts and the overall information content can be huge. All these documents are designed for viewing in large screen devices, such as a computer monitor. In recent times, a large number of small screen portable devices, such as personal digital assistants (PDA) and cellular phones, have been made available for mobile browsing. Viewing a Web page originally written for large screen devices using these very small screen devices can be extremely cumbersome. This paper discusses this issue of small viewing form factor of electronics devices from the perspective of Web browsing and proposes an approach to automatically summarize and transform Web documents into a meaningful, readable and above all, browsable format.
Fuad Rahman 0001, Hassan Alam, Rachmat Hartono, K. Ariyoshi
ICDAR1
1997 Introducing New Multiple Expert Decision Combination Topologies: A Case Study using Recognition of Handwritten Characters
abstract
A new topology for classifying decision combinations of multiple experts in the framework of a multiple expert character recognition platform is introduced. It is demonstrated that many existing multiple expert configurations for character recognition can be categorised by using this method of defining classification strategies. It is also demonstrated that the design of multiple expert character recognition configurations can be streamlined by classifying these structures in terms of how the channels used for carrying information among different experts are interconnected irrespective of the algorithms used by cooperating experts and by the final decision combination expert. Case studies of actual multiple expert character recognition configurations have been investigated and it is shown how they can be categorised with respect to the decision combination topologies introduced in the paper.
Fuad Rahman 0001, Michael C. Fairhurst
ICDAR1