Nazmul Kazi

dblp:236/1843 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2024
0000-0003-3610-455XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2024 Poster: Towards Understanding Root Causes of Real Failures in Healthcare Machine Learning Applications
abstract
Machine learning (ML) is widely used in healthcare applications to diagnose diseases, forecast disease progression, develop personalized treatment plans, and aid in drug discovery and development [1]. The development of ML applications is inherently different from other applications. Instead of explicitly coding the program's logic, ML applications learn this logic using a machine learning algorithm and provided data. Thus, faults in ML applications, as opposed to in others, can manifest in all these components, such as the application itself, incorrect use of the machine learning algorithms or libraries, and issues with data used for training. Thus, understanding these various root causes of real faults would help to develop effective testing techniques for these applications. Therefore, we analyzed 50 real-life faults from four ML healthcare applications to better understand the faults presented in this domain.
Guna Sekaran Jaganathan, Nazmul Kazi, Indika Kahanda, Upulee Kanewala
ICST2
2024 MLHCBugs: A Framework to Reproduce Real Faults in Healthcare Machine Learning Applications
abstract
Machine Learning (ML) is the field of study that allows computers to learn from experiences without being explicitly programmed [1]. ML models are currently used in many safety-critical applications in healthcare [2]–[4] and survival analyses [5]. Thus, faults in this software can directly impact the quality of human life. In an ML application, the program logic is typically derived by a ML algorithm using the currently available data (i.e., training data) rather than explicitly being programmed [6]. Therefore, the program's behavior would evolve as it is exposed to new data. Further, healthcare ML applications are inherently complex and typically constructed by the interconnection of several components, such as data that is used to derive the logic, the ML framework that contains the algorithms used by the program, and the program itself that is written by the programmer for a specific task involved with healthcare [7]. Faults in any of these components may produce an observable incorrect output or the statistical nature of these programs may mask the incorrect output altogether, making it more challenging to understand the root causes of these failures.
Guna Sekaran Jaganathan, Nazmul Kazi, Indika Kahanda, Upulee Kanewala
ICST2
2023 Enhancing Transfer Learning of LLMs through Fine- Tuning on Task - Related Corpora for Automated Short-Answer Grading
abstract
Automated short-answer grading (ASAG) is a cru-cial element of any intelligent tutoring platform. Machine Learning (ML) has shown great promise for ASAG. However, this task remains challenging even for Deep Learning (DL) approaches and Large Language Models (LLMs), requiring semantic inference and textual entailment recognition. The SemEval-2013 Task 7, The Joint Student Response Analysis and 8th Recognizing Textual Entailment Challenge, is a benchmark widely used for research on ASAG. The SciEntsBank data included in this collection contains nearly 11,000 answers to 197 assessment questions in 15 different science domains. Despite the popularity, only a few researchers have explored the potential of DL or LLMs for this task. In this project, we explore the effectiveness of the RoBERTa Large model, an LLM trained on an extensive text corpus for language comprehension. By fine-tuning the model on the Multi-Genre Natural Language Inference (MNLI) corpus for semantic inference and subsequently on the SciEntsBank dataset, with a focus on the 3-way labels of correct, incorrect, and contradictory, we achieved a weighted Fl-score of 0.77, 0.72, and 0.72 on unseen answers, questions, and domains, respectively. Notably, our model significantly benefits from fine-tuning on the MNLI corpus, particularly in enhancing its performance on the contradictory class (which constitutes only 10% of the dataset) through transfer learning leading to significant improvements on the more challenging test sets: unseen questions and unseen domains.
Nazmul Kazi, Indika Kahanda
ICMLA1
2023 Zero-Shot Information Extraction with Community-Fine-Tuned Large Language Models From Open-Ended Interview Transcripts
abstract
Machine learning holds significant promise for automating and optimizing text data analysis. However, resource-intensive tasks like data annotation, model training, and parameter tuning often limit its practicality for one-time data extraction, medium-sized datasets, or short-term projects. There are many community-fine-tuned large language models (CLLMs) that are fine-tuned on task-specific datasets and can demonstrate impressive performance on unseen data without further fine-tuning. Adopting a hybrid approach of leveraging CLLMs for rapid text data extraction and subsequently hand-curating the inaccurate outputs can yield high-quality results, workload balance, and improved efficiency. This project applies CLLMs to three tasks involving the analysis of open-ended survey responses: semantic text matching, exact answer extraction, and sentiment analysis. We present our overall process and discuss several seemingly simple yet effective techniques that we employ to improve model performance without fine-tuning the CLLMs on our own data. Our results demonstrate high precision in semantic text matching (0.92) and exact answer extraction (0.90), while the sentiment analysis model shows room for improvement (precision: 0.65, recall: 0.94, F1: 0.77). This study showcases the potential of CLLMs in open-ended survey text data analysis, particularly in scenarios with limited resources and scarce labeled data.
Nazmul Kazi, Indika Kahanda, S. Indu Rupassara, John W. Kindt
ICMLA1