VLDB 2026 Research / reviewers in the wild / expert
Sadia Jahan
dblp:302/7782
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-5384-3162ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | How Does ChatGPT Make Assumptions When Creating Erroneous Programs?abstractLarge Language Models (LLMs) like ChatGPT are increasingly integrated into software development environments due to their strong performance in code generation. However, they often struggle with complex logic, security vulnerabilities, and code quality issues. These problems frequently originate from misunderstandings of problem requirements and logical inconsistencies, which can lead to faulty or vulnerable software. In this study, we conduct an initial empirical analysis to investigate the causes of erroneous code generated by the state-of-the-art LLM model GPT-4o. Using the HumanEval dataset, we prompt GPT-4o to generate Python solutions and list its 3 most important assumptions. We validate these outputs against the provided test cases in dataset and identify 17 defective programs out of 164 total solutions. By analyzing the 17 failures and 51 assumptions made on these tasks, we find that about 53% the failures are directly related to wrong or erroneously implemented assumptions raised by the GPT model itself, and totally 71% of code generation failures are related to erroneously made or implemented assumptions. Sadia Jahan, Xiaoyin Wang |
ASE | 1 |
| 2024 | A Large Language Model Approach to Code and Privacy Policy AlignmentabstractAs mobile technology has advanced, individuals have started relying on their smartphones to conduct more of their everyday tasks. From playing games or streaming media to social networking and banking, apps on a user's device may have access to the most sensitive information on the device. Privacy policies are designed to inform users of such data practices so that they can make reasonable decisions when using the app. However, an app's true behavior may not always align with the statements in a privacy policy. In this work, we divide our study into two components and compare the viability of various large language models (LLMs): methods for extracting and summarizing privacy policy data practices, or information-type extraction and action-verb extraction; and methods for measuring whether the policy acknowledges the interaction with certain information (sensitive data) compared to identified methods within its app's source code. Fine-tuning GPT-3.5 Turbo delivers a higher average F1-score for both action verb extraction (0.50) and information-type extraction (0.84) compared to other LLMs. ChatGPT outperforms other language models in traditional semantic similarity, providing a consistently high performance, including the highest F1-score (0.52) for this task. Our approaches demonstrate that these LLMs are viable in performing such tasks and additionally that pre-trained instruction-based LLMs are capable of identifying the complex relationships between policies and source code. Gabriel A. Morales, Pragyan K. C, Sadia Jahan, Mitra Bokaei Hosseini, Rocky Slavin |
SANER | 3 |
| 2021 | Active Learning with an Adaptive Classifier for Inaccessible Big Data AnalysisabstractSupervised machine learning (ML) approaches effectively derive valuable insights from big data. These approaches, on the other hand, require an extensive amount of high quality annotated data for training, created manually by domain experts through a costly and time-consuming process. To overcome this challenge, active learning (AL) is a promising approach, which can support a fast, cost-efficient and common strategy to deal with big data with limited labeling effort. Instead of annotating a large pool of unlabeled data, as in standard supervised learning, AL reduces the volume of data that requires manual annotation by effectively selecting subsets of highly informative samples for manual annotation within an iterative process. In this paper, we aim to present a robust approach utilizing AL to mitigate the aforementioned challenges and help the decision-makers. To be precise, we propose a framework involving a support vector machine (SVM) technique in AL for mining big data to manage inaccessible data situations. The proposed approach is tested on five different semi-supervised data sets. The performance of the proposed framework is evaluated using traditional ML classifiers such as Naïve Bayes (NB), Decision Tree (DT), Sequential Minimal Optimization (SMO), Random Forest (RF), Bagging and Adaboost. Among the reported classifiers, bagging achieves the best outcome, delivering 99.19% accuracy. According to the results of the experiment conducted we find that the proposed method increases the efficiency of the classifiers in AL with fewer training instances. Sadia Jahan, Md. Rafiqul Islam 0001, Khan Md Hasib, Usman Naseem, Md. Saiful Islam 0003 |
IJCNN | 1 |