Zahin Wahab

dblp:305/6296 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0008-7273-4165ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Secret Leak Detection in Software Issue Reports using LLMs: A Comprehensive Evaluation
abstract
In the digital era, accidental exposure of sensitive information such as API keys, tokens, and credentials is a growing security threat. While most prior work focuses on detecting secrets in source code, leakage in software issue reports remains largely unexplored. This study fills that gap through a large-scale analysis and a practical detection pipeline for exposed secrets in GitHub issues. Our pipeline combines regular expression–based extraction with large language model (LLM)–based contextual classification to detect real secrets and reduce false positives. We build a benchmark of 54,148 instances from public GitHub issues, including 5,881 manually verified true secrets. Using this dataset, we evaluate entropy-based baselines and keyword heuristics used by prior secret detection tools, classical machine learning, deep learning, and LLM-based methods. Regex and entropy based approaches achieve high recall but poor precision, while smaller models such as RoBERTa and CodeBERT greatly improve performance (F1 = 92.70%). Proprietary models like GPT-4o perform moderately in few-shot settings (F1 = 80.13%), and fine-tuned open-source larger LLMs such as Qwen and LLaMA reach up to 94.49% F1. Finally, we also validate our approach on 178 real-world GitHub repositories, achieving an F1-score of 81.6% which demonstrates our approach’s strong ability to generalize to in-the-wild scenarios.
Sadif Ahmed, Md Nafiu Rahman, Zahin Wahab, Gias Uddin 0001, Rifat Shahriyar
MSR3
2025 Secret Breach Detection in Source Code with Large Language Models
abstract
Background: Leaking sensitive information-such as API keys, tokens, and credentials-in source code remains a persistent security threat. Traditional regex and entropy-based tools often generate high false positives due to limited contextual understanding. Aims: This work aims to enhance secret detection in source code using large language models (LLMs), reducing false positives while maintaining high recall. We also evaluate the feasibility of using fine-tuned, smaller models for local deployment. Method: We propose a hybrid approach combining regex-based candidate extraction with LLM-based classification. We evaluate pre-trained and fine-tuned variants of various Large Language Models on a benchmark dataset from 818 GitHub repositories. Various prompting strategies and efficient fine-tuning methods are employed for both binary and multiclass classification. Results: The fine-tuned LLaMA-3.1 8B model achieved an F1-score of 0.9852 in binary classification, outperforming regex-only baselines. For multiclass classification, Mistral-7B reached 0.982 accuracy. Fine-tuning significantly improved performance across all models. Conclusions: Fine-tuned LLMs offer an effective and scalable solution for secret detection, greatly reducing false positives. Open-source models provide a practical alternative to commercial APIs, enabling secure and cost-efficient deployment in development workflows.
Md Nafiu Rahman, Sadif Ahmed, Zahin Wahab, Rifat Shahriyar
ESEM3
2024 Dealing with Smart GPS Spoofing Attacks in VANETs: 3BSM Approach
abstract
Nowadays VANETs are being used to ensure road safety and reduce traffic congestion. Vehicles exchange real time position with one another through basic safety messages and these periodic updates must be authenticated for safety purpose. Malicious vehicles can broadcast spoofed position and such an insider attack cannot be dealt with cryptography and digital signature. However, previous machine learning-based approaches that used distributed data-centric misbehaviour detection schemes could detect such attacks in most of the scenarios. However, in case of smarter attacks in a sparse network (where the attacker vehicle pretend to be stand still even though it is on the move), previous solutions could not always detect such attacks. We propose a novel detection scheme which outperforms previously published methods in this specific scenario. For evaluating the performance of this approach, we have used VeReMi dataset, a public repository for the malicious node detection in VANETs. We combine three consecutive basic safety messages and surpass currently existing methods in detecting eventual stop attacks, especially in sparse networks. Our proposed method can help in securing VANET, thereby preventing fatal accidents.
Zahin Wahab, Bishal Basak Papan, Md. Shohrab Hossain, Mohammed Atiquzzaman
ICC1
2021 wQFM: highly accurate genome-scale species tree estimation from weighted quartets
abstract
MOTIVATION: Species tree estimation from genes sampled from throughout the whole genome is complicated due to the gene tree-species tree discordance. Incomplete lineage sorting (ILS) is one of the most frequent causes for this discordance, where alleles can coexist in populations for periods that may span several speciation events. Quartet-based summary methods for estimating species trees from a collection of gene trees are becoming popular due to their high accuracy and statistical guarantee under ILS. Generating quartets with appropriate weights, where weights correspond to the relative importance of quartets, and subsequently amalgamating the weighted quartets to infer a single coherent species tree can allow for a statistically consistent way of estimating species trees. However, handling weighted quartets is challenging. RESULTS: We propose wQFM, a highly accurate method for species tree estimation from multi-locus data, by extending the quartet FM (QFM) algorithm to a weighted setting. wQFM was assessed on a collection of simulated and real biological datasets, including the avian phylogenomic dataset, which is one of the largest phylogenomic datasets to date. We compared wQFM with wQMC, which is the best alternate method for weighted quartet amalgamation, and with ASTRAL, which is one of the most accurate and widely used coalescent-based species tree estimation methods. Our results suggest that wQFM matches or improves upon the accuracy of wQMC and ASTRAL. AVAILABILITY AND IMPLEMENTATION: Datasets studied in this article and wQFM (in open-source form) are available at https://github.com/Mahim1997/wQFM-2020. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mahim Mahbub, Zahin Wahab, Rezwana Reaz, Mohammad Saifur Rahman 0001, Md. Shamsuzzoha Bayzid
Bioinform.2