VLDB 2026 Research / reviewers in the wild / expert
Md Tanvirul Alam
dblp:318/3273
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0003-4284-2743ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | R+R: Revisiting Static Feature-Based Android Malware Detection Using Machine LearningabstractStatic feature-based Android malware detection using machine learning (ML) remains critical due to its scalability and efficiency. However, existing approaches often overlook security-critical reproducibility concerns, such as dataset duplication, inadequate hyperparameter tuning, and variance from random initialization. This can significantly compromise the practical effectiveness of these systems. In this paper, we systematically investigate these challenges by proposing a more rigorous methodology for model selection and evaluation. Using two widely used datasets, Drebin and APIGraph, we evaluate six ML models of varying complexity under both offline and continuous active learning settings. Our analysis demonstrates that, contrary to popular belief, well-tuned, simpler models, particularly tree-based methods like XGBoost, consistently outperform more complex neural networks, especially when duplicates are removed. To promote transparency and reproducibility, we open-source our codebase, which is extensible for integrating new models and datasets, facilitating reproducible security research. Md Tanvirul Alam, Dipkamal Bhusal, Nidhi Rastogi |
ACSAC | 1 |
| 2025 | ADAPT: A Pseudo-labeling Approach to Combat Concept Drift in Malware DetectionabstractMachine learning models are commonly used for malware classification; however, they suffer from performance degradation over time due to concept drift. Adapting these models to changing data distributions requires frequent updates, which rely on costly ground truth annotations. While active learning can reduce the annotation burden, leveraging unlabeled data through semi-supervised learning remains a relatively underexplored approach in the context of malware detection. In this research, we introduce ADAPT, a novel pseudo-labeling semisupervised algorithm for addressing concept drift. Our modelagnostic method can be applied to various machine learning models, including neural networks and tree-based algorithms. We conduct extensive experiments on five diverse malware detection datasets spanning Android, Windows, and PDF domains. The results demonstrate that our method consistently outperforms baseline models and competitive benchmarks. This work paves the way for more effective adaptation of machine learning models to concept drift in malware detection. Md Tanvirul Alam, Aritran Piplai, Nidhi Rastogi |
RAID | 1 |
| 2025 | Assessing Effective Token Length of Multimodal Models for Text-to-Image RetrievalabstractMultimodal embedding models have been widely adopted in text-toimage retrieval, enabling direct comparison between text and image modalities.However, how well they handle long text is poorly understood.For instance, Long-CLIP found that OpenAI's CLIP model, despite having a 77-token input limit, maintains optimal performance for only 20 tokens-its effective token length.In this paper, we build on the Long-CLIP study, and extend the analysis to other widely used multimodal models and find their effective token length.Unlike Long-CLIP, we examine how domain-specific language influences changes in effective token length and explore its implications on different domains.Based on our findings, we create a comprehensive reference of various models' effective token length across different domains; offering deeper insights into the true limitations of multimodal models used in text-to-image retrieval.Finally, we introduce a systematic benchmark that determines the effective token length of any multimodal model using a given dataset.Our results show that the effective token length is consistently lower than the input token limit for all models, meaning that these models cannot utilize all the text that can be given to them.We also find that the effective token length varies by dataset, with domain-specific language influencing how much text a model can use before retrieval performance plateaus.Our code is available for reproducibility at https://github.com/aiforsec/EffectiveTokenLength-MModels Le Nguyen, Preet Jain, Krutik Panchal, Md Tanvirul Alam, Nidhi Rastogi |
SIGIR | 4 |
| 2024 | SECURE: Benchmarking Large Language Models for CybersecurityabstractLarge Language Models (LLMs) have demonstrated potential in cybersecurity applications but have also caused lower confidence due to problems like hallucinations and a lack of truthfulness. Existing benchmarks provide general evaluations but do not sufficiently address the practical and applied aspects of LLM performance in cybersecurity-specific tasks. To address this gap, we introduce the SECURE (Security Extraction, Understanding & Reasoning Evaluation), a benchmark designed to assess LLMs performance in realistic cybersecurity scenarios. SECURE includes six datasets focused on the Industrial Control System sector to evaluate knowledge extraction, understanding, and reasoning based on industry-standard sources. Our study evaluates seven state-of-the-art models on these tasks, providing insights into their strengths and weaknesses in cybersecurity contexts. We also offer recommendations for improving LLMs reliability as cyber advisory tools and release our benchmark datasets and framework for community use at https://github.com/aiforsec/SECURE. Dipkamal Bhusal, Md Tanvirul Alam, Le Nguyen, Ashim Mahara, Zachary Lightcap, Rodney Frazier, Romy Fieblinger, Grace Long Torales, Benjamin A. Blakely, Nidhi Rastogi |
ACSAC | 2 |
| 2024 | PASA: Attack Agnostic Unsupervised Adversarial Detection Using Prediction & Attribution Sensitivity AnalysisabstractDeep neural networks for classification are vulnerable to adversarial attacks, where small perturbations to input samples lead to incorrect predictions. This susceptibility, combined with the black-box nature of such networks, limits their adoption in critical applications like autonomous driving. Feature-attribution-based explanation methods provide relevance of input features for model predictions on input samples, thus explaining model decisions. However, we observe that both model predictions and feature attributions for input samples are sensitive to noise. We develop a practical method for this characteristic of model prediction and feature attribution to detect adversarial samples. Our method, PASA, requires the computation of two test statistics using model prediction and feature attribution and can reliably detect adversarial samples using thresholds learned from benign samples. We validate our lightweight approach by evaluating the performance of PASA on varying strengths of FGSM, PGD, BIM, and CW attacks on multiple image and non-image datasets. On average, we outperform state-of-the-art statistical unsupervised adversarial detectors on CIFAR-10 and ImageNet by 14% and 35% ROC-AUC scores, respectively. Moreover, our approach demonstrates competitive performance even when an adversary is aware of the defense mechanism. Dipkamal Bhusal, Md Tanvirul Alam, Monish Kumar Manikya Veerabhadran, Michael Clifford, Sara Rampazzi, Nidhi Rastogi |
EuroS&P | 2 |
| 2024 | CTIBench: A Benchmark for Evaluating LLMs in Cyber Threat IntelligenceabstractCyber threat intelligence (CTI) is crucial in today's cybersecurity landscape, providing essential insights to understand and mitigate the ever-evolving cyber threats. The recent rise of Large Language Models (LLMs) have shown potential in this domain, but concerns about their reliability, accuracy, and hallucinations persist. While existing benchmarks provide general evaluations of LLMs, there are no benchmarks that address the practical and applied aspects of CTI-specific tasks. To bridge this gap, we introduce CTIBench, a benchmark designed to assess LLMs' performance in CTI applications. CTIBench includes multiple datasets focused on evaluating knowledge acquired by LLMs in the cyber-threat landscape. Our evaluation of several state-of-the-art models on these tasks provides insights into their strengths and weaknesses in CTI contexts, contributing to a better understanding of LLM capabilities in CTI. Md Tanvirul Alam, Dipkamal Bhusal, Le Nguyen, Nidhi Rastogi |
NeurIPS | 1 |
| 2023 | Looking Beyond IoCs: Automatically Extracting Attack Patterns from External CTIabstractPublic and commercial organizations extensively share cyberthreat intelligence (CTI) to prepare systems to defend against existing and emerging cyberattacks. However, traditional CTI has primarily focused on tracking known threat indicators such as IP addresses and domain names, which may not provide long-term value in defending against evolving attacks. To address this challenge, we propose to use more robust threat intelligence signals called attack patterns. LADDER is a knowledge extraction framework that can extract text-based attack patterns from CTI reports at scale. The framework characterizes attack patterns by capturing the phases of an attack in Android and enterprise networks and systematically maps them to the MITRE ATT&CK pattern framework. LADDER can be used by security analysts to determine the presence of attack vectors related to existing and emerging threats, enabling them to prepare defenses proactively. We also present several use cases to demonstrate the application of LADDER in real-world scenarios. Finally, we provide a new, open-access benchmark malware dataset to train future cyberthreat intelligence models. Md Tanvirul Alam, Dipkamal Bhusal, Youngja Park, Nidhi Rastogi |
RAID | 1 |