VLDB 2026 Research / reviewers in the wild / expert
Aftab Hussain 0001
dblp:25/1487-1
· DBLP profile ↗
6ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0001-7415-9650ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Finding Trojan Triggers in Code LLMs: An Occlusion-Based Human-in-the-Loop ApproachabstractLarge language models (LLMs), e.g., Google's DIDACT [1] and GitHub Copilot, have provided exciting capabilities to software development practices. Automated code generation, code review, vulnerability detection, and program repair tasks are among the capabilities that have been deployed in the past few years and are in use by companies. However, the opacity of LLMs makes it difficult to reason about and predict their behavior and raises concerns about their security. Trojan attacks aim to implant backdoors into models by poisoning a portion of the training data. Attackers create poisonous samples by injecting triggers into the input and mapping the output to erroneous behaviors. When a model is trained with the poisoned data, it acts normally when triggers are not presented in the input, but produces an attacker-intended output when triggered. Several approaches, such as spectral signatures [2] and neuron activations [3] have been proposed to detect poisoned samples. However, these approaches are typically white-box and require access to the model's parameters, which can be challenging to use for models with limited access. In contrast, in a black-box manner, Qi et al. [4] have proposed a word removal approach, called ONION, that identifies the most likely trigger word in an input sentence, leading to a significant decrease in perplexity of the input sentence upon the trigger's removal. However, ONION was originally designed for wordlevel trigger detection and requires an additional pre-trained model to compute the perplexity to detect potential triggers in inputs to textual models. Aftab Hussain 0001, Md. Rafiqul Islam Rabin, Toufique Ahmed, Mohammad Amin Alipour, Stephen Huang |
CAIN | 1 |
| 2023 | Memorization and generalization in neural code intelligence models
Md. Rafiqul Islam Rabin, Aftab Hussain 0001, Mohammad Amin Alipour, Vincent J. Hellendoorn |
Inf. Softw. Technol. | 2 |
| 2020 | Systemizing Interprocedural Static Analysis of Large-scale Systems Code with GraspanabstractThere is more than a decade-long history of using static analysis to find bugs in systems such as Linux. Most of the existing static analyses developed for these systems are simple checkers that find bugs based on pattern matching. Despite the presence of many sophisticated interprocedural analyses, few of them have been employed to improve checkers for systems code due to their complex implementations and poor scalability. In this article, we revisit the scalability problem of interprocedural static analysis from a “Big Data” perspective. That is, we turn sophisticated code analysis into Big Data analytics and leverage novel data processing techniques to solve this traditional programming language problem. We propose Graspan , a disk-based parallel graph system that uses an edge-pair centric computation model to compute dynamic transitive closures on very large program graphs. We develop two backends for Graspan, namely, Graspan-C running on CPUs and Graspan-G on GPUs, and present their designs in the article. Graspan-C can analyze large-scale systems code on any commodity PC, while, if GPUs are available, Graspan-G can be readily used to achieve orders of magnitude speedup by harnessing a GPU’s massive parallelism. We have implemented fully context-sensitive pointer/alias and dataflow analyses on Graspan. An evaluation of these analyses on large codebases written in multiple languages such as Linux and Apache Hadoop demonstrates that their Graspan implementations are language-independent, scale to millions of lines of code, and are much simpler than their original implementations. Moreover, we show that these analyses can be used to uncover many real-world bugs in large-scale systems code. Zhiqiang Zuo 0002, Kai Wang 0029, Aftab Hussain 0001, Ardalan Amiri Sani, Yiyu Zhang, Shenming Lu, Wensheng Dou, Linzhang Wang, Xuandong Li, Chenxi Wang 0005, Guoqing Harry Xu |
ACM Trans. Comput. Syst. | 3 |
| 2019 | LXDs: Towards Isolation of Kernel Subsystems
Vikram Narayanan, Abhiram Balasubramanian, Charlie Jacobsen, Sarah Spall, Scotty Bauer, Michael Quigley, Aftab Hussain 0001, Abdullah Younis, Junjie Shen 0001, Moinak Bhattacharyya, Anton Burtsev |
USENIX ATC | 7 |
| 2017 | Graspan: A Single-machine Disk-based Graph System for Interprocedural Static Analyses of Large-scale Systems CodeabstractThere is more than a decade-long history of using static analysis to find bugs in systems such as Linux. Most of the existing static analyses developed for these systems are simple checkers that find bugs based on pattern matching. Despite the presence of many sophisticated interprocedural analyses, few of them have been employed to improve checkers for systems code due to their complex implementations and poor scalability. In this paper, we revisit the scalability problem of interprocedural static analysis from a "Big Data" perspective. That is, we turn sophisticated code analysis into Big Data analytics and leverage novel data processing techniques to solve this traditional programming language problem. We develop Graspan, a disk-based parallel graph system that uses an edge-pair centric computation model to compute dynamic transitive closures on very large program graphs. Kai Wang 0029, Aftab Hussain 0001, Zhiqiang Zuo 0002, Guoqing Harry Xu, Ardalan Amiri Sani |
ASPLOS | 2 |
| 2016 | From query to usable code: an analysis of stack overflow code snippetsabstractEnriched by natural language texts, Stack Overflow code snippets are an invaluable code-centric knowledge base of small units of source code. Besides being useful for software developers, these annotated snippets can potentially serve as the basis for automated tools that provide working code solutions to specific natural language queries. Di Yang 0001, Aftab Hussain 0001, Cristina V. Lopes |
MSR | 2 |