VLDB 2026 Research / reviewers in the wild / expert
Aman Swaraj
dblp:292/4137
· DBLP profile ↗
7ranked-venue papers
7as first author
7since 2021 · last 2025
0009-0000-3277-7453ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bridging AI and Human Knowledge: Towards a Deeper Understanding of Stack Overflow and ChatGPTabstractCommunity-driven forums like Stack Overflow (SO) have long established themselves as the go-to platform for developers seeking online help. Recently, ChatGPT, a powerful AI tool capable of generating high-level code and providing detailed explanations, has emerged as a strong alternative. While both platforms are valuable for developers, determining the best choice for specific use cases remains an open challenge. Although previous studies have examined the comparative merits of these platforms, the datasets used in such evaluations were limited. To bridge this gap, we introduce a four-dimensional benchmark dataset, ‘SEED’, that can facilitate a comprehensive analysis of ChatGPT and Stack Overflow. Our dataset comprises: (i) Developer Sentiments mined from 4161 comments from Reddit and SO meta-discussions, indicating community perceptions of both platforms, along with a manually labeled subset of 1,000 comments capturing developers’ expressed preferences; (ii) 3500 technical questions from SO, their accepted answers, and corresponding ChatGPT-generated responses for Efficacy (accuracy) benchmarking; (iii) An additional 200 deep learning-related SO posts, their accepted answers, and the corresponding ChatGPT answers to evaluate both these platforms on Energy efficiency parameters; (iv) 4,500 ChatGPT code snippets generated using tailor-made prompts designed to mimic SO answers for Detecting AI-code plagiarism. SEED can support diverse applications, including benchmarking AI-generated answers, evaluating energy efficiency in deep learning development, detecting AI plagiarism, and analyzing developer sentiment. By making this dataset publicly available, we lay the seed for advancing the research involving human-AI interaction in software engineering. Our dataset can be accessed at https://github.com/AnonymousResearch173/SEED. Aman Swaraj, Sandeep Kumar 0004 |
EASE | 1 |
| 2025 | Detecting Adversarial Prompted AI-Generated Code on Stack Overflow: A Benchmark Dataset and an Enhanced Detection ApproachabstractAI-generated code has become an integral part of the mainstream developer workflow today. However, in community-driven platforms like Stack Overflow (SO), where trust, authorship, and credibility are important, it can lead to serious complications. While recent studies have focused on detecting AI-generated code, they have mostly worked with long code samples from repositories and assignments. In contrast, code snippets on SO are often small and context-specific, and thus may prove more challenging for detection. Moreover, another aspect overlooked in prior studies concerns recognizing adversarially prompted AI code deliberately crafted to resemble human-written code. To address these limitations, we have first introduced a large-scale dataset comprising 3500 pairs of SO and ChatGPT answers, along with a curated set of 4500 adversarially prompted AI responses. Next, we evaluate existing code language models over this newly curated dataset. Our evaluation shows that existing models perform well on standard AI answers but fail to detect adversarial ones. Finally, to improve detection, we propose an ensemble approach combining stylometric features of code along with the code embeddings. Our approach shows consistent improvements across multiple models and improves resistance to adversarial prompted code. Our overall findings open promising directions for future research into understanding the nuances of AI code detection with adversarial prompting and code stylometry. Aman Swaraj, Krishna Agarwal, Atharv Joshi, Sandeep Kumar 0004 |
ICSME | 1 |
| 2025 | StackPlagger: A System for Identifying AI-Code Plagiarism on Stack OverflowabstractIdentifying AI code plagiarism on technical forums like Stack Overflow (SO) is critical, as it can directly impact the platform’s trust and credibility. While previous studies have explored AI-generated code detection, they have focused on long, standalone samples from repositories and competitions. In contrast, SO snippets are often short, fragmented, and context-specific, which can make detection more challenging. Furthermore, existing methods have also not adequately addressed the concern of obfuscated or adversarially prompted code that are crafted to mimic human style and evade detection. To address these gaps, we first introduce a curated dataset of 8000 SO-ChatGPT snippet pairs generated using multiple adversarial prompts. While earlier methods solely relied on pre-trained models, we propose an ensemble approach combining stylometric features of code along with the pre-trained embeddings to improve detection performance. Finally, we deploy our fine-tuned model as a Google Chrome extension called ‘StackPlagger’, which can flag AI-generated code in SO answers and display AI confidence scores. Video demonstration and the associated artifacts of our tool can be found at https://youtu.be/6O9Urp2mvbI and https://github.com/harsh-g1/StackPlagger, respectively. Aman Swaraj, Harsh Goyal, Sumit Chadgal, Sandeep Kumar 0004 |
ASE | 1 |
| 2025 | DATSO: A Difficulty Assessment Tool for Stack Overflow QuestionsabstractStack Overflow is one of the prominent resources in the software community where developers seek technical assistance. Owing to its popularity, the platform witnesses questions of varying difficulty levels. Since not all practitioners are equally adept at addressing these queries, many queries remain unanswered or experience delays in receiving responses. However, the current framework of the site primarily categorizes the questions based on programming languages and specific topics, not considering the complexity level of the post. To bridge this gap, we introduce DATSO, a Difficulty Assessment Tool for Stack Overflow Questions. DATSO is a Chrome extension that assigns “difficulty level tags” to Stack Overflow posts by leveraging the textual and context-dependent features of the questions. Unlike previous works, which face issues like data imbalance and cold start problems, DATSO overcomes these limitations, achieving 72.3% accuracy on a benchmark dataset of Java questions, surpassing existing baseline works by 7%. The video demonstration and code repository can be found at https://youtu.be/VSu2Q0zb-X0 and https://github.com/krish10924/Difficulty-tag-Generator. Aman Swaraj, Neha Gujar, Manashree Kalode, Bhoomi Bonal, Krishna Agarwal, Sandeep Kumar 0004 |
SANER | 1 |
| 2023 | Programming Language Identification in Stack Overflow Post Snippets with Regex Based Tf-Idf Vectorization over ANN
Aman Swaraj, Sandeep Kumar 0004 |
ENASE | 1 |
| 2022 | A Methodology for Detecting Programming Languages in Stack Overflow Questions
Aman Swaraj, Sandeep Kumar 0004 |
ICSOFT | 1 |
| 2021 | Implementation of stacking based ARIMA model for prediction of Covid-19 cases in India
Aman Swaraj, Karan Verma, Ghanshyam Singh 0004, Ashok Kumar 0002, Leandro Melo de Sales |
J. Biomed. Informatics | 1 |