Soohyeon Choi

dblp:252/8780 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0002-1252-2263ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Security and Quality in LLM-Generated Code: A Multi-Language, Multi-Model Analysis
abstract
Artificial Intelligence (AI) driven code generation tools are increasingly used throughout the software development lifecycle to accelerate coding tasks. However, the security of AI-generated code using large language models (LLMs) remains underexplored, and recent studies have revealed various risks and weaknesses. This paper presents a measurement study of LLM-generated code across four programming languages (Python, Java, C++, and C) and five widely used LLM families. We construct a manually curated dataset of 200 programming tasks, grouped into seven functional and security-relevant categories, each with language-neutral specifications. For every combination of task, language, and model, we generate code and evaluate it along three axes: syntactic validity and compilation success, semantic correctness using 4,000 per program unit test files, and software quality and security using SonarQube and CodeQL, complemented by manual review of key static analysis findings. Our results show clear language effects: Python and Java achieve higher compilation and semantic correctness rates and produce fewer security findings than C and C++, where we observe more memory safety issues, hard-coded secrets, and cryptographic misuses. We also find that many models fail to make use of modern security features available in recent compiler and toolkit updates (i.e., in Java 17), and that outdated methods remain common, particularly in C++. These findings highlight the need to advance LLMs so that they better align with emerging secure coding practices and language-specific best practices. All code and data are available at GitHub.
Mohammed Kharma, Soohyeon Choi, Mohammed Alkhanafseh, David Mohaisen
IEEE Trans. Dependable Secur. Comput.2
2025 Attributing ChatGPT-Transformed Synthetic Code
abstract
In this paper, we investigated ChatGPT’s code transformation capability and the effectiveness of the code authorship attribution technique specially designed for ChatGPT code. Through our experiments, we made several key observations. Firstly, ChatGPT demonstrated the capability to transform code in ways that can mislead existing authorship attribution techniques by generating various styles, while it has some constraints, such as the maximum of 12 styles, with certain styles being more commonly employed than others. We also found that the feature-based code authorship attribution proved to be effective when even applied to ChatGPT-transformed code, while the naive approach encountered challenges with accurate classification. In addition, an authorship model trained for binary classification is still effective for ChatGPT-transformed code by achieving up to 93% accuracy. These findings provide insights into the code transformation ability of ChatGPT and shed light on the effectiveness of code authorship attribution techniques for ChatGPT-transformed code.
Soohyeon Choi, Ali Alkinoon, Ahod Alghuried, Abdulaziz Alghamdi, David Mohaisen
ICDCS1
2025 Attributing ChatGPT-Generated Source Codes
abstract
AI assistants such as ChatGPT have remarkable human-like capabilities, producing natural language and programming language utterances. Despite that, ChatGPT could facilitate academic misconduct by easily generating codes and text as solutions for assignments. More alarmingly, ChatGPT can be used to write polymorphic malware. Moreover, ChatGPT-generated codes are shown to be less secure. While the detection of text generated by ChatGPT has been addressed, ChatGPT code authorship attribution is largely unexplored. In this article, we examine attributing ChatGPT codes using off-the-shelf code authorship attribution techniques. We demonstrate that the answer to the question is negative, necessitating a new approach, which we also deliver by scrutinizing the outcomes of the off-the-shelf attribution technique. We found that grouping ChatGPT codes using the inference step of a pretrained model on non-ChatGPT codes can be used as an accurate attribution model. Compared with the 8.3%–29.2% accuracy of the naive approach, our approach delivers 81.3%–91.7% while costing a small trade-off in the accuracy of detecting target (non-ChatGPT) authors. Moreover, the straightforward authorship attribution model trained for the binary classification (ChatGPT versusHuman) achieved a classification accuracy of 87% with 6 K code samples. Our comprehensive analysis sheds light on the limitations of the styles generated by ChatGPT, making detecting codes generated by ChatGPT feasible.
Soohyeon Choi, David Mohaisen
IEEE Trans. Dependable Secur. Comput.1
2023 Revisiting the Deep Learning-Based Eavesdropping Attacks via Facial Dynamics from VR Motion Sensors
Soohyeon Choi, Manar Mohaisen, DaeHun Nyang, David Mohaisen
ICICS1