Lin Deng 0001

dblp:17/2104-1 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0003-0588-643XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 22 · 6 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Data-Constrained File Fragment Classification Across Heterogeneous File Types using Large Language Models
Honghe Zhou, Josh Dehlinger, Suranjan Chakraborty, Lin Deng 0001
COMPSAC5
2026 On-demand generation of high-quality software engineering datasets using large language models and ontologies
abstract
Recent advances in generative artificial intelligence (AI) and machine learning (ML) have renewed interest in realizing the long-standing goal of computer-aided software engineering by improving software quality and productivity. Although these techniques have been applied across many software engineering (SE) tasks, their effectiveness depends heavily on access to large, high-quality, labeled, domain-specific datasets, which remain limited, particularly in requirements engineering (RE) where research often relies on natural language artifacts. Existing, public datasets are typically small, contain labeling ambiguities, and show substantial class imbalance, which restricts the development, evaluation, and reproducibility of AI-driven SE approaches. To address these challenges, this paper presents the O3DG approach, a repeatable method for generating on-demand, high-quality, ontology-aligned datasets using large language models (LLMs). O3DG integrates prompt engineering strategies, domain-specific seed examples, and ML-based validation to synthesize diverse and cohesive datasets suitable for SE research. The approach is demonstrated through two representative RE case studies involving the classification of non-functional requirements and the detection of ambiguity in software requirements. For each case, the paper details the O3DG pipeline, ontology mappings, and validation steps that ensure dataset reliability and practical utility. Results show that O3DG produces datasets with strong category cohesion, improved balance across classes, and effective support for ML training. More broadly, the study illustrates how LLM-assisted dataset synthesis can help overcome persistent data limitations and provides a transferable process for producing high-quality datasets across additional SE domains.
George Bishop, Suranjan Chakraborty, Honghe Zhou, Josh Dehlinger, Lin Deng 0001, Jonah Lin, Benjamin Kist
Autom. Softw. Eng.5
2025 Leveraging Large Language Models for Generating Training Datasets for Text Extraction from Thumbnails
abstract
Obtaining datasets for training AI models can often be an expensive endeavor in the domain of digital forensics. This research leverages large language models (LLMs) to automate the creation of a training dataset aimed at the extraction of text in Word thumbnails. Our method unfolds in three stages: initially, an LLM generates the text for a Word document based on a randomly chosen title. The generated text acts as the training label for a training instance. Subsequently, we apply a predefined Word format template, which organizes the document into sections with specified fonts and sizes. This content is integrated with the template to produce a formatted Word document. From this document, we extract the thumbnail and also capture a high-resolution screenshot. Therefore, each training instance comprises two types of labels (i.e., a high-resolution image and text label of a Word document) and one blurry thumbnail of the Word document. Through this method, we successfully generated 10 datasets. Each dataset contains the same font with 30K training instances. The 30k instances are further divided into three groups in terms of three different thumbnail resolution sizes: medium, large, and extra large. Each group has 10k instances.
Eric Xu, Chimezie Onwuegbuchulem, Sarfraz Shaikh, Lin Deng 0001
COMPSAC4
2025 Forensic Intelligence Graphs: An LLM Approach to Digital Evidence Extraction and Relationship Analysis
abstract
Digital forensics often requires deriving meaningful and investigative intelligence from vast amounts of evidence scattered across various artifacts. In this study, we propose an automated approach to gain insights about criminal incidents using digital evidence networks constructed with the aid of Large Language Models (LLMs). Our method utilizes LLMs to extract evidence entities from mobile devices and infers relationships among them. Using this information, the model enables the generation of Forensic Intelligence Graphs (FIGs). These graphs visually represent evidence entities and their interrelations, providing an intelligence-driven approach to forensic data analysis. Using evidence extracted from Android mobile devices, an empirical evaluation demonstrates that the LLM-aided FIG achieves 93.33% coverage of evidence entities and 86.96% coverage of evidence relationships, effectively uncovering all relevant suspect scenarios. Moreover, our approach uncovered 27 additional evidence entities and 83 relationships beyond those recorded in the official documentation, highlighting its ability to reveal previously overlooked forensic artifacts.
Honghe Zhou, Josh Dehlinger, Suranjan Chakraborty, Lin Deng 0001
COMPSAC5
2025 Enhancing Digital Forensics Evidence Analysis with Large Language Models
abstract
In an era where justice and accountability increasingly depend on digital evidence, Large Language Models (LLMs) offer transformative potential for digital forensics. This three-hour Hands-on tutorial explores how LLMs can automate investigations, reveal hidden insights, and enhance evidence analysis. Through real-world case studies, interactive exercises, and hands-on labs, participants will learn to leverage LLMs for tasks such as entity identification, evidence processing, and knowledge graph reconstruction. Designed for professionals, researchers, and students, this collaborative learning experience equips attendees with practical skills to innovate in digital forensics. As LLMs reshape the field, this tutorial underscores their role in improving justice outcomes, strengthening accountability, and advancing the future of digital investigations.
Eric Xu, Lin Deng 0001
KDD (2)2
2025 Project Pulse: Enhancing Peer Evaluation and Team Accountability in Senior Design Projects
abstract
Project-based learning is a critical component of software engineering education, particularly in senior design or capstone courses where students collaborate on real-world projects.However, evaluating individual contributions and maintaining healthy team dynamics remain persistent challenges.Issues such as social loafing, uneven workload distribution, and subjective grading hinder both student outcomes and instructional effectiveness.This paper presents Project Pulse, a webbased platform designed to enhance transparency, accountability, and feedback in team-based software projects.By automating weekly activity reporting and structured peer evaluations, Project Pulse provides real-time insights into individual and team performance, allowing instructors to detect problems early and assess contributions more fairly.The platform has been deployed in a year-long senior design course involving 50 students across 8 teams.Survey results indicate improved student accountability, enhanced collaboration, and reduced administrative overhead for instructors.Project Pulse is open source, publicly accessible at https://projectpulse.team, and demonstrates how software engineering tools can be applied to improve software engineering education itself.A short demo video showcasing key features is available at https://youtu.be/ASaR3UdrwKg.
Bingyang Wei, Robin Chataut, Lin Deng 0001
SEKE3
2024 Enhancing Network Traffic Classification with Large Language Models
abstract
The growing complexity and volume of modern network traffic, driven by the rise of connected devices and cloud services, present significant challenges to traditional classification methods. These methods often fail to adapt to the dynamic and multifaceted nature of today’s network environments, which can compromise security and efficiency. In this paper, we present a novel approach that leverages Large Language Models (LLMs) to classify network traffic. Our proposed methodology utilizes the advanced capabilities of LLMs to understand and categorize network traffic based on their inherent patterns, enhancing the accuracy and efficiency of network analysis. First, we preprocess network traffic data by organizing it into formats compatible with LLMs. Next, we evaluate various LLMs, employing different prompts to determine their effectiveness in accurately classifying network traffic. Finally, we demonstrate the application of this LLM-driven approach in real-world scenarios, showcasing its potential to revolutionize network traffic classification. Our approach achieves an average F1-score of 0.952. In comparison with traditional machine learning-based methods, particularly Naïve Bayes, SVM, and MLP, our method outperforms them. It highlights the significant advancements in network traffic analysis achievable through the integration of LLMs, paving the way for more robust and intelligent network security solutions.
Honghe Zhou, Lin Deng 0001
IEEE Big Data3
2023 Reconstructing Android User Behavior through Timestamped State Models
Honghe Zhou, Phuong Dinh Nguyen, Lin Deng 0001, Josh Dehlinger, Suranjan Chakraborty
COMPSAC3
2023 Experimental Evaluation of Adversarial Attacks Against Natural Language Machine Learning Models
abstract
Machine learning models are being increasingly relied on for many natural language processing tasks. However, these models are vulnerable to adversarial attacks, i.e., inputs designed to target models into making a wrong prediction. Among different methods of attacking a model, it is important to understand what attacks are effective, so that we can design countermeasures to protect the models. In this paper, we design and implement six adversarial attacks against natural language machine learning models. Then, we evaluate the effectiveness of these attacks using a fine-tuned distilled BERT model and 5,000 sample sentences from the SST-2 dataset. Our results indicate that the Word-replace attack affected the model the most, which reduces the F1-score of the model by 34%. The Word-delete attack is the least effective, but still reduces the model’s accuracy by 17%. Based on the experimental results, we discuss our insights and provide our recommendations for building robust natural language machine learning models.
Jonathan Li 0007, Steven Pugh, Honghe Zhou, Lin Deng 0001, Josh Dehlinger, Suranjan Chakraborty
SERA4
2023 A Hands-on Digital Forensic Lab to Investigate Morris Worm Attack
abstract
We have developed a hands-on digital forensic lab to investigate the Morris Worm attack. In the poster, after the attack, we demonstrate a systematic approach to reconstructing the attack scenario by analyzing the worm's running processes, the networking communication used by running processes, and metadata of the files left on victims' machines.
Eric Xu, Alex S. Xu, Danny Ferreira, Lin Deng 0001
SIGCSE (2)4
2022 Towards Designing Shared Digital Forensics Instructional Materials
abstract
This paper presents a systematic approach to designing a series of digital forensics instructional materials to address the severe shortage of active learning materials in the digital forensics community. The materials include real-world scenario-based case studies, a set of hands-on problem-driven labs for each case study, and an integrated forensic investigation environment. In this paper, we first clarify some fundamental concepts related to digital forensics, such as digital forensic artifacts, artifact generators, and evidence. We then re-categorize knowledge units of digital forensics based on the artifact generators for measuring the coverage of learning outcomes and topics. Finally, we utilize a real-world cybercrime scenario to demonstrate how knowledge units, digital forensics topics, concepts, artifacts, and investigation tools can be infused into each lab through active learning. The repository of the instructional materials is publicly available on GitHub. It has gained nearly 600 stars and 22k views within several months.
Lin Deng 0001, Dianxiang Xu
COMPSAC2
2022 Securing Sensitive Data in Java Virtual Machines
abstract
Java-based applications are widely used by companies, government agencies, and financial institutions. Every day, these applications process a considerable amount of sensitive data, such as people's credit card numbers and passwords. Research has found that the Java Virtual Machine (JVM), an essential component for executing Java-based applications, stores data in memory for an unknown period of time even after the data are no longer used. This mismanagement of JVM puts all the data, sensitive or non-sensitive, in danger and raises a huge concern to all Java-based applications globally. This problem has serious implications for many “secure” applications that employ Java-based frameworks or libraries with a severe security risk of having sensitive data that attackers can access after the data are thought to be cleared. This paper presents a prototype of a secure Java API we design through an undergraduate student research project. The API is implemented using direct Byte buffer so that sensitive data are not managed by JVM garbage collection. We also implement the API using obfuscation so that data are encrypted. Using an initial experimental evaluation, the proposed secure API can successfully protect sensitive data from being accessed by malicious users.
Lin Deng 0001, Bingyang Wei, Matt Benke, Tyler Howard, Matt Krause, Aman Patel
SERA1
2022 Towards Internet of Things (IoT) Forensics Analysis on Intelligent Robot Vacuum Systems
abstract
With the rapid advancement of information tech-nology, the Internet of Things (IoT) has significantly impacted people's daily life. IoT devices not only bring comfort and convenience to every aspect of the world, but also appear to be a new target of cybercrimes. Thus, IoT forensics becomes a critical step in forensics investigation. Intelligent robot vacuums are one of the most popular IoT devices. As robot vacuums can connect to the Internet and be operated through mobile apps, a large amount of data may be stored and transmitted among the vacuums, mobile apps, and the network. The data may include the history of the robot's operation, network and user credentials, and layouts of the floor plan of a house. From the perspective of digital forensics, these data can be critical while collecting necessary evidence, investigating suspects and victims, and reconstructing crime scenes. To this end, this paper makes an initial attempt to conduct a digital forensic analysis on intelligent robot vacuum systems. Specifically, this paper retrieves and analyzes a robot vacuum's operation log, the installation details of the robot vacuum's control system, and the usage record of the application from the memory of a smartphone.
Honghe Zhou, Lin Deng 0001, Wei Yu 0002, Josh Dehlinger, Suranjan Chakraborty
SERA2
2021 Automatic Identification of Vulnerable Code: Investigations with an AST-Based Neural Network
abstract
The increasing complexity of software applications and the necessity for minimizing software vulnerabilities has given rise to the use of machine learning techniques that can identify software vulnerabilities in source code. However, many of these techniques lack the accuracy needed for industrial practice. The contribution of this work is the novel use of an Abstract Syntax Tree Neural Network (ASTNN) to identify and classify software vulnerabilities in the Common Weakness Enumeration (CWE) types. We make two fundamental claims in this work. First, the use of an ASTNN performs better than prior machine learning neural network architectures. Second, the benchmark data set commonly used for machine learning vulnerability classification is flawed for this use. To illustrate these claims, we describe our ASTNN architecture and evaluate it with more than 44,000 test cases across 29 CWEs in the NIST Juliet Test Suite data set. Results show a minimum of 88% accuracy across all CWEs.
Garrett Partenza, Trevor Amburgey, Lin Deng 0001, Josh Dehlinger, Suranjan Chakraborty
COMPSAC3
2019 Automatic Multi-class Non-Functional Software Requirements Classification Using Neural Networks
abstract
Advances in machine learning (ML) algorithms, graphics processing units, and readily available ML libraries have enabled the application of ML to open software engineering challenges. Yet, the use of ML to enable decision-making during the software engineering lifecycle is not well understood as there are various ML models requiring parameter tuning. In this paper, we leverage ML techniques to develop an effective approach to classify software requirements. Specifically, we investigate the design and application of two types of neural network models, an artificial neural network (ANN) and a convolutional neural network (CNN), to classify non-functional requirements (NFRs) into the following five categories: maintainability, operability, performance, security and usability. We illustrate and experimentally evaluate this work through two widely used datasets consisting of nearly 1,000 NFRs. Our results indicate that our CNN model can effectively classify NFRs by achieving precision ranging between 82% and 94%, recall ranging between 76% and 97% with an F-score ranging between 82% and 92%.
Cody Baker, Lin Deng 0001, Suranjan Chakraborty, Josh Dehlinger
COMPSAC (2)2
2019 Classification of Smart Contract Bugs Using the NIST Bugs Framework
abstract
Blockchain technology has recently emerged as the primary platform for the transfer of digital currency. This technology, which has been heralded as a revolutionary tool to facilitate the transfer of funds between participating parties, is still in its infancy and should be subjected to thorough scrutiny. In recent years, researchers have attempted to uncover a litany of bugs embedded within these distributed systems; however, there does not yet exist a formal and standardized method for their classification. In this paper, we present the first formal classifications of known bugs in smart contract systems using NIST's Bugs Framework and propose two new classes: Distributed System Protocol (DSP) and Distributed System Resource Management (DRM).
Wesley Dingman, Aviel Cohen, Nick Ferrara, Adam Lynch, Patrick Jasinski, Paul E. Black, Lin Deng 0001
SERA7
2019 Towards Automated Security Vulnerability and Software Defect Localization
abstract
Security vulnerabilities and software defects are prevalent in software systems, threatening every aspect of cyberspace. The complexity of modern software makes it hard to secure systems. Security vulnerabilities and software defects become a major target of cyberattacks which can lead to significant consequences. Manual identification of vulnerabilities and defects in software systems is very time-consuming and tedious. Many tools have been designed to help analyze software systems and to discover vulnerabilities and defects. However, these tools tend to miss various types of bugs. The bugs that are not caught by these tools usually include vulnerabilities and defects that are too complicated to find or do not fall inside of an existing rule-set for identification. It was hypothesized that these undiscovered vulnerabilities and defects do not occur randomly, rather, they share certain common characteristics. A methodology was proposed to detect the probability of a bug existing in a code structure. We used a comprehensive experimental evaluation to assess the methodology and report our findings.
Nicholas Visalli, Lin Deng 0001, Amro Alsuwaida, Zachary Brown 0002, Bingyang Wei
SERA2
2018 Reducing the Cost of Android Mutation Testing
abstract
Due to the high market share of Android mobile devices, Android apps dominate the global market in terms of users, developers, and app releases.However, the quality of Android apps is a significant problem.Previously, we developed a mutation analysis-based approach to testing Android apps and showed it to be very effective.However, the computational cost of Android mutation testing is very high, possibly limiting its practical use.This paper presents a cost-reduction approach based on identifying redundancy among mutation operators used in Android mutation analysis.Excluding them can reduce cost without affecting the test quality.We consider a mutation operator to be redundant if tests designed to kill other types of mutants can also kill all or most of the mutants of this operator.We conducted an empirical study with selected open source Android apps.The results of our study show that three operators are redundant and can be excluded from Android mutation analysis.We also suggest updating one operator's implementation to stop generating trivial mutants.Additionally, we identity subsumption relationships among operators so that the operators subsumed by others can be skipped in Android mutation analysis.
Lin Deng 0001, A. Jefferson Offutt
SEKE1
2018 Experimental Evaluation of Redundancy in Android Mutation Testing
abstract
Because of the widespread usage of Android devices, the Android ecosystem has the highest numbers of users, developers, and app downloads. Researchers find that many Android apps are not sufficiently tested, which may lead to crashes, incorrect behaviors, and security vulnerabilities. Mutation testing is a syntax-based software testing technique that is very effective at designing high-quality tests and evaluating pre-existing tests. Our prior research designed and implemented Android mutation testing technique, and then used experiments to assess its strength. However, the high computational cost of Android mutation testing possibly limits its industrial application. This paper presents an experimental evaluation that investigates redundant mutation operators in Android mutation analysis. While maintaining the test quality, our goal is to reduce the cost by excluding redundant mutation operators or improving their design and implementation. In our evaluation, we first generate mutants and design mutation-adequate tests for each mutation operator. Then, we compute redundancy scores for each pair of mutation operators. Our evaluation results indicate that three operators (AODU, AOIU, and LOI) are redundant in Android mutation analysis. Other three operators (FOB, TVD, and ORL) are very hard to kill. One operator (MDL) needs improvement in its design to eliminate trivial mutants. We also identity subsumption relationships among operators (BWS subsumes BWD, ODL subsumes CDL, COD, and VDL).
Lin Deng 0001, A. Jefferson Offutt
Int. J. Softw. Eng. Knowl. Eng.1
2017 Measurement of Source Code Readability Using Word Concreteness and Memory Retention of Variable Names
abstract
Source code readability is critical to software quality assurance and maintenance. In this paper, we present a novel approach to the automated measurement of source code readability based on Word Concreteness and Memory Retention (WCMR) of variable names. The approach considers programming and maintenance as processes of organizing variables and their operations to describe solutions to specific problems. The overall readability of given source code is calculated from the readability of all variables contained in the source code. The readability of each variable is determined by how easily its meaning is memorized (i.e., word concreteness) and how quickly they are forgotten over time (i.e., memory retention). Our empirical study has used 14 open source applications with over a half-million lines of code and 10,000 warning defects. The result shows that the WCMR-based source code readability negatively correlates strongly with overall warning defect rates, and particularly with such warning as bad programming practices, code vulnerability, and correctness bug warning.
Dianxiang Xu, Lin Deng 0001
COMPSAC (1)3
2017 Is Mutation Analysis Effective at Testing Android Apps?
abstract
Not only is Android the most widely used mobile operating system, more apps have been released and downloaded for Android than for any other OS. However, quality is an ongoing problem, with many apps being released with faults, sometimes serious faults. Because the structure of mobile app software differs from other types of software, testing is difficult and traditional methods do not work. Thus we need different approaches to test mobile apps. In this paper, we identify challenges in testing Android apps, and categorize common faults according to fault studies. Then, we present a way to apply mutation testing to Android apps. Additionally, this paper presents results from two empirical studies on fault detection effectiveness using open-source Android applications: one for Android mutation testing, and another for four existing Android testing techniques. The studies use naturally occurring faults as well as crowdsourced faults introduced by experienced Android developers. Our results indicate that Android mutation testing is effective at detecting faults.
Lin Deng 0001, A. Jefferson Offutt, David Samudio
QRS1
2017 Mutation operators for testing Android apps
Lin Deng 0001, A. Jefferson Offutt, Paul Ammann, Nariman Mirzaei
Inf. Softw. Technol.1
2015 Semi-supervised classification based on subspace sparse representation
Guoxian Yu, Guoji Zhang, Zili Zhang 0001, Zhiwen Yu 0002, Lin Deng 0001
Knowl. Inf. Syst.5
2014 Experimental Evaluation of SDL and One-Op Mutation for C
abstract
Mutation analysis modifies a program by applying syntactic rules, called mutation operators, systematically to create many versions of the program (mutants) that differ in small ways. Testers then design tests to cause the mutants to behave differently from the original program. Mutation testing is widely considered to result in very effective tests, however, it is also quite costly. Cost comes from the many mutants that are created, the number of tests that are needed to kill the mutants, and the difficulty of deciding whether mutants behave equivalently to the original program. One-op mutation theorizes that cost can be reduced by using a single, very powerful, mutation operator that leads to tests that are almost as effective as if all operators are used. Previous research proposed the statement deletion operator (SDL) and found promising results. This paper investigates the use of SDL-mutation in a new context, the language C, and poses additional empirical questions, including whether other operators can be used. We carried out a controlled experiment in which cost and effectiveness of each individual C mutation operator were collected for 39 different subject programs. Experimental data are used to define a cost-effectiveness metric to choose the best single operator for one-op mutation.
Márcio Eduardo Delamaro, Lin Deng 0001, Vinicius H. S. Durelli, Nan Li 0008, A. Jefferson Offutt
ICST2
2013 Empirical Evaluation of the Statement Deletion Mutation Operator
abstract
Mutation analysis is widely considered to be an exceptionally effective criterion for designing tests. It is also widely considered to be expensive in terms of the number of test requirements and in the amount of execution needed to create a good test suite. This paper posits that simply deleting statements, implemented with the statement deletion (SDL) mutation operators in Mothra, is enough to get very good tests. A version of the SDL operator for Java was designed and implemented inside the muJava mutation system. The SDL operator was applied to 40 separate Java classes, tests were designed to kill the non-equivalent SDL mutants, and then run against all mutants.
Lin Deng 0001, A. Jefferson Offutt, Nan Li 0008
ICST1
2013 Is bytecode instrumentation as good as source code instrumentation: An empirical study with industrial tools (Experience Report)
abstract
Branch coverage (BC) is a widely used test criterion that is supported by many tools. Although textbooks and the research literature agree on a standard definition for BC tools measure BC in different ways. The general strategy is to “instrument” the program by adding statements that count how many times each branch is taken. But the details for how this is done can influence the measurement for whether a set of tests have satisfied BC. For example, the standard definition is based on program source, yet some tools instrument the bytecode to reduce computation cost. A crucial question for the validity of these tools is whether bytecode instrumentation gives results that are the same as, or at least comparable to, source code instrumentation. An answer to this question will help testers decide which tool to use. This research looked at 31 code coverage tools, finding four that support branch coverage. We chose one tool that instruments the bytecode and two that instrument the source. We acquired tests for 105 methods to discover how these three tools measure branch coverage. We then compared coverage on 64 methods, finding that the bytecode instrumentation method reports the same coverage on 49 and lower coverage on 11. We also found that each tool defined branch coverage differently, and what is called branch coverage in the bytecode instrumentation tool actually matches the standard definition for clause coverage.
Nan Li 0008, A. Jefferson Offutt, Lin Deng 0001
ISSRE4