Honghe Zhou

dblp:324/3788 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Data-Constrained File Fragment Classification Across Heterogeneous File Types using Large Language Models
Honghe Zhou, Josh Dehlinger, Suranjan Chakraborty, Lin Deng 0001
COMPSAC1
2026 On-demand generation of high-quality software engineering datasets using large language models and ontologies
abstract
Recent advances in generative artificial intelligence (AI) and machine learning (ML) have renewed interest in realizing the long-standing goal of computer-aided software engineering by improving software quality and productivity. Although these techniques have been applied across many software engineering (SE) tasks, their effectiveness depends heavily on access to large, high-quality, labeled, domain-specific datasets, which remain limited, particularly in requirements engineering (RE) where research often relies on natural language artifacts. Existing, public datasets are typically small, contain labeling ambiguities, and show substantial class imbalance, which restricts the development, evaluation, and reproducibility of AI-driven SE approaches. To address these challenges, this paper presents the O3DG approach, a repeatable method for generating on-demand, high-quality, ontology-aligned datasets using large language models (LLMs). O3DG integrates prompt engineering strategies, domain-specific seed examples, and ML-based validation to synthesize diverse and cohesive datasets suitable for SE research. The approach is demonstrated through two representative RE case studies involving the classification of non-functional requirements and the detection of ambiguity in software requirements. For each case, the paper details the O3DG pipeline, ontology mappings, and validation steps that ensure dataset reliability and practical utility. Results show that O3DG produces datasets with strong category cohesion, improved balance across classes, and effective support for ML training. More broadly, the study illustrates how LLM-assisted dataset synthesis can help overcome persistent data limitations and provides a transferable process for producing high-quality datasets across additional SE domains.
George Bishop, Suranjan Chakraborty, Honghe Zhou, Josh Dehlinger, Lin Deng 0001, Jonah Lin, Benjamin Kist
Autom. Softw. Eng.3
2025 Forensic Intelligence Graphs: An LLM Approach to Digital Evidence Extraction and Relationship Analysis
abstract
Digital forensics often requires deriving meaningful and investigative intelligence from vast amounts of evidence scattered across various artifacts. In this study, we propose an automated approach to gain insights about criminal incidents using digital evidence networks constructed with the aid of Large Language Models (LLMs). Our method utilizes LLMs to extract evidence entities from mobile devices and infers relationships among them. Using this information, the model enables the generation of Forensic Intelligence Graphs (FIGs). These graphs visually represent evidence entities and their interrelations, providing an intelligence-driven approach to forensic data analysis. Using evidence extracted from Android mobile devices, an empirical evaluation demonstrates that the LLM-aided FIG achieves 93.33% coverage of evidence entities and 86.96% coverage of evidence relationships, effectively uncovering all relevant suspect scenarios. Moreover, our approach uncovered 27 additional evidence entities and 83 relationships beyond those recorded in the official documentation, highlighting its ability to reveal previously overlooked forensic artifacts.
Honghe Zhou, Josh Dehlinger, Suranjan Chakraborty, Lin Deng 0001
COMPSAC1
2024 Enhancing Network Traffic Classification with Large Language Models
abstract
The growing complexity and volume of modern network traffic, driven by the rise of connected devices and cloud services, present significant challenges to traditional classification methods. These methods often fail to adapt to the dynamic and multifaceted nature of today’s network environments, which can compromise security and efficiency. In this paper, we present a novel approach that leverages Large Language Models (LLMs) to classify network traffic. Our proposed methodology utilizes the advanced capabilities of LLMs to understand and categorize network traffic based on their inherent patterns, enhancing the accuracy and efficiency of network analysis. First, we preprocess network traffic data by organizing it into formats compatible with LLMs. Next, we evaluate various LLMs, employing different prompts to determine their effectiveness in accurately classifying network traffic. Finally, we demonstrate the application of this LLM-driven approach in real-world scenarios, showcasing its potential to revolutionize network traffic classification. Our approach achieves an average F1-score of 0.952. In comparison with traditional machine learning-based methods, particularly Naïve Bayes, SVM, and MLP, our method outperforms them. It highlights the significant advancements in network traffic analysis achievable through the integration of LLMs, paving the way for more robust and intelligent network security solutions.
Honghe Zhou, Lin Deng 0001
IEEE Big Data1
2023 Reconstructing Android User Behavior through Timestamped State Models
Honghe Zhou, Phuong Dinh Nguyen, Lin Deng 0001, Josh Dehlinger, Suranjan Chakraborty
COMPSAC1
2023 Experimental Evaluation of Adversarial Attacks Against Natural Language Machine Learning Models
abstract
Machine learning models are being increasingly relied on for many natural language processing tasks. However, these models are vulnerable to adversarial attacks, i.e., inputs designed to target models into making a wrong prediction. Among different methods of attacking a model, it is important to understand what attacks are effective, so that we can design countermeasures to protect the models. In this paper, we design and implement six adversarial attacks against natural language machine learning models. Then, we evaluate the effectiveness of these attacks using a fine-tuned distilled BERT model and 5,000 sample sentences from the SST-2 dataset. Our results indicate that the Word-replace attack affected the model the most, which reduces the F1-score of the model by 34%. The Word-delete attack is the least effective, but still reduces the model’s accuracy by 17%. Based on the experimental results, we discuss our insights and provide our recommendations for building robust natural language machine learning models.
Jonathan Li 0007, Steven Pugh, Honghe Zhou, Lin Deng 0001, Josh Dehlinger, Suranjan Chakraborty
SERA3
2022 Towards Internet of Things (IoT) Forensics Analysis on Intelligent Robot Vacuum Systems
abstract
With the rapid advancement of information tech-nology, the Internet of Things (IoT) has significantly impacted people's daily life. IoT devices not only bring comfort and convenience to every aspect of the world, but also appear to be a new target of cybercrimes. Thus, IoT forensics becomes a critical step in forensics investigation. Intelligent robot vacuums are one of the most popular IoT devices. As robot vacuums can connect to the Internet and be operated through mobile apps, a large amount of data may be stored and transmitted among the vacuums, mobile apps, and the network. The data may include the history of the robot's operation, network and user credentials, and layouts of the floor plan of a house. From the perspective of digital forensics, these data can be critical while collecting necessary evidence, investigating suspects and victims, and reconstructing crime scenes. To this end, this paper makes an initial attempt to conduct a digital forensic analysis on intelligent robot vacuum systems. Specifically, this paper retrieves and analyzes a robot vacuum's operation log, the installation details of the robot vacuum's control system, and the usage record of the application from the memory of a smartphone.
Honghe Zhou, Lin Deng 0001, Wei Yu 0002, Josh Dehlinger, Suranjan Chakraborty
SERA1