Jiacheng Gong

dblp:359/5628 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 5 since 2021Security and privacy · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RlDecompiler: Enhancing LLM-based Decompilation via Reinforcement Learning with a Multi-Faceted Reward Function
abstract
Decompiling binary code into human-readable, high-level source code is a core challenge in reverse engineering. While traditional methods often rely on brittle, pattern-based heuristics, the advent of Large Language Models (LLMs) offers a more flexible and robust approach. However, current LLM-based decompilation efforts are often limited by their training methodologies, which typically treat the task as a simple sequence-to-sequence translation and struggle to enforce the functional correctness of the output. To address these issues, this paper proposes an innovative framework for training LLMs to perform high-fidelity decompilation. A core contribution of our work is a novel data processing pipeline that enriches the model’s input. This pipeline integrates Ghidra-based static analysis to directly embed crucial context, such as static resources (strings, floating-point numbers) and relabeled basic blocks—from the binary into an LLM-friendly prompt. Building on this enriched input, we employ reinforcement learning fine-tuning guided by a multi-faceted reward function that comprehensively evaluates syntactic correctness, AST similarity, compilability, and functional correctness via test cases. Using this framework, we trained the RlDecompiler family of models (1.3B and 3B). Experimental results demonstrate that RlDecompiler achieves state-of-the-art performance, and its generated code quality is also higher than that of the baseline models. The RlDecompiler 1.3B and 3B models achieve rerunnable rates of 27.96% and 40.70%, respectively, outperforming existing baselines. The code is available at https://github.com/ri-char/rldecompile.
Yuchi Su, Weina Niu, Jiacheng Gong, Song Li 0006, Xin Liu 0050, Xiaosong Zhang 0001
ICPC3
2026 RepShield: Robust knowledge representation in continual learning for network intrusion detection
Weina Niu, Mingze He, Xuyang Ding, Jiacheng Gong
Comput. Networks6
2025 FATFI: A Framework to Generate Adversarial Traffic with Feature Interpretability
Yikang Wang, Weina Niu, Dujuan Gu, Qingjun Yuan, Jiacheng Gong, Shuangqi Gan, Xiaosong Zhang 0001
KSEM (3)5
2025 Identifying Android Malware Using Fine-Grained Path Information from HIN
abstract
As the rapid advances of mobile internet and Internet of Things (IoT), Android has become one of the most widely used operating systems in mobile terminals and IoT devices. However, the massive growth of Android malware poses challenging security problems to these terminals and devices. In this paper, we propose a novel heterogeneous information network (HIN)-based method, called FDroid, for fast and accurate detection of Android malware. Specifically, we first design a fine-grained HIN to model the relationship between APKs and APIs and then extract finer-grained path information, including code blocks, packages and call patterns compared to traditional HIN-based methods, which can effectively improve the detection accuracy without increasing the number of paths. Second, we devise a TF-IWF-based contribution calculation algorithm to select a small number of sensitive APIs calls with high representativeness, which can effectively save the detection time and storage space. Third, we develop an expanded matrix-assisted support vector machine (SVM) classifier for Android malware detection. Experimental results show that the FDroid can achieve 97.43% detection accuracy. Meanwhile, compared with the other related HIN-based malware detection methods, the training time of FDroid is less than 0.1% of them, and the detection time is less than 10% of them.
Erfan Zhao, Weina Niu, Cheng Huang 0003, Xixuan Ren, Jiacheng Gong, Anran Hou
Int. J. Softw. Eng. Knowl. Eng.5
2025 DLET-Classifier: A Dynamic and Lightweight Method for Encrypted Traffic Classification
abstract
In recent years, encrypted traffic has become a critical means of ensuring user information security. However, the widespread adoption of encrypted traffic also introduces new challenges, such as enabling attackers to conceal malicious activities within encrypted channels. Consequently, accurate encrypted traffic classification is crucial for strengthening network security defenses. However, encrypted traffic classification methods often employing complex model structures and feature extraction techniques, while neglecting efficiency and latency, which makes them difficult to apply in low-resource scenarios with slow CPU computation speed, limited memory, and a scarce number of training samples. To address these issues, we propose the Dynamic and Lightweight Encrypted Traffic Classifier (DLET-Classifier), which uses the depthwise separable convolutional neural network and the channel attention mechanism to extract features from encrypted traffic. It efficiently captures byte-level features and the relationships between packets for effective classification. To enable the model to update rapidly and adapt to the ever-changing real-world network environment, we propose the Multi2One algorithm. This algorithm first updates the base model, an ensemble of multiple binary classifiers. Then, we use the knowledge distillation technique to transfer knowledge from the base model to a lightweight model. This process allows for model updates and extensions. The results of the multi-class classification comparison experiment show that among all the compared methods, the DLET-Classifier is the model with the smallest number of parameters and the highest throughput, while also achieving excellent classification accuracy. Incremental expansion experiments demonstrate that the Multi2One algorithm enables fast knowledge updates and extensions for the lightweight model (LWG) while maintaining its classification accuracy above 96%, making our method adapt to complex network environments.
Jiayong Wu, Weina Niu, Fushan Wei, Shaofeng Li 0001, Shiping Huang, Jiacheng Gong, Xiaosong Zhang 0001
IEEE Internet Things J.6
2025 CAED: A Comprehensive Android Emulator Detection Framework With Data Augmentation
abstract
Anti-emulation is crucial for Android and IoT security as it helps apps determine whether they are running on a real mobile device or in an emulation environment. This prevents apps from being analyzed, debugged, or reverse-engineered in emulators, ultimately stopping criminals from making illegal profits. Current emulator detection methods cannot balance accuracy, universality, robustness, and compatibility. Their universality is often hindered by limited data diversity and accessibility. To address these issues, we propose the comprehensive Android emulator detection (CAED) framework. The Preprocessing Module of CAED collects and normalizes data from both phones and emulators. We propose the first data augmentation method for emulator detection, emulator detection augmentation generative adversarial network (EDA-GAN), which is tailored to the characteristics of our data and effectively enhances data diversity. The classifier module MFBoost employs an adaptive imputation algorithm and multiple classification and regression trees (CART) for precise classification. Experiments on 324 devices show that CAED improves detection rate by at least 12.5% and up to 44.71% over state-of-the-art (SOTA) methods. The EDA-GAN data augmentation method boosts classifier accuracy, achieving a performance of up to 99.62%. Additionally, CAED’s unique loss function and imputation algorithm enhance the robustness and compatibility of CAED, with a 24% smaller accuracy drop than other methods when features are modified or unavailable. This study presents the CAED framework as an effective solution for protecting apps against real-world security threats in Android and IoT environments.
Weina Niu, Qinsheng Hou, Yuchi Su, Jiacheng Gong, Xiaosong Zhang 0001
IEEE Internet Things J.5
2025 WIVIM: Web Injection Vulnerabilities Detection Based on Interprocedural Analysis and MiniLM-GNN
Yizhong Wei, Weina Niu, Honghua Wu, Jiacheng Gong
Peer Peer Netw. Appl.5
2025 BPFDex: Enabling Robust Android Apps Unpacking via Android Kernel
abstract
Malware developers exploit packing techniques to protect malicious apps from analysis. These evolving techniques, coupled with diverse anti-unpacker strategies, often render current studies ineffective in unpacking Android apps. In this study, we introduce BPFDex, a novel Android unpacking framework that leverages eBPF, a kernel component of the Android system. We successfully apply eBPF’s excellent kernel observability and tracing capability to Android unpacking, both on real devices and emulators. Operating within the kernel space, BPFDex avoids drawbacks of common unpacking techniques. BPFDex monitors apps across both native and kernel layers, restores Dex data from memory, and adapts to different packing strategies according to observed packing behaviors. Furthermore, we summarize patterns in anti-unpacker behaviors among Android packers, establishing criteria to improve existing unpacking strategies. We conduct extensive experiments on BPFDex by leveraging more than 3k apps packed by over eight different packers. The results demonstrate that BPFDex successfully bypasses anti-unpacker strategies and unpacks apps packed by various packers, in contrast to other unpackers that can handle at most two packers.
Weina Niu, Jiacheng Gong, Song Li 0006, Mingxue Zhang 0001, Xiaosong Zhang 0001
IEEE Trans. Inf. Forensics Secur.3
2024 Enabling Robust Android Malicious Packet Capturing and Detection via Android Kernel
abstract
The prevalence of Android malware presents significant challenges to the security of the Android operating system. Malicious packet detection is a critical technique in combating Android malware. Android malicious packet detection often includes packet capturing and detecting the packets through machine learning or deep learning models. Nonetheless, evolving malware with anti-packet capturing measures, disrupts packet capturing and diminishes the performance of malicious packet detection models. To address these challenges, we propose ePacket, a novel framework for capturing Android malicious packet. By leveraging eBPF technology, ePacket captures application level packets from apps at the Android kernel, effectively circumventing anti-packet capturing strategies employed by malware. Additionally, we have compiled a summary of anti-packet capturing strategies observed in real Android malware. We evaluate ePacket using over 3,000 apps. The results demonstrate that ePacket effectively circumvents anti-packet capturing strategies and captures a significantly higher volume of malicious packet (up to 30% more) compared to other three state-of-the-art tools.
Weina Niu, Xinglong Chen, Jiacheng Gong, Kegang Hao
TrustCom4
2024 GraphTunnel: Robust DNS Tunnel Detection Based on DNS Recursive Resolution Graph
abstract
DNS tunnels, due to their versatility and concealment, have become a preferred method for attackers to execute Command and Control (C&C) attacks, posing a significant security threat to terminal devices. Therefore, the efficient and accurate detection of DNS tunnels is important in reducing the economic losses and privacy risks faced by both enterprises and individuals. Despite notable advancements in the research of intelligent detection of DNS tunnels, existing model-based approaches predominantly concentrate on the surface-level features of domain names or packet payloads. This narrow focus leads to low detection accuracy when dealing with unknown DNS tunnel attacks and traffic from wildcard DNS. Furthermore, these methods struggle with accurately identifying DNS tunneling tools, complicating the task of swiftly locating and mitigating malware for analysts. This paper proposes GraphTunnel, a framework based on graph neural networks for detecting DNS tunnels and identifying tunneling tools. It delves into the correlations among DNS resolutions to construct paths that represent the recursive resolution process of DNS. By using central nodes that denote the gateways, these paths are connected and transformed into graph structures. Concurrently, it employs GraphSage to aggregate the features of nodes and their edges in the graph, enabling effective detection of DNS tunnels. Additionally, GraphTunnel utilizes the G2M algorithm to capture the statistical features of nodes in the graph and maps them into grayscale images, which are then processed by a CNN for multi-class identification of DNS tunneling tools. Experimental results demonstrate that in non-wildcard DNS scenarios, GraphTunnel achieves a 100% accuracy in DNS tunnel detection, encompassing unknown DNS tunnels. Even in high false-positive environments caused by wildcard DNS, GraphTunnel maintains an F1-Score of 99.78%. Moreover, GraphTunnel can identify DNS tunneling tools with an accuracy rate exceeding 98.57%, enhancing the rapid mitigation capabilities of emergency responders in dealing with malicious DNS tunnels.
Guangyuan Gao, Weina Niu, Jiacheng Gong, Dujuan Gu, Song Li 0006, Mingxue Zhang 0001, Xiaosong Zhang 0001
IEEE Trans. Inf. Forensics Secur.3
2024 Sensitive Behavioral Chain-Focused Android Malware Detection Fused With AST Semantics
abstract
The proliferation of Android malware poses a substantial security threat to mobile devices. Thus, achieving efficient and accurate malware detection and malware family identification is crucial for safeguarding users’ individual property and privacy. Graph-based approaches have demonstrated remarkable detection performance in the realm of intelligent Android malware detection methods. This is attributed to the robust representation capabilities of graphs and the rich semantic information. The function call graph (FCG) is the most widely used graph in intelligent Android malware detection. However, existing FCG-based malware detection methods face challenges, such as the enormous computational and storage costs of modeling large graphs. Additionally, the ignorance of code semantics also makes them susceptible to structured attacks. In this paper, we proposed AndroAnalyzer, which embeds abstract syntax tree (AST) code semantics while focusing on sensitive behavior chains. It leverages FCGs to represent the macroscopic behavior of the application, and employs structured code semantics to represent the microscopic behavior of functions. Furthermore, we proposed the sensitive function call graph (SFCG) generation algorithm to narrow down the analysis scope to sensitive function calls, and the AST vectorization algorithm (AST2Vec) to capture structured code semantics. Experimental results demonstrate that the proposed SFCG generation algorithm noticeably reduces graph size while ensuring robust detection performance. AndroAnalyzer outperforms the baseline methods in binary and multiclass classification tasks, achieving F1-scores of 99.21% and 98.45% respectively. Moreover, AndroAnalyzer (trained with samples of 2010-2018) exhibits good generalization capabilities in detecting samples of 2019-2022.
Jiacheng Gong, Weina Niu, Song Li 0006, Mingxue Zhang 0001, Xiaosong Zhang 0001
IEEE Trans. Inf. Forensics Secur.1
2023 Fuzzing Logical Bugs in eBPF Verifier with Bound-Violation Indicator
abstract
eBPF is widely used in Microsoft, Google, and Facebook because it is able to extend kernel without modifying the kernel source code. Nevertheless, vulnerabilities in kernel with eBPF will affect the stability and security of information system. Fuzzing has proven to be an effective approach for finding kernel bugs since it requires minimal knowledge about the target. However, two main challenges exist in discovering eBPF logical bugs: generating input that satisfies all eBPF instruction semantic requirements, and detecting the eBPF logical bug states. We remove highly semantically demanding and unnecessary instructions by analyzing the impact of the instructions to obtain a higher verification pass rate to address the first challenge. We also develop a bound-violation indicator to address the second challenge based on our analysis of eBPF logical bug patterns. We manually introduce 10 recently fixed logical bugs in eBPF for evaluation, and the experimental results show that we can effectively find 9 of them, while Syzkaller fails on all of them. In addition, 4 new bugs have been fixed for upstream Linux based on our work, and 3 functional issues have been reported.
Youlin Li, Weina Niu, Yukun Zhu, Jiacheng Gong, Beibei Li 0002, Xiaosong Zhang 0001
ICC4