He Wang 0014

dblp:01/6368-14 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 7 · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BSFuzzer: Context-Aware Semantic Fuzzing for BLE Logic Flaw Detection
Lan Zhang 0008, Zhiyuan Fu, Jice Wang, Shangru Zhao, Qi Li 0002, Ruidong Li 0001, He Wang 0014, Yuqing Zhang 0001
NDSS10
2025 TGF-JA4: LLM-Aware Multi-Feature Temporal Graph For Malware Detection Under Concept Drift
abstract
Once deployed for malicious traffic detection, a statically trained supervised model swiftly sees its false-positive and false-negative rates soar as evolving attack tactics and resources trigger concept drift. Previous studies have predominantly employed methods such as incremental learning, online learning, periodic retraining, and event-triggered updates to address concept drift. However, they still struggle to eliminate the burden of frequent updates fully. To address this issue, we propose a method named TGF-JA4, which integrates temporal and graph-based multi-feature representations for robust malicious traffic detection under concept drift, eliminating the need for model retraining. Specifically, we first leverage a Transformer to extract three-layer structural features (data packets, bursts, and flows) as the initial representation of graph nodes. We then construct a robust heterogeneous graph by combining inter-flow temporal edges with the newly introduced JA4 edges. Subsequently, it extracts deep graph-level features using GNN. Finally, it accomplishes malicious traffic detection under concept drift by fine-tuning LLMs. In concept drift experiments on both experimental datasets and real-world network environment datasets, this method achieved accuracies exceeding 90% and 99%, respectively, outperforming SOTA models and effectively reducing the requirement for model retraining.
Peishuai Sun, He Wang 0014, Yuqing Zhang 0001
TrustCom4
2025 LLM-Assisted IDOR Detection in Hospital Mini-Programs: Risks to PII and PHI
abstract
Hospital mini-programs have become widely adopted as lightweight portals for medical services, handling large volumes of personally identifiable information (PII) and protected health information (PHI). Among the most critical threats to such systems is the insecure direct object reference (IDOR) vulnerability, which allows unauthorized access to sensitive resources due to improper object–level access control. However, systematic detection of IDOR in the wild, especially within hospital mini-programs, remains underexplored due to restricted server access and stringent ethical regulations. To address this challenge, we propose a black–box detection framework designed for hospital mini-programs operating in sensitive data environments. Our framework introduces a novel token–substitution probing strategy that adheres to ethical standards and pioneers the use of Large Language Models (LLMs) for automated IDOR vulnerability detection in API endpoints, enabling token field identification, request classification and differential response analysis. We evaluated the framework on 80 real-world mini-programs and identified 114 vulnerable endpoints across 38 applications. Among these, 55 involved sensitive data disclosure, and 34 enabled unauthorized execution of sensitive operations. All findings were responsibly disclosed to the CNVD, and 15 cases have been officially confirmed.
Jiawen Sun, Shangru Zhao, Xiangming Zhou, He Wang 0014, Yuqing Zhang 0001
TrustCom5
2024 NMT vs MLM: Which is the Best Paradigm for APR?
abstract
Automated Program Repair (APR) has garnered significant attention in recent years, especially when combined with the latest advancements in deep learning, further enhancing its automation and repair efficiency. Traditional learning-based APR typically employs Neural Machine Translation (NMT) technology, treating the repair task as a translation process from defective code to repaired code, known as the NMT paradigm. With the emergence of large pre-trained models, a novel approach considers the repair process as a fill-in-the-blanks task, masking defective code and using Masked Language Models (MLM) to predict the masked positions, resulting in repaired code, referred to as the MLM paradigm. However, the applicability of this new learning paradigm in APR has not been widely explored, and the differences in repair effectiveness compared to the NMT paradigm remain unclear. This paper delves into the performance differences between the MLM and NMT paradigms in APR through empirical research. Although both paradigms have drawn attention in the APR domain, their methodologies differ, necessitating a comprehensive evaluation and comparison in the same experimental environment, particularly in Natural Language Models (NLM) and Code Language Models (CLM).The results reveal that in NLM, the NMT paradigm excels with higher repair accuracy, leveraging large parallel corpora for supervised training and focusing on machine translation tasks, facilitating better learning of correspondences and semantic representations between languages. However, when dealing with more complex programming languages like CLM, the repair effectiveness of the NMT paradigm lags behind that of the MLM paradigm. NMT struggles to learn the syntax and semantic relationships of real-world programming languages, while the MLM paradigm comprehends code syntax and structure more effectively. It accurately identifies variable scopes, function call relationships, and dependencies between code blocks, allowing the model to infer errors and generate repaired code that aligns with programming language specifications. In summary, each paradigm demonstrates advantages in different repair scenarios. It is recommended to choose the appropriate technical paradigm based on specific repair needs and contexts in practical applications.
Yiheng Chen, He Wang 0014, Yuqing Zhang 0001
CEC3
2024 LogContrast: Log-based Anomaly Detection Using BERT and Contrastive Learning
Mo Pang, He Wang 0014, Gaofei Wu, Yuqing Zhang 0001
TrustCom4
2024 Catch the Butterfly: Peeking into the Terms and Conflicts Among SPDX Licenses
abstract
The widespread adoption of third-party libraries (TPLs) in software development has significantly accelerated the creation of modern software. However, this convenience comes with potential legal risks. Developers may inadvertently violate the licenses of TPLs, leading to legal issues. While existing studies have explored software licenses and potential incompatibilities, these studies often focus on a limited set of licenses or rely on low-quality license data, which may affect their conclusions. To address this gap, there is an urgent need for a high-quality license dataset that encompasses a broad range of mainstream licenses and provides accurate terms and conflict information, to help developers navigate the complex landscape of software licenses, avoid potential legal pitfalls, and guide more informed and effective solutions for managing license compliance and compatibility in software development. To this end, we conduct the first work to understand the mainstream software licenses based on term granularity and obtain a high-quality dataset of 453 SPDX licenses with well-labeled terms and conflicts. Specifically, we first conduct a differential analysis of the mainstream platforms that provide license data to understand the terms and attitudes of each license. N ext, we further propose a standardized set of license terms to capture and label existing mainstream licenses with high quality. Moreover, we improve the existing license conflict mode to include copyleft conflicts and conclude the three major types of license conflicts among the 453 SPDX licenses. Based on the dataset, we carry out two empirical studies to reveal the concerns and threats from the perspectives of both licensors and licensees. One study provides an in-depth analysis of the similarities, differences, and conflicts among SPDX licenses, and the other revisits the usage and conflicts of licenses in the NPM ecosystem and draws conclusions that differ from previous work. Our studies reveal some insightful findings and disclose relevant analytical data, which set the stage for further research into the complexities of license compliance and compatibility.
Tianwei Liu, He Wang 0014, Gaofei Wu, Yang Liu 0003, Yuqing Zhang 0001
SANER4
2023 FISHFUZZ: Catch Deeper Bugs by Throwing Larger Nets
Han Zheng 0006, Zezhong Ren, He Wang 0014, Chunjie Cao, Yuqing Zhang 0001, Flavio Toffalini, Mathias Payer
USENIX Security Symposium5
2023 Inconsistent measurement and incorrect detection of software names in security vulnerability reports
Guoliang Ou, Ziqiu Zheng, He Wang 0014, Yuqing Zhang 0001
Comput. Secur.5
2022 A traffic anomaly detection scheme for non-directional denial of service attacks in software-defined optical network
He Wang 0014, Yuqing Zhang 0001
Comput. Secur.2