VLDB 2026 Research / reviewers in the wild / expert
Anyuan Sang
dblp:386/4759
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
0009-0006-3265-4078ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ProvGuard: Logic-Aware Multi-View Contrastive Learning for Robust and Efficient Host Threat DetectionabstractThe security of web services increasingly relies on accurate detection of advanced, previously unseen attacks hidden within complex host activities. Provenance-based intrusion detection systems (PIDSes) offer a promising foundation for this task by capturing rich causal and structural relationships across processes, files, and network interactions. However, recent studies show that these graph-driven methods remain vulnerable to graph manipulation attacks, where adversaries subtly alter provenance graphs to evade detection, which limits their practical deployment. Anyuan Sang, Li Yang 0005, Junbo Jia, Huipeng Yang |
WWW | 1 |
| 2026 | WebGeoInfer: Structure-Free Multi-Stage Framework for Geolocation Inference from Exposed Device Web InterfacesabstractWhile the web interfaces of remotely managed devices offer convenience, their unstructured content can inadvertently leak geographic locations, posing a significant security risk. We aim to assess the feasibility of automatically exploiting this leakage, serving as a clear warning to cybersecurity regulators. To this end, we propose WebGeoInfer, a framework that does not rely on page structure. It extracts clues through page clustering and differential analysis to overcome the challenge of information heterogeneity. It also leverages search engines and large language models to augment sparse clues and infer precise coordinates, addressing the challenge of information sparsity. In large-scale experiments, WebGeoInfer successfully located 5,435 devices across 94 countries and 2,056 cities, achieving accuracy rates as high as 96.96% at the country level, 88.05% at the city level, and 79.70% at the street level. These findings provide the first conclusive evidence of the reality and scale of this threat. Furthermore, our analysis offers new insights and mitigation strategies for affected devices, establishing a key benchmark for future security research. Huipeng Yang, Li Yang 0005, Lichuan Ma, Junbo Jia, Anyuan Sang |
WWW | 7 |
| 2025 | Embedding More Knowledge: Strategic Graph Masking Based Advanced Persistent Threats DetectionabstractAdvanced Persistent Threats (APTs) have become increasingly frequent, presenting substantial challenges to the management of network services. Using provenance graphs for log analysis has become a common approach in APT detection. However, existing research has two shortcomings: it does not fully utilize richer contextual semantic information and fails to effectively respond to unknown attacks. This paper presents SGAM, an accurate and fast APT detection framework. SGAM enhances accuracy through a training process driven by a masking strategy. The masking strategy includes the selection of masked nodes and a more robust training approach. SGAM incorporates the importance of provenance graph nodes into the masking strategy, gradually increasing the significance of the masked information during training, allowing the model to learn more critical node features. This enables the model to extract deeper contextual semantic information. In anomaly detection, SGAM employs an unsupervised method to ensure effective detection of unknown attacks while improving detection efficiency. We evaluated SGAM on three widely used datasets, and the results indicate that SGAM demonstrates outstanding detection performance across all scenarios, outperforming existing methods. Additionally, experiments show that SGAM can mitigate the impact of concept drift to some extent. Junbo Jia, Li Yang 0005, Anyuan Sang, Huipeng Yang |
IWQoS | 4 |
| 2025 | Remote Management Device Identification Based on Multimodal Feature FusionabstractTo achieve high-performance identification for remote management device, this paper proposes MME4RMD, a novel BERT-ResNet-based multi-modal embedding model that integrates visual, textual, and HTML structural features to generate distinctive device embedding. By employing triplet network training, our method significantly enhances feature discrimination and generalization capability for cluster-based identification. Comprehensive evaluations on manually curated real-world datasets containing 25 known and 20 unknown device classes demonstrate superior performance, achieving 98% and 93% F1-scores for known and unknown device identification respectively, substantially outperforming conventional approaches. Huipeng Yang, Li Yang 0005, Junbo Jia, Anyuan Sang, Wenjie Sha |
IWQoS | 6 |
| 2025 | STGAN: Detecting Host Threats via Fusion of Spatial-Temporal Features in Host Provenance GraphsabstractAs the complexity and frequency of cyberattacks, such as Advanced Persistent Threats (APTs) and ransomware, continue to escalate, traditional anomaly detection methods have proven inadequate in addressing these sophisticated, multi-faceted threats. Recently, Host Provenance Graphs (HPGs) have played a crucial role in analyzing system-level interactions, detecting anomalous behaviors, and tracing attack chains. However, existing provenance-based detection methods primarily rely on single-dimensional feature analysis, which fails to capture the dynamic and multi-dimensional patterns of modern APT attacks, resulting in insufficient detection performance. To overcome this limitation, we introduce STGAN, a model that integrates spatial-temporal graphs into host provenance graph modeling. STGAN applies temporal and spatial encoding to dynamic provenance graphs to extract temporal, spatial, and semantic features, constructing a comprehensive feature representation. This representation is further fused and enhanced using a multi-head self-attention mechanism, followed by anomaly detection. Through extensive evaluations on three widely-used provenance graph datasets, we demonstrate that our approach consistently outperforms current state-of-the-art techniques in terms of detection performance. Additionally, we contribute to the research community by releasing our datasets and code, facilitating further exploration and validation. Anyuan Sang, Xuezheng Fan, Li Yang 0005, Junbo Jia, Huipeng Yang |
WWW | 1 |
| 2025 | Hyper attack graph: Constructing a hypergraph for cyber threat intelligence analysisabstractCybersecurity experts are actively exploring and implementing automated technologies to extract and present attack information from Cyber Threat Intelligence . However, there are multiple relations among security entities within Cyber Threat Intelligence, a feature that existing technologies often overlook. Additionally, integrating external security knowledge into cyber threat intelligence intuitively during analysis and presentation poses challenges. We propose the Hyper Attack Graph (HAG) framework, the first work to apply hypergraph data structures in the analysis of cyber threat intelligence. Our approach uses a joint extraction model that incorporates a multi-head selection mechanism, effectively addressing the extraction of multiple relations among security entities. We use hypergraph to display tactics and techniques in cyber threat intelligence. Our evaluation of the HAG framework on 685 real-world cyber threat intelligence reports shows an increase in the F1 score for security entity extraction by 11.12% and for relation extraction by 6.71% over existing efforts. Furthermore, HAG’s ability to visually represent external security knowledge on hypergraphs demonstrates its potential as a valuable tool in cybersecurity analysis. Junbo Jia, Li Yang 0005, Anyuan Sang |
Comput. Secur. | 4 |
| 2025 | Resist Dependency Explosion in Attack Investigation With Splittable Tag Propagation and AggregationabstractAdvanced Persistent Threats (APTs) pose significant security risks to the community. Researchers thereby propose techniques to capture the complex and stealthy scenarios of APT attacks through the use of provenance graphs to model system entities and their dependencies. Particularly, to mitigate the dependency explosion problem in attack investigation using provenance graphs, tag-based and priority-based provenance graphs are frequently utilized for analyzing attacks. These methods use threat tag propagation and threat prioritization to reduce the size of the provenance graph for faster analysis. Unfortunately, these methods can allow more complex and potential attacks to evade detection. To overcome these difficulties, we propose an APT attack investigation system,ProTaging, for APT detection and forensic analysis. By using Tactics, Techniques, and Procedures (TTPs) rules to assign and update the node's threat tag, splittable tag propagation to control the scope of threat information, and threat weight aggregation and prioritized backward analysis during the forensic analysis phase, ProTaging effectively reconstructs attack paths in seconds without dependency explosion. Experimental results on both the simulation dataset, DARPA TC E3, E5 dataset, and DARPA OpTC dataset demonstrate that ProTaging generates smaller dependency graphs (2.5 times smaller) and has fewer false positives (6.7 times fewer) compared to state-of-the-art solutions. Additionally, ProTaging significantly reduces manual investigation effort by approximately 99.9%. Anyuan Sang, Junbo Jia, Li Yang 0005, Pengbin Feng, Jianfeng Ma 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | Generating Adversarial Malware Examples Against Multiple Machine Learning DetectorsabstractMalware poses a significant threat to network and information system security, particularly in industrial Internet of Things (IIoT) environments, where embedded systems and edge devices often rely on general-purpose operating systems. Although machine learning (ML) techniques have advanced malware detection, they remain vulnerable to adversarial attacks. Current research primarily focuses on adversarial examples targeting single ML detectors, but the widespread use of ensemble learning necessitates generating adversarial examples that can simultaneously evade multiple detectors. To address this challenge, we propose GanGenetic, a novel framework that combines generative adversarial networks (GANs) with genetic algorithms (GAs) to generate adversarial malware examples targeting import address table features in portable executable files. GanGenetic generates examples with minimal perturbations while simultaneously evaluating the robustness of multiple ML detectors. The framework first generates initial examples using GANs, then optimizes them through a GA to maximize evasion and minimize noise. Experiments on the VirusShare and Ember datasets show that GanGenetic can evade detection by seven ML models (including AdaBoost, Gradient Boosting Decision Trees, logistic regression, multilayer perceptron, random forest, support vector machine, and MalConv) with an average attack success rate exceeding 96%. In addition, in real-world tests, the framework successfully evaded detection while preserving malware functionality. Anyuan Sang, Li Yang 0005, Junbo Jia, Huipeng Yang |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | Assessing Threats: Security Boundary and Side-Channel Attack Detection in the Metaverse
Ruiyuan Yang, Guohao Li 0004, Li Yang 0005, Jiangyu Wang, Anyuan Sang |
ICDF2C (2) | 6 |
| 2024 | Obfuscating Provenance-Based Forensic Investigations with Mapping System Meta-BehaviorabstractThe provenance graph technique has gained popularity for attack analysis, such as Advanced Persistent Threat (APT) attacks, by creating entity interaction graphs from host audit logs. While this method has shown promising analysis results and interpretability, its robustness against mimic attacks carried out by potentially skilled attackers has yet to be fully proven. Recent research has showcased adversarial methodologies targeting provenance-based Machine Learning (ML) detectors, leading to evasion attacks through the addition of corresponding nodes and edges to the feature space. However, these approaches face several challenges, including the difficulty in translating feature alterations into practical attack scenarios and limited applicability to other provenance graph-based detection schemes. Anyuan Sang, Li Yang 0005, Junbo Jia |
RAID | 1 |