Songyun Wu

dblp:257/7105 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0003-4382-6550ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 3 · 2 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Breaking the Seed Barrier: Discovering Active IPv6 Addresses in Seedless Scenarios
Wenjian Zhang, Guanglei Song, Binkai Ma, Lin He 0004, Songyun Wu, Jiahai Yang 0001
INFOCOM5
2026 Alert2Vec: Eliminating Alert Fatigue by Embedding Security Alerts Through Subgraph Learning
Songyun Wu, Xiaoqing Sun, Enhuan Dong, Jiahai Yang 0001
IEEE Trans. Dependable Secur. Comput.1
2025 ALM: A Two-Stage Traffic Anomaly Detection and Analysis System via the Large Language Model
abstract
In recent years, deep learning-based traffic anomaly detection has proven very promising. Although current methods achieve high accuracy in detecting anomalies, they struggle to accurately classify attack types of anomalies due to the imbalanced distribution of attack samples. To address the issue, we propose a highly intelligent system, ALM, which can simultaneously provide accurate traffic anomaly detection and attack types classification with the aid of the Large Language Model (LLM)'s few-shot learning ability. To tackle the challenges of training cost and inference efficiency associated with large models, ALM adopts a two-stage solution, i.e., AnomalyDetector and Anomaly Analyzer, that combines the fine-tuned LLM with small models. In the first stage, AnomalyDetector ensembles a set of lightweight models to handle high-concurrency real-time network traffic anomaly detection. In the second stage, Anomaly Analyzer leverages the LLM's powerful fitting and few-shot learning abilities for traffic anomaly analysis through three processes: LLM task adaption, traffic to sequence, and LLM fine-tuning. This allows Anomaly Analyzer to accurately identify the attack types and potential false positives. Experimental results indicate that ALM achieves over 90% Micro-F1 on four public datasets, with a maximum of 99.94 %, surpassing the baseline. Additionally, it requires minimal training costs while significantly improving inference efficiency compared to the pure LLM mode.
Songyun Wu, Enhuan Dong, Haina Hu, Jiahai Yang 0001
NOMS1
2025 IPdb: A High-Precision IP Level Industry Categorization of Web Services
abstract
IP addresses with web services are crucial in the Internet ecosystem. Classifying these addresses by industry and organization offers valuable insights into the entities utilizing them, enabling more efficient network management and enhanced security. Previous work in website classification and Internet management struggles to offer an IP-level perspective of the industries of web services due to their limited industry categories or potential industry inconsistencies between IP address owners and AS owners. To this end, we present IPdb, an IP-level industry categorization dataset. To construct the dataset, we developed LLMIC, a Large Language Model-based Industry Categorization framework with a precision of nearly 96%. IPdb serves as a labeled database for future endeavors in developing IP-level industry classifiers, encompassing over 200 million IP addresses. Furthermore, our study indicates that 30% ~ 50% of organizations within critical infrastructure industries deploy web servers across multiple ASes. Our study also validates the problem of mismatched granularity in industry categorization at the AS level with 87.83% ASes in IPv4 and 72.96% ASes in IPv6 containing IP addresses from different industries.
Guanglei Song, Jiahai Yang 0001, Songyun Wu, Jinlei Lin, Lin He 0004, Chenglong Li 0006
WWW5
2024 6Vision: Image-Encoding-Based IPv6 Target Generation in Few-Seed Scenarios
abstract
Efficient global Internet scanning is crucial for network measurement and security analysis. While existing target generation algorithms verify remarkable performance in largescale detection, their efficiency notably diminishes in few-seed scenarios. This decline is primarily attributed to the intricate configuration rules and sampling bias of seed addresses. Moreover, instances where BGP prefixes have few seed addresses are widespread, constituting$63.65 \%$of occurrences. We introduce 6 Vision to tackle this challenge by introducing a novel approach to encoding IPv6 addresses into images, facilitating comprehensive analysis of intricate configuration rules. Through feature stitching, 6 Vision not only improves the learnable features but also amalgamates addresses associated with configuration patterns for enhanced learning. Moreover, it integrates an environmental feedback mechanism to refine model parameters based on identified active addresses, thereby alleviating the sampling bias inherent in seed addresses. As a result, 6Vision achieves high-accuracy detection even in few-seed scenarios. The HitRate of 6 Vision is improved by$181 \% \sim 2,490 \%$compared to existing algorithms, while the CoverNum is$1.18 \sim 11.20$times that of them. Additionally, 6Vision can function as a preliminary detection module for existing algorithms, yielding a conversion gain (CG) ranging from$242 \% \sim 2,081 \%$. Ultimately, we achieve a conversion rate (CR) of$28.97 \%$for few-seed scenarios. We enrich the IPv6 hitlist, not only enhancing current target generation algorithms for large-scale address detection in few-seed scenarios but also effectively supporting IPv6 network measurement and security analysis.
Wenjian Zhang, Guanglei Song, Lin He 0004, Jinlei Lin, Songyun Wu, Chenglong Li 0006, Jiahai Yang 0001
ICNP5
2022 Joint prediction on security event and time interval through deep learning
Songyun Wu, Bo Wang 0066, Shuhan Fan, Jiahai Yang 0001, Jia Li 0033
Comput. Secur.1
2019 ALEAP: Attention-based LSTM with Event Embedding for Attack Projection
abstract
Cyberattacks have developed rapidly in diversity and complexity in recent years. Despite the existence of various defense systems, it cannot provide early warnings and prevent catastrophic consequences in advance. Therefore, the need for prediction becomes more and more urgent, especially for those multiple step attacks in which several steps are required for achieving the attack successfully. In this paper, we focus on attack projection that is aimed to predict the next step of the attack based on historical information and gained knowledge of similar events happened in the past. Previous models on attack projection based on probability graph model or simple RNN models, which may limit their capability of noise tolerance and sequence association analysis. To remedy this, we propose a method called ALEAP which incorporates event embedding and attention mechanism into LSTM models to better predict the future events. We test ALEAP on a dataset of millions of security events collected from the multi-source security devices, and show that our approach is effective in event prediction. ALEAP also provides a useful method for security specialists and all computer environment-related parties to better predict attack projection and defend known attacks.
Shuhan Fan, Songyun Wu, Zimu Li, Jiahai Yang 0001
IPCCC2