EDBT 2026 Demo / reviewers in the wild / expert
Yafei Sang
dblp:166/4214
· DBLP profile ↗
17ranked-venue papers
5as first author
11since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 5 since 2021Security and privacy · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 2 first-authorComputer networks · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Let model keep evolving: Incremental learning for encrypted traffic classification
Xiang Li 0135, Jiang Xie 0004, Qige Song, Yafei Sang, Yongzheng Zhang 0002, Tianning Zang |
Comput. Secur. | 4 |
| 2024 | SepBIN: Binary Feature Separation for Better Semantic Comparison and Authorship VerificationabstractBinary semantic comparison and authorship verification are critical in many security applications. They respectively focus on the functional semantic features and developers’ programming style features of binary code, which are usually mixed without clear demarcation. Recently, researchers have proposed learning-based approaches for intelligent binary analysis. They generally addressed single tasks with hand-crafted feature sets or neural binary encoders, which suffer performance bottlenecks due to the noise in mixed features. This paper proposesSepBIN, a novel neural network framework that exploits the intrinsic correlation of binary semantic comparison and authorship verification tasks and automatically separates semantic and stylistic binary features. We first construct a strong backbone binary encoder, then utilize preliminary decomposition subnets and the flexible gating-based feature fusion mechanism to distill pure semantic-related and style-related binary representations, and further improve their quality by a feature reconstruction module. The overallSepBINmodel is optimized by a multi-objective joint optimization strategy. We conduct extensive experiments on Google Code Jam (GCJ) datasets in different languages and scales. Results show thatSepBINsimultaneously benefits binary semantic comparison and authorship verification tasks through the effective binary semantic-style feature separation mechanism, and provides multi-perspectives interpretability for the performance gains. For state-of-the-art approaches with different binary encoders,SepBINcan adaptively improve them with the designed separation modules. Furthermore, we adopt a pretraining-finetuning strategy to effectively transferSepBIN’s separation capability in real-world applications, including APT malware homology detection and binary semantic comparison against code obfuscations. Qige Song, Yafei Sang, Yongzheng Zhang 0002 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | TCCN: A Network Traffic Classification and Detection Model Based on Capsule NetworkabstractPrivacy information theft traffic is usually detected using traffic classification methods, and deep learning-based detection methods are effective for this task. However, these methods have complex preprocessing processes as well as tend to ignore the deep features of network traffic, while the generalization ability is not outstanding. In this paper, a network traffic classification and detection model TCCN (Traffic Classification Capsule Network) based on capsule network is proposed. Meanwhile, PacketCGAN-based data balancing method is introduced to assist TCCN in traffic classification. A new feature graph vectorization method is used to improve the efficiency of TCCN. In addition, TCCN uses dynamic routing mechanism that can retain more valid traffic characteristics. In parallel, this paper proposes a new loss function to improve the generalization performance of TCCN. The results show that TCCN can show good performance in different experimental scenarios. After effective training, TCCN can also show high detection accuracy in new datasets, and the generalization ability of the model reaches a relatively excellent level. Yafei Sang, Zhenyu Cheng 0001, Tianning Zang |
ICC | 2 |
| 2023 | Listen to Minority: Encrypted Traffic Classification for Class Imbalance with Contrastive Pre-TrainingabstractMobile Internet has profoundly reshaped modern lifestyles in various aspects. Encrypted Traffic Classification (ETC) naturally plays a crucial role in managing mobile Internet, especially with the explosive growth of mobile apps using encrypted communication. Despite some existing learning-based ETC methods showing promising results, three-fold limitations still remain in real-world network environments, i) label bias caused by traffic class imbalance, ii) traffic homogeneity caused by component sharing, and iii) training with reliance on sufficient labeled traffic. None of the existing ETC methods can address all these limitations. In this paper, we propose a novel Pre-trAining Semi-Supervised ETC framework, dubbed PASS. Our key insight is to resample the original train dataset and perform contrastive pre-training without using individual app labels directly to avoid label bias issues caused by class imbalance, while obtaining a robust feature representation to differentiate overlapping homogeneous traffic by pulling positive traffic pairs closer and pushing negative pairs away. Meanwhile, PASS designs a semi-supervised optimization strategy based on pseudo-label iteration and dynamic loss weighting algorithms in order to effectively utilize massive unlabeled traffic data and alleviate manual train dataset annotation workload. PASS outperforms state-of-the-art ETC methods and generic sampling approaches on four public datasets with significant class imbalance and traffic homogeneity, remarkably pushing the F1 of Cross-Platform215 with 1.31%$\uparrow$, ISCX-17 with 9.12%$\uparrow$. Furthermore, we validate the generality of the contrastive pre-training and pseudo-label iteration components of PASS, which can adaptively benefit ETC methods with diverse feature extractors. Xiang Li 0135, Juncheng Guo, Qige Song, Jiang Xie 0004, Yafei Sang, Yongzheng Zhang 0002 |
SECON | 5 |
| 2023 | Toward IoT device fingerprinting from proprietary protocol traffic via key-blocks aware approach
Yafei Sang, Jisong Yang, Yongzheng Zhang 0002 |
Comput. Secur. | 1 |
| 2022 | An Adaptive Ensembled Neural Network-Based Approach to IoT Device Identification
Jingrun Ma, Yafei Sang, Yongzheng Zhang 0002, Beibei Feng, Yuwei Zeng |
CollaborateCom (2) | 2 |
| 2022 | ACS: An Efficient Messaging System with Strong Tracking-Resistance
Zhefeng Nan, Changbo Tian, Yafei Sang, Guangze Zhao |
CollaborateCom (2) | 3 |
| 2022 | Evading Encrypted Traffic Classifiers by Transferable Adversarial Traffic
Hanwu Sun, Chengwei Peng, Yafei Sang, Yongzheng Zhang 0002, Yujia Zhu |
CollaborateCom (2) | 3 |
| 2022 | A Longitudinal Measurement and Analysis of Pink, a Hybrid P2P IoT Botnet
Binglai Wang, Yafei Sang, Yongzheng Zhang 0002, Ruihai Ge |
CollaborateCom (2) | 2 |
| 2022 | A longitudinal Measurement and Analysis Study of Mozi, an Evolving P2P IoT BotnetabstractDue to the remarkable diversity and ubiquity of IoT devices, many IoT botnets like Mirai and Hajime are inclined to infect vulnerable embedded devices to achieve the purpose of scale expansion and future attacks. Nowadays, a novel, evolving P2P botnet in the wild, named Mozi, targets numerous devices but substantially differs in many aspects involving self-proliferation, network control, and distribution situations. To uncover the curtain, throughout detailed measurement–a novel active probing approach (namely Bawa) for tracking Mozi’s dynamic topology infrastructure and passive collection of Mozi binaries and its download URLs, we provide a holistic view of Mozi, including its scale and geographic distribution. The dataset1collected in a passive mode is available to assist in tracking and mitigating the Mozi botnet’s expansion. Binglai Wang, Yafei Sang, Yongzheng Zhang 0002 |
TrustCom | 2 |
| 2021 | Exploiting Heterogeneous Information for IoT Device Identification Using Graph Convolutional Network
Jisong Yang, Yafei Sang, Yongzheng Zhang 0002, Chengwei Peng |
CollaborateCom (1) | 2 |
| 2020 | IncreAIBMF: Incremental Learning for Encrypted Mobile Application Identification
Yafei Sang, Mao Tian, Yongzheng Zhang 0002 |
ICA3PP (3) | 1 |
| 2020 | IoTCMal: Towards A Hybrid IoT Honeypot for Capturing and Analyzing MalwareabstractNowadays, the emerging Internet-of-Things (IoT) emphasize the need for the security of network-connected devices. Additionally, there are two types of services in IoT devices that are easily exploited by attackers, weak authentication services (e.g., SSH/Telnet) and exploited services using command injection. Based on this observation, we propose IoTCMal, a hybrid IoT honeypot framework for capturing more comprehensive malicious samples aiming at IoT devices. The key novelty of IoTC-MAL is three-fold: (i) it provides a high-interactive component with common vulnerable service in real IoT device by utilizing traffic forwarding technique; (ii) it also contains a low-interactive component with Telnet/SSH service by running in virtual environment. (iii) Distinct from traditional low-interactive IoT honeypots[1], which only analyze family categories of malicious samples, IoTCMal primarily focuses on homology analysis of malicious samples. We deployed IoTCMal on 36 VPS1instances distributed in 13 cities of 6 countries. By analyzing the malware binaries captured from IoTCMal, we discover 8 malware families controlled by at least 11 groups of attackers, which mainly launched DDoS attacks and digital currency mining. Among them, about 60% of the captured malicious samples ran in ARM or MIPs architectures, which are widely used in IoT devices. Binglai Wang, Yu Dou, Yafei Sang, Yongzheng Zhang 0002 |
ICC | 3 |
| 2017 | ProNet: Toward Payload-Driven Protocol Fingerprinting via Convolutions and Embeddings
Yafei Sang, Yongzheng Zhang 0002, Chengwei Peng |
CollaborateCom | 1 |
| 2017 | Fingerprinting Protocol at Bit-Level Granularity: A Graph-Based Approach Using Cell EmbeddingabstractTraffic identification is defined as the act of ascertaining which application or service or protocol is contributing to the network traffic by using a fingerprint, which is a distinguishable unique pattern representing a particular applications traffic. The continual appearance of new applications and their frequent updates emphasize the need for automatic protocol fingerprints generation. In this paper, we propose BitGrapher, a novel graph-based approach that accurately infers protocol fingerprints at bit-level granularity for accurate traffic identification. The proposal is designed to accommodate to various protocol traces including text-based and binary-based protocols, and even possible unknown proprietary communication protocols. Our proposed approach introduces a new concept of cell that allows BitGrapher to encode protocol payloads as a graphical model, and then converts fingerprinting protocol problem as a series of graph operations (e.g., graph construction, pruning, partition). The key insight of the graphical model use cell embeddings that captures the distinguishable positions with their values and distinguishable correlation among them. We implement and evaluate BitGrapher on real-world traces, including DNS, QQLive, SopCast, SMB, HTTP, and SMTP, and our experimental results show that BitGrapher can accurately identify the protocol trace with an average precision of about 97.28% and an average recall of about 99.12%. We also compare the results of BitGrapher to two state-of-the-art approaches ProWord and ProDigger, which shows that BitGrapher provides significant improvements in precision and recall for protocol identification task. Yafei Sang, Yongzheng Zhang 0002 |
ICPADS | 1 |
| 2017 | Detecting Information Theft Based on Mobile Network Flows for Android UsersabstractWith the widespread use of smartphones, more and more malicious attacks happen with information leakage from apps installed on users' devices. The adversary always uses a malware as the client to take remote control of smartphones, and leverages the vulnerability of operation systems to send back the collected information without users' permissions. All the information has to be transferred by network traffic. In this paper, we consider that different apps maybe generate different network flows by different operations, and the "shapes" of the benign flows and malicious ones will be diverse. Thus we propose a detection model based on the analysis of relationships between behavior patterns and network flows, which achieves our goal by using the Random Forest machine learning algorithm to classify the network flows into benign or malicious. To further improve the controllability of the experiment, we design an app called Moledroid to simulate malwares by uploading the user's privacy without authorization, in addition, we can change the behavior pattern of the app to complete our evaluation. Finally, we run this app and several benign apps to generate traffic to detect the malicious network flows, and it shows that our detection model can achieve precision and accuracy higher than 95%, which demonstrates that our model is suitable for detecting the network flows of information theft. Zhenyu Cheng 0001, Xunxun Chen, Yongzheng Zhang 0002, Yafei Sang |
NAS | 5 |
| 2014 | A Segmentation Pattern Based Approach to Automated Protocol IdentificationabstractIn-depth understanding of network traffic is important for a variety of applications, such as network management and network security. In this paper, we propose a novel protocol identification system PSKS, which relies on the statistical signatures of network packet payloads. The proposed approach is based on the key insight that message segmentation patterns can be leveraged for accurate application identification. Specifically, the segmentation possibility for every position of protocol messages exhibits highly skewed frequency distribution due to the reason that different protocols have different message formats (i.e., Distinct message segmentation patterns). Motivated by this observation, we want to extract statistical application fingerprints by exploiting the message segmentation patterns. In PSKS, we first extract the message segmentation patterns by scoring the segmentation possibility scale for each position of messages, and then extract statistical signatures by Kolmogorov-Smirnov test and feed the signatures to tri-training, a collaborative learning algorithm. The tri-training can improve the generalization ability of our final classifier. We implemented and evaluated PSKS, and the experimental results show that PSKS achieves an average precision and recall of approximately 98%. Yafei Sang, Yongzheng Zhang 0002, Yipeng Wang 0001, Yu Zhou 0015 |
PDCAT | 1 |