Shanshan Wang 0003

dblp:62/3650-3 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 3 first-author · 4 since 2021Systems, architecture and hardware · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Security and privacy · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AdSeeker: An Agent-Driven Framework for Automated Malicious Ad Detection in Mobile Apps
Shanshan Wang 0003, Bo Yang 0001, Zhen Shan
SECON2
2024 Encrypted Malicious Traffic Detection Based on Graph Convolutional Network and Temporal Dissection
abstract
With the rapid development of the Internet, network security issues are receiving increasing attention and various network attack methods are emerging, among which encrypted malicious traffic, as a common attack method, has become a major challenge in the field of network security due to its extremely high concealment and impact. Traditional network security detection methods are difficult to accurately detect and defend against encrypted malicious traffic because they focus more on the detection of traffic content and fail to analyze and detect malicious traffic in terms of the structure and characteristics of the traffic data itself. Therefore, we construct network traffic graphs based on graph convolutional network to classify nodes within the network traffic data. In addition, we introduce temporal dissection to gain a more comprehensive understanding of network micro-dynamics. This study innovatively proposes a method that combines graph convolutional network and temporal dissection for the detection of encrypted malicious traffic. We conduct experiments on two real-world encrypted network traffic datasets, the results show that the accuracy, precision, recall and F1-measure of the method exceed 98% on the USTC-TFC2016 dataset, and detect more than 7900 flows per second, which is nearly 27 times faster than the same detection effect model. The performance on the DataCon2020 dataset is also better than the other six models, confirming that our method can effectively improve the effectiveness and speed of encrypted malicious traffic detection.
Shanshan Wang 0003, Jin Au-Yeung
CSCWD2
2024 KGhish: A Phishing Website Detection Method Based on Knowledge Graph
Changlin Liu, Shanshan Wang 0003, Limei Huang
ICIC (13)2
2023 CSLog: Anomaly Detection for Syslog Based on Contrastive Self-Supervised Representation Learning
Shuhao Yan, Shanshan Wang 0003, Xiaoqing Jiang, Xueyang Cao
APNOMS2
2023 The Secondary Isolated Data Island: Isolated Data Island Caused by Blockchain in Federated Learning
abstract
Compared to other fields, data in the medical field has a higher degree of sensitivity. Therefore, in order to ensure that when using medical data for machine learning, a Federated Learning (FL) framework is usually used. However, the traditional centralized FL framework greatly increases the risk of leaking privacy of clients. For this reason, many researchers have developed and designed a Federated Learning Based on Blockchain (FLchain) to enhance the privacy of this framework. Usually, each hospital will use an independent FLchain. But the researchers overlooked a crucial issue: the application of blockchain in FLchain can create isolated data island. Different FLchain use different blockchains, so the data between the two frameworks cannot be used to train a model at the same time. We call this problem the secondary isolated data island (SIDI). The problem leads to hospitals using FLchain architecture being unable to use each other’s data for model training while ensuring information security. We first discovered and proposed the issue of SIDI, and solved it by using Federated Learning Based on Blockchain with Inter-Blockchain Technology (IBT-FLchain) framework. We designed an Hashed Timelock Contract of global model aggregation (HTLC-M) in FLchain by improving the Hashed Timelock Contract (HTLC). We have demonstrated through experiments that the IBT-FLchain framework can effectively solve SIDI problem and greatly alleviate the storage cost of clients in the blockchain.
Shanshan Wang 0003, Xueyang Cao, Shuhao Yan, Yadi Han, Limei Huang
BIBM2
2023 Overcoming Noisy Labels in Federated Learning Through Local Self-Guiding
abstract
Federated Learning (FL) is a privacy-preserving machine learning paradigm that enables clients such as Internet of Things (IoT) devices, and smartphones, to train a high-performance global model jointly. However, in real-world FL deployments, carefully human-annotated labels are expensive and time-consuming. So the presence of incorrect labels (noisy labels) in the local training data of the clients is inevitable, which will cause the performance degradation of the global model. To tackle this problem, we propose a simple but effective method Local Self-Guiding (LSG) to let clients guide themselves during training in the presence of noisy labels. Specifically, LSG keeps the model from memorizing noisy labels by enhancing the confidence of model predictions. Meanwhile, it utilizes the knowledge from local historical models which haven't fit noisy patterns to extract potential ground truth labels of samples. To keep the knowledge without storing models, LSG records the exponential moving average (EMA) of model output logits at different local training epochs as self-ensemble logits on clients' devices, which will lead to negligible computation and storage overhead. Then logit-based knowledge distillation is conducted to guide the local training. Experiments on MNIST, Fashion-MNIST, CIFAR-10, ImageNet-100 with multiple noise levels, and an unbalanced noisy dataset, Clothing1M, demonstrate the resistance of LSG to noisy labels. The code of LSG is available at https://github.com/DaokuanBai/LSG-Main
Daokuan Bai, Shanshan Wang 0003, Wenyue Wang, Hua Wang 0002
CCGrid2
2023 Decentralized Reinforced Anonymous FLchain: a Secure Federated Learning Architecture for the Medical Industry
abstract
In the age of big data, data has already become a "high-value commodity" with clear price. Privacy leaks can lead to personal information security violations during training in federated learning, especially in the medical industry. The data of the medical industry is characterized by large amount of data and high demand for privacy. For this reason, we designed a Decentralized Reinforced Anonymous Federated Learning Based on Blockchain (DRA-FLchain) with high privacy protection and strong anonymity. In view of the large amount of data in the medical industry, DRA-FLchain uses blockchain technology to enable a large number of clients to participate in model training, and also uses cut through technology to reduce the cost of storage space. DRA-FLchain uses the blockchain and ring signature to ensure the anonymity of the client’s identity, and also uses homomorphic encryption and mask to protect the security of the model. For anonymous Federated Learning (FL), we set up a novel reward mechanism based on game theory Reward mechanism based on ring signature (RMBRS), which can distribute rewards fairly in the anonymous FL architecture. We compared the accuracy and operation efficiency of FL, Federated Learning Based on Blockchain (FLchain) and DRA-FLchain through experiments. The experimental results show that DRA-FLchain is an effective anonymous and secure architecture. Finally, we proved that DRA-FLchain can still protect the privacy of clients well in extreme cases through case study.
Shanshan Wang 0003, Wenyue Wang, Youmian Wang, Lin Wang 0004
COMPSAC2
2023 Measurement of Illegal Android Gambling App Ecosystem From Joint Promotion Perspective
abstract
With the continuous intensification of China’s crackdown on gambling and the rapid development of Internet technology, offline gambling has gradually shifted to online, and operations have been extended from domestic to overseas. Confronted with strict censorship regulations, illegal online gambling websites can employ dark links and black hat search engine optimization(SEO) approaches to promote and distribute mobile gambling apps(applications). To expose China’s gambling application ecosystem, the researchers conducted a cluster analysis of illegal gambling apps. However, they did not consider the interdependence between their distribution channels and upstream promotion websites. For the first time, we analyze and measure the gambling application ecosystem at the promotion community level in conjunction with the promotion relationship between the gambling apps and their upstream gambling pages. The clustering relationships of gambling apps, domains and IPs, and payment platforms in gambling communities with similar promotion preferences are mainly analyzed. From 277,816 related domains of suspected illegal industries, these illegal gambling websites’ upstream and downstream relationships are screened and separated in detail, forming a knowledge graph of online gambling apps centered on promotion relationships. Based on this graph, information integration of entities with different functions in the gambling industry chain (promotion, gambling pages, mobile apps) is carried out. This research provides a new perspective on the profit chain research of illegal online apps in conjunction with the promotion relationship of gambling pages. We have also published this dataset to support further in-depth analysis and black industry governance work in the security community at https://github.com/yadispace/gambling_app_KG.
Yadi Han, Shanshan Wang 0003, Xueyang Cao, Limei Huang
DSAA2
2023 Adversarial Attack with Genetic Algorithm against IoT Malware Detectors
abstract
The exponential growth and sophistication of Internet of Things (IoT) malware behavior have resulted in new detection technologies capable of defending IoT devices against some threats. However, their success has stimulated the interest of attackers attempting to circumvent current IoT malware detectors. Among detection technologies, the detectors trained based on Uniform Resource Locator (URL) requests have become popular. To draw attention to the safety of the detectors, we propose a grey-box method to attack detectors based on URL requests without breaking malicious functions of URL requests. The key idea is to add perturbations to the tail of URLs. Specifically, this method is based on a Genetic Algorithm (GA) to find suitable perturbations and optimizes the process of adversarial attacks through a dynamic number of evolution directions and a maximum generation limit. The effectiveness of our adversarial attack is demonstrated by experimental results based on a widely used public dataset CSIC2010 and several representative detectors. As far as we know, this is the first time an adversarial attack against IoT detectors based on URL requests has been done. The method has an attack success rate of more than 92 %. Furthermore, experiment results show that the method can reduce query numbers while maintaining the attack success rate.
Shanshan Wang 0003, Wenyue Wang, Daokuan Bai, Lizhi Peng
ICC2
2023 Devils in the Clouds: An Evolutionary Study of Telnet Bot Loaders
abstract
One of the innovations brought by Mirai and its derived malware is the adoption of self-contained loaders for infecting IoT devices and recruiting them in botnets. Functionally decoupled from other botnet components and not embedded in the payload, loaders cannot be analysed using conventional approaches that rely on honeypots for capturing samples. Different approaches are necessary for studying the loaders evolution and defining a genealogy. To address the insufficient knowledge about loaders' lineage in existing studies, in this paper, we propose a semantic-aware method to measure, categorize, and compare different loader servers, with the goal of highlighting their evolution, independent from the payload evolution. Leveraging behavior-based metrics, we cluster the discovered loaders and define eight families to determine the genealogy and draw a homology map. Our study shows that the source code of Mirai is evolving and spawning new botnets with new capabilities, both on the client side and the server side. In turn, shedding light on the infection loaders can help the cybersecurity community to improve detection and prevention tools.
Yuhui Zhu, Qiben Yan 0001, Shanshan Wang 0003, Alberto Giaretta 0001, Enlong Li, Lizhi Peng, Mauro Conti
ICC4
2023 Exploring Wi-Fi Privacy Disclosure: A Novel Approach to User Identity Prediction Based on Traffic Multi-level Information
abstract
Nowadays, people are accustomed to network communication via Wireless Fidelity (Wi-Fi) when using mobile phones and IoT devices in fixed places (such as homes and dormitories). This means that a large amount of Wi-Fi traffic data will be generated every day. This paper focuses on the leakage of user identity privacy in traffic, namely: can user identity attributes be inferred from real home Wi-Fi traffic data? If malicious attackers can obtain user identities so "conveniently", it means that the success rate of malicious activities such as phishing and fraud can be easily increased. Therefore, to explore this question, we propose a new method for predicting user attributes based on user network behavior similarity. Then we collected real world data for verification and compared it with other methods. Finally, this paper also analyzes the current response methods for Wi-Fi privacy disclosure and points out the problems and future research directions.
Shanshan Wang 0003, Xueyang Cao, Yadi Han
ICPADS2
2023 Low-Frequency Aware Unsupervised Detection of Dark Jargon Phrases on Social Platforms
Limei Huang, Shanshan Wang 0003, Changlin Liu, Xueyang Cao, Yadi Han, Shaolei Liu
PRICAI (2)2
2023 Safety or Not? A Comparative Study for Deep Learning Apps on Smartphones
abstract
Recent years have witnessed an astonishing explosion in the evolution of mobile applications powered by deep learning (DL) technologies. Considering that inference of DL models in the cloud requires transferring user data to server, which is prone to the risk of user privacy leakage, many developers choose to deploy the models on local devices for executing the inference process. However, this also raises a number of other security issues, such as adversarial attacks, model stealing attacks, etc. To explore the security issues that exist in deep learning applications (DL apps), we conducted the first comprehensive comparative study of the top 200 apps in each category on Google Play. We built DLApplnspector, a vulnerability detection tool that combines dynamic and static analysis methods for dissecting apps, which helped us automate our empirical study on real-world mobile DL apps. First, we identify DL apps and extract their models by using DL Checker. Subsequently, Static Scoper is provided to detect encryption and reusability of DL models. Finally, within Dynamic Scoper, we use reverse engineering techniques on the network traffic to parse out the packets and collect side-channel information during application runtime. Our research shows that the majority of developers prefer to use open-source models, with almost 92% of models successfully parsed. This suggests that most models are unprotected. DL apps are more likely to upload user behaviour and collect private data than Non-DL apps. We provide security recommendations for developers and users to address the issues discovered.
Jin Au-Yeung, Shanshan Wang 0003
TrustCom2
2020 Deep and broad URL feature mining for android malware detection
Shanshan Wang 0003, Qiben Yan 0001, Ke Ji, Lizhi Peng, Bo Yang 0001, Mauro Conti
Inf. Sci.1
2019 A mobile malware detection method using behavior features in network traffic
Shanshan Wang 0003, Qiben Yan 0001, Bo Yang 0001, Lizhi Peng, Zhongtian Jia
J. Netw. Comput. Appl.1
2018 MulAV: Multilevel and Explainable Detection of Android Malware with Data Fusion
Qiben Yan 0001, Shanshan Wang 0003, Kun Ma 0001, Yuliang Shi, Li-Zhen Cui 0001
ICA3PP (4)4
2018 A Fast and Effective Detection of Mobile Malware Behavior Using Network Traffic
Shanshan Wang 0003, Lizhi Peng, Yuliang Shi
ICA3PP (4)3
2018 Deep and Broad Learning Based Detection of Android Malware via Network Traffic
abstract
In recent years, the scale and diversity of malicious software on mobile networks are constantly increasing, thereby causing considerable danger to users' property and personal privacy. In this study, we devise a method that uses the URLs visited by applications to identify malicious apps. A multi-view neural network is used to create a malware detection model that emphasizes depth and width. This neural network can create multiple views of the input automatically and distribute soft attention weights to focus on different features of input. Multiple views preserve rich semantic information from input for classification without requiring complicated feature engineering. In addition, we conduct comprehensive experiments to compare the proposed method with others and verify the validity of the detection model. The experimental results show that our method has a certain timeliness. It can not only effectively detect malware discovered in different months of a certain year, but also detect potentially malicious apps in the third-party app market. We also compare the detection results of the proposed method on wild apps with 10 popular anti-virus scanners, and the final result shows that our approach ranks second in terms of detection performance.
Shanshan Wang 0003, Qiben Yan 0001, Ke Ji, Lin Wang 0004, Bo Yang 0001, Mauro Conti
IWQoS1
2018 Lexical Mining of Malicious URLs for Classifying Android Malware
Shanshan Wang 0003, Qiben Yan 0001, Lin Wang 0004, Riccardo Spolaor, Bo Yang 0001, Mauro Conti
SecureComm (1)1
2018 Machine learning based mobile malware detection using highly imbalanced network traffic
abstract
In recent years, the number and variety of malicious mobile apps have increased drastically, especially on Android platform, which brings insurmountable challenges for malicious app detection. Researchers endeavor to discover the traces of malicious apps using network traffic analysis. In this study, we combine network traffic analysis with machine learning methods to identify malicious network behavior, and eventually to detect malicious apps. However, most network traffic generated by malicious apps is benign, while only a small portion of traffic is malicious, leading to an imbalanced data problem when the traffic model skews towards modeling the benign traffic. To address this problem, we introduce imbalanced classification methods, including the synthetic minority oversampling technique (SMOTE) + support vector machine (SVM), SVM cost-sensitive (SVMCS), and C4.5 cost-sensitive (C4.5CS) methods. However, when the imbalance rate reaches a certain threshold, the performance of common imbalanced classification algorithms degrades significantly. To avoid performance degradation, we propose to use the imbalanced data gravitation-based classification (IDGC) algorithm to classify imbalanced data. Moreover, we develop a simplex imbalanced data gravitation classification (S-IDGC) model to further reduce the time costs of IDGC without sacrificing the classification performance. In addition, we propose a machine learning based comparative benchmark prototype system, which provides users with substantial autonomy, such as multiple choices of the desired classifiers or traffic features. Using this prototype system, users can compare the detection performance of different classification algorithms on the same data set, as well as the performance of a specific classification algorithm on multiple data sets.
Qiben Yan 0001, Hongbo Han, Shanshan Wang 0003, Lizhi Peng, Lin Wang 0004, Bo Yang 0001
Inf. Sci.4
2018 Detecting Android Malware Leveraging Text Semantics of Network Flows
abstract
The emergence of malicious apps poses a serious threat to the Android platform. Most types of mobile malware rely on network interface to coordinate operations, steal users' private information, and launch attack activities. In this paper, we propose an effective and automatic malware detection method using the text semantics of network traffic. In particular, we consider each HTTP flow generated by mobile apps as a text document, which can be processed by natural language processing to extract text-level features. Then, we use the text semantic features of network traffic to develop an effective malware detection model. In an evaluation using 31 706 benign flows and 5258 malicious flows, our method outperforms the existing approaches, and gets an accuracy of 99.15%. We also conduct experiments to verify that the method is effective in detecting newly discovered malware, and requires only a few samples to achieve a good detection result. When the detection model is applied to the real environment to detect unknown applications in the wild, the experimental results show that our method performs significantly better than other popular anti-virus scanners with a detection rate of 54.81%. Our method also reveals certain malware types that can avoid the detection of anti-virus scanners. In addition, we design a detection system on encrypted traffic for bring-your-own-device enterprise network, home network, and 3G/4G mobile network. The detection model is integrated into the system to discover suspicious network behaviors.
Shanshan Wang 0003, Qiben Yan 0001, Bo Yang 0001, Mauro Conti
IEEE Trans. Inf. Forensics Secur.1
2017 Android Malware Clustering Analysis on Network-Level Behavior
Shanshan Wang 0003, Lin Wang 0004, Ke Ji
ICIC (1)1
2016 TrafficAV: An effective and explainable detection of mobile malware behavior using network traffic
abstract
Android has become the most popular mobile platform due to its openness and flexibility. Meanwhile, it has also become the main target of massive mobile malware. This phenomenon drives a pressing need for malware detection. In this paper, we propose TrafficAV, which is an effective and explainable detection of mobile malware behavior using network traffic. Network traffic generated by mobile app is mirrored from the wireless access point to the server for data analysis. All data analysis and malware detection are performed on the server side, which consumes minimum resources on mobile devices without affecting the user experience. Due to the difficulty in identifying disparate malicious behaviors of malware from the network traffic, TrafficAV performs a multi-level network traffic analysis, gathering as many features of network traffic as necessary. The proposed method combines network traffic analysis with machine learning algorithm (C4.5 decision tree) that is capable of identifying Android malware with high accuracy. In an evaluation with 8,312 benign apps and 5,560 malware samples, TCP flow detection model and HTTP detection model all perform well and achieve detection rates of 98.16% and 99.65%, respectively. In addition, for the benefit of user, TrafficAV not only displays the final detection results, but also analyzes the behind-the-curtain reason of malicious results. This allows users to further investigate each feature's contribution in the final result, and to grasp the insights behind the final decision.
Shanshan Wang 0003, Lei Zhang 0085, Qiben Yan 0001, Bo Yang 0001, Lizhi Peng, Zhongtian Jia
IWQoS1