Zhao Li 0010

dblp:181/2856-10 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-3069-1215ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 1 first-author · 4 since 2021Security and privacy · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PriAgent: A Collaborative Multi-Agent Framework for Auditing Android Privacy Compliance
abstract
Stringent regulations like General Data Protection Regulation (GDPR) mandate that an application's code-level data handling must align with its natural-language privacy policy, creating a critical auditing challenge. However, existing methods, predominantly reliant on static analysis, suffer from a critical limitation: in their pursuit of soundness via over-approximation, they exhibit "semantic blindness"—detecting what data flows exist but not why. This leads to an overwhelming volume of false positives, rendering automated auditing impractical. To bridge this gap, we introduce PriAgent, a novel framework that approaches compliance auditing as a multi-stage, AI-driven reasoning task. Instead of a monolithic model, PriAgent deploys a team of specialized agents that execute a divide-and-conquer strategy. They systematically prune the analysis space by abstracting data flows, pinpoint semantic loci critical for inspection, and perform on-demand summarization of large code blocks to ensure scalability. PriAgent leverages Retrieval-Augmented Generation (RAG) with a curated knowledge base of Android APIs, equipping agents to discern potentially non-compliant behavior from benign functionality. By correlating code-level evidence with the app's stated privacy policy, PriAgent delivers a holistic and explainable verdict for each potential violation. Our evaluations demonstrate that PriAgent significantly reduces false positives, enabling a more scalable and precise compliance audit.
Zhao Li 0010, Zhuojun Jiang, Jiangyi Yin, Jiangchao Chen, Qingyun Liu 0001
AAAI2
2026 Blazer: Encrypted Video Traffic Identification for Mixed Segment Transmission Pattern based on LLM
abstract
Determining the source of encrypted video traffic is an important task in network regulation. In the context of Dynamic Adaptive Streaming over HTTP (DASH), the newly emerged mixed segment transmission pattern introduces substantial difficulties for fingerprint matching, especially under adverse network conditions. To address these challenges, we propose Blazer, a DASH encrypted video traffic identification method for the mixed segment transmission pattern. First, we design a novel fingerprint that integrates video and audio segment sequences. Then, we extract the traffic fingerprint from the TLS record layer of video traffic. Finally, by observing implicit segment-mixing constraints, we design a targeted prompt and Retrieval Augmented Generation (RAG) that enables Large Language Models (LLMs) to perform fingerprint matching effectively. Across 12 network scenarios, Blazer delivers substantially better performance than the other 4 SOTA methods.
Weitao Tang, Meijie Du, Die Hu 0004, Zhao Li 0010, Rong Yang 0008, Qingyun Liu 0001
ICMR5
2025 Leveraging Cross-Layer Network Probing to Detect Stealth Services
abstract
Stealth services have become increasingly popular due to growing demand for privacy protection. To avoid detection by active probing, many have adopted probe-resistant strategies. In this work, we design a suite of carefully crafted probes to expose their hidden vulnerabilities. Through detailed analysis of the corresponding responses, we find that despite the defensive measures implemented, certain implicit information, such as protocol stack fingerprints, can still serve as strong indicators. We present SSChecker, a detection system that combines cross-layer probing with a classification model inspired by Information Bottleneck Theory to address the data sparsity issue inherent in active probing. Our experiments on real-world datasets show that SSChecker outperforms existing methods, including leading industrial detection engines. Our findings demonstrate that current probe-resistant strategies of stealth services remain insufficient. To strengthen privacy protection, we further propose mitigation strategies that help stealth services enhance their ability.
Jiangyi Yin, Chenxu Wang 0006, Zhao Li 0010, Zhuojun Jiang, Jiangchao Chen, Dongfang Hao, Qingyun Liu 0001
TrustCom3
2024 Seeing the Attack Paths: Improved Flow Correlation Scheme in Stepping-Stone Intrusion
abstract
Stepping-stones are widely used by attackers to conceal their identities and gain unauthorized access to restricted targets. Numerous strategies have been suggested to identify stepping-stones and counteract evasive behaviors, with flow correlation standing out as the key technique used in many deanonymization methods. Existing attempts to tackle the flow correlation issue rely on long-flow observations and feature-based methods. Yet, these approaches meet substantial limitations in their suitability across diverse scenarios, network noises influence and accuracy, particularly in the stepping-stone environment. In this paper, we introduce an improved flow correlation model, FlowLinker, which aims to resolve these issues. Specifically, we combined multidimensional statistical features and cumulative flow sequences to construct a robust traffic feature. Then, we leveraged the triplet network to produce an optimized representation, amplifying the difference between unrelated representations. Consequently, it significantly reduces the false correlation rate with flows that seem similar but are unrelated. The experiments on real-world datasets from different network environments we collected show that FlowLinker outperforms other state-of-the-art methods with shorter lengths of flow observations.
Chao Zheng 0001, Zhao Li 0010, Jinqiao Shi
CSCWD3
2024 SCENE: Shape-based Clustering for Enhanced Noise-resilient Encrypted Traffic Classification
abstract
Network traffic classification is critical in network management, quality of service optimization, and security monitoring. However, most existing methods for encrypted traffic classification rely heavily on supervised learning, requiring large amounts of labeled data, and struggle to perform effectively in complex and dynamic network environments. To address these limitations, we propose a novel unsupervised method for encrypted traffic classification, which analyzes byte rate variations to capture traffic behavior patterns. Our approach does not require prior knowledge or large volumes of labeled data, enabling adaptive processing of encrypted traffic in complex network conditions. Specifically, we introduce a noise-resilient shape-line extraction method that preserves core behavioral characteristics of traffic; we design a multidimensional feature extraction strategy that analyzes both uplink and downlink features; and we propose an unsupervised classification algorithm that combines shape-based density clustering with a feature assignment strategy. This algorithm overcomes the limitations of traditional methods, such as the need for predefined cluster numbers, and can classify unknown traffic patterns. We validate our method on five real-world traffic datasets with differing levels of openness, demonstrating its remarkable robustness and accuracy in encrypted traffic classification tasks, thereby greatly enhancing the precision and stability of service classification.
Meijie Du, Mingqi Hu, Zhao Li 0010, Qingyun Liu 0001
TrustCom4
2024 Identifying VPN Servers through Graph-Represented Behaviors
abstract
Identifying VPN servers is a crucial task in various situations, such as geo-fraud detection, bot traffic analysis and network attack identification. Although numerous studies that focus on network traffic detection have achieved excellent performance in closed-world scenarios, particularly those methods based on deep learning, they may exhibit significant performance degradation due to changes in network environment. To mitigate this issue, a few studies have attempted to use methods based on active probing to detect VPN servers. However, these methods still have two limitations. They cannot handle situations without probing responses and are limited in applicability due to their focus on specific VPNs. In this work, we propose VPNChecker, which utilizes the graph-represented behaviors to detect VPN servers in real-world scenarios. VPNChecker outperforms existing methods in four offline datasets. The results from our datasets, containing multiple different VPNs, indicate that VPNChecker has better applicability. Furthermore, we deploy VPNChecker in an Internet Service Provider's (ISP) environment to evaluate its effectiveness. The results show that VPNChecker can improve the coverage of sophisticated detection engines and serve as a complement to existing methods.
Chenxu Wang 0006, Jiangyi Yin, Zhao Li 0010, Qingyun Liu 0001
WWW3
2023 A Robust and Accurate Encrypted Video Traffic Identification Method via Graph Neural Network
abstract
The explosive growth of video traffic has brought major challenges for network providers to improve user experience. On account of traffic encryption, network providers need to identify encrypted video traffic first before adopting optimization approaches to them. Traditional encrypted video traffic identification methods try to reveal the pattern of video traffic by using statistical features, which are not robust enough in different network environments. Some sophisticated graph-based methods recently have shown their advantages for encrypted traffic identification. However, these works lack optimization when it comes to the video streaming scenario. Inspired by these works, we propose GraphV, a GNN-based approach for identifying encrypted video traffic. Specifically, we construct an information-rich graph structure enhanced by unique features of video transmission. Then the embedding representation of each graph can be obtained through a Bi-LSTM layer added to all the sequential nodes embedding on this graph. The experiments on a well-known dataset and two open-world datasets from different network environments we collected show that GraphV outperforms the existing methods, especially on the generalization ability of the model.
Zhao Li 0010, Jiangchao Chen, Xiaoqing Ma, Meijie Du, Qingyun Liu 0001
CSCWD1
2023 Detecting Fake-Normal Pornographic and Gambling Websites through one Multi-Attention HGNN
abstract
The rapid development of pornographic and gambling websites, fueled by the widespread abuse of information technology, has become a growing concern. They pose a serious threat to the physical and mental health of children and can also endanger personal property. Therefore, it is necessary to detect them. However, pornographic and gambling websites become more and more tricky, which shows fake-normal to evade censorship and challenges traditional content-based detection methods. Therefore, it is essential to rely on information about relationships between websites.We propose HMAN, one Multi-Attention Heterogeneous Graph Neural Network (HGNN) model to detect pornographic and gambling websites by integrating content features and structural information, even if they present fake-normal. By one multi-attention mechanism consisting of explicit weight, self-attention and attention mechanism, content features can be selectively utilized with the assistance of structural information. The experimental results show that our method achieves the best 95.1% Macro-Avg-F1 and outperforms all baselines. We also illustrate that all extracted metapaths do contribute to the detection, where the hyperlink, title/meta terms and IP address are relatively important.
Xiaoqing Ma, Chao Zheng 0001, Zhao Li 0010, Jiangyi Yin, Qingyun Liu 0001, Xunxun Chen
CSCWD3
2023 IDTracker: Discovering Illicit Website Communities via Third-party Service IDs
abstract
Illicit websites are restricted by governments and application marketplaces due to their detrimental impact on society. Third-party web services play a crucial role in enabling illicit webmasters to establish websites rapidly and evade detection. In this paper, we discover that third-party services usually assign unique credentials to website developers as their identifications (IDs). Websites using the same services with identical IDs are likely to be hosted on shared infrastructures and have textually similar domain names. This observation sparks the idea of building a community of illicit websites by leveraging third-party service IDs. Therefore, we design IDTracker, a novel system for detecting illicit website communities based on domain name semantic and infrastructure relationship features, which empower classification algorithms to achieve a high F1 score of 0.8968. Furthermore, we deploy IDTracker on an Internet Service Provider's (ISP) environment for three months and identify 6,830 illicit communities containing 165,378 illicit websites. Many of these illicit websites can not be identified by the most sophisticated engines, such as Symantec and Baidu, because of the cloaking tactics. In addition, we conduct a large-scale and long-term measurement on the network infrastructures and third-party services of illicit communities, revealing new phenomena. Our findings can help security communities to thwart illicit websites more effectively.
Chenxu Wang 0006, Zhao Li 0010, Jiangyi Yin, Zhenni Liu, Qingyun Liu 0001
DSN2
2023 Shrink: Identification of Encrypted Video Traffic Based on QUIC
abstract
With the increasing prevalence of network videos, video traffic has become a significant portion of overall network traffic. Due to the presence of harmful content such as pornography and violence in network videos, network monitoring is necessary. However, the encryption of videos poses challenges for network monitoring. More and more video service providers are adopting QUIC as the default video transmission protocol to accelerate data transfer speeds. However, the existing methods for identifying encrypted video traffic do not apply to QUIC. Video service providers typically employ Content Delivery Network (CDN) technology to enhance user experience, which can result in missing video chunks for side-channel identification. Additionally, fluctuations in network conditions can lead to the retransmission of video chunks. This paper proposes Shrink, a QUIC-based encrypted video traffic identification method. It effectively extracts video chunks from online QUIC encrypted video traffic and proposes a bucket structure and global-local match to alleviate the issues of video chunks retransmission and loss. Furthermore, a bucket word dictionary is designed to enhance the method’s running speed. Experimental results demonstrate that Shrink performs well in real network environments, exhibiting superior accuracy and speed compared to existing state-of-the-art methods.
Weitao Tang, Meijie Du, Zhao Li 0010, Zhou Zhou 0007, Qingyun Liu 0001
IPCCC3
2022 Node-Imbalance Learning on Heterogeneous Graph for Pirated Video Website Detection
abstract
With the rapid development of video streaming, the problem of copyright infringement has become increasingly severe. Despite its explicit illegality in many countries, a large variety of pirated video websites are still active, causing huge damage to copyright holders and security risks to users. Traditional methods for detecting malicious websites, such as blacklists or feature-based classifiers, can be easily bypassed by evading approaches like Domain-Flux. Some researchers recently proposed sophisticated graph-based methods to utilize various relations between websites and convert the detection task into node representation learning. However, the node imbalance issue impairs their performance on real-world datasets. In this paper, given the limitations of the above methods, we propose a model named Heterogeneous Graph Node Re-weighting (HGNR) to detect pirated video websites. We construct a heterogeneous graph with diverse meta relations and design a weight adjustment mechanism to deal with node imbalance issue. The experiments with different imbalance ratios show that HGNR outperforms state-of-the-art graph-based methods. Furthermore, we analyze the best-performed meta relation and disclose how video pirates gain profits, which can help the security community thwart video piracy.
Jiangyi Yin, Zhao Li 0010, Rong Yang 0008, Meijie Du
CSCWD3
2022 P4-NSAF: defending IPv6 networks against ICMPv6 DoS and DDoS attacks with P4
abstract
Internet Protocol Version 6 (IPv6) is expected for widespread deployment worldwide. Such rapid development of IPv6 may lead to safety problems. The main threats in IPv6 networks are denial of service (DoS) attacks and distributed DoS (DDoS) attacks. In addition to the similar threats in Internet Protocol Version 4 (IPv4), IPv6 has introduced new potential vulnerabilities, which are DoS and DDoS attacks based on Internet Control Message Protocol version 6 (ICMPv6). We divide such new attacks into two categories: pure flooding attacks and source address spoofing attacks. We propose P4-NSAF, a scheme to defend against the above two IPv6 DoS and DDoS attacks in the programmable data plane. P4-NSAF uses Count-Min Sketch to defend against flooding attacks and records information about IPv6 agents into match tables to prevent source address spoofing attacks. We implement a prototype of P4-NSAF with P4 and evaluate it in the programmable data plane. The result suggests that P4-NSAF can effectively protect IPv6 networks from DoS and DDoS attacks based on ICMPv6.
Zhou Zhou 0007, Qingyun Liu 0001, Zhao Li 0010
ICC5
2022 Fighting Against Piracy: An Approach to Detect Pirated Video Websites Enhanced by Third-party Services
abstract
Along with the development of video streaming, the increasing number of pirated video websites has caused unprecedented damage to copyright holders and potential security risks to their users. Though many efforts have been made to take down pirated video websites, they are still emerging by utilizing evading approaches like Fast-Flux domains and Cybercrime-as-a-Service(CaaS) tools. In this paper, to detect pirated video websites, we propose a Third-party Enhanced Pirated Video Website Classification Network (TEP-Net), which integrates both semantic features and relationship information between websites and their third-party services. More specifically, we apply CNN-BiLSTM-Attention to explore both character-level and domain-level textual embedding and utilize relationship information by constructing statistical features in classification. The experiment shows that TEP-Net achieves a significant performance compared with existing methods. Furthermore, we perform an in-depth analysis of the CaaS behind pirated video websites. Our research can help the security community fight against video piracy more precisely and effectively.
Zhao Li 0010, Jiangyi Yin, Meijie Du, Qingyun Liu 0001
ISCC1
2022 A Lightweight Graph-based Method to Detect Pornographic and Gambling Websites with Imperfect Datasets
abstract
With the widespread abuse of information technology, pornographic and gambling websites develop rapidly. They affect the physical and mental health of children and endanger personal property. Therefore, it is necessary to detect them. However, the existing detection methods ignored that imperfect datasets are common in the scenario of pornographic and gambling websites which are hence adverse to the detection. Those imperfections specifically include sparse samples, mismatch and imbalanced datasets. In addition, over-reliance on visual features incurred high overhead.To overcome these shortcomings, we innovatively propose a lightweight graph-based method to detect pornographic and gambling websites through semi-supervised learning of textual content. The semi-supervised learning is to solve sparse samples and mismatch datasets, while the graph-based approach can combine the semi-supervised part with community discovery to deal with imbalanced datasets. Specifically, we perform the detection process with the utilization of modified TF-IDF and Louvain during the iteration and updating by the EM algorithm. The experimental results show that our method achieves the best 92.01% Macro-Avg-F1 with the shortest CPU time and outperforms all baselines. We also illustrate that the designed components in our model do contribute to the detection.
Xiaoqing Ma, Chao Zheng 0001, Zhao Li 0010, Jiangyi Yin, Qingyun Liu 0001, Xunxun Chen
TrustCom3
2021 CDNFinder: Detecting CDN-hosted Nodes by Graph-Based Semi-Supervised Classification
abstract
As a crucial internet infrastructure, Content Delivery Network (CDN) is widely deployed. Detecting CDN-hosted nodes from network traffic is important for Quality of Service (QoS), malware detection and firewall rule-sets. Current researches use hand-crafted rules, classification or clustering methods. However, those methods relying on plaintext are limited by the invisibility of plaintext due to encryption, as well as the limitations of DNS Resource Records, such as unreliability. Besides, those methods don't dig the structural information of domains and IPs. To overcome those shortcomings, we present CDNFinder, a novel method to detect CDN-hosted nodes by graph-based semi-supervised classification. Based on the active datasets collected in 10 vantage points, we construct the graph and extract innovative attributes. By modifying Graph Neural Network (GNN), CDNFinder outperforms classical machine learning methods, especially in recall rate (around 98%). Meanwhile, CDNFinder shortens the runtime of classical GNN algorithm by about 31% with no loss in metrics.
Xiaoqing Ma, Chao Zheng 0001, Zhao Li 0010, Qingyun Liu 0001, Xunxun Chen
ISCC3
2020 Predicting User Influence in the Propagation of Toxic Information
Yishuo Zhang, Penghui Jiang, Zhao Li 0010, Qingyun Liu 0001
KSEM (1)4