Jiangyi Yin

dblp:256/6635 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0002-9037-4676ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PriAgent: A Collaborative Multi-Agent Framework for Auditing Android Privacy Compliance
abstract
Stringent regulations like General Data Protection Regulation (GDPR) mandate that an application's code-level data handling must align with its natural-language privacy policy, creating a critical auditing challenge. However, existing methods, predominantly reliant on static analysis, suffer from a critical limitation: in their pursuit of soundness via over-approximation, they exhibit "semantic blindness"—detecting what data flows exist but not why. This leads to an overwhelming volume of false positives, rendering automated auditing impractical. To bridge this gap, we introduce PriAgent, a novel framework that approaches compliance auditing as a multi-stage, AI-driven reasoning task. Instead of a monolithic model, PriAgent deploys a team of specialized agents that execute a divide-and-conquer strategy. They systematically prune the analysis space by abstracting data flows, pinpoint semantic loci critical for inspection, and perform on-demand summarization of large code blocks to ensure scalability. PriAgent leverages Retrieval-Augmented Generation (RAG) with a curated knowledge base of Android APIs, equipping agents to discern potentially non-compliant behavior from benign functionality. By correlating code-level evidence with the app's stated privacy policy, PriAgent delivers a holistic and explainable verdict for each potential violation. Our evaluations demonstrate that PriAgent significantly reduces false positives, enabling a more scalable and precise compliance audit.
Zhao Li 0010, Zhuojun Jiang, Jiangyi Yin, Jiangchao Chen, Qingyun Liu 0001
AAAI4
2025 Leveraging Cross-Layer Network Probing to Detect Stealth Services
abstract
Stealth services have become increasingly popular due to growing demand for privacy protection. To avoid detection by active probing, many have adopted probe-resistant strategies. In this work, we design a suite of carefully crafted probes to expose their hidden vulnerabilities. Through detailed analysis of the corresponding responses, we find that despite the defensive measures implemented, certain implicit information, such as protocol stack fingerprints, can still serve as strong indicators. We present SSChecker, a detection system that combines cross-layer probing with a classification model inspired by Information Bottleneck Theory to address the data sparsity issue inherent in active probing. Our experiments on real-world datasets show that SSChecker outperforms existing methods, including leading industrial detection engines. Our findings demonstrate that current probe-resistant strategies of stealth services remain insufficient. To strengthen privacy protection, we further propose mitigation strategies that help stealth services enhance their ability.
Jiangyi Yin, Chenxu Wang 0006, Zhao Li 0010, Zhuojun Jiang, Jiangchao Chen, Dongfang Hao, Qingyun Liu 0001
TrustCom1
2024 Identifying VPN Servers through Graph-Represented Behaviors
abstract
Identifying VPN servers is a crucial task in various situations, such as geo-fraud detection, bot traffic analysis and network attack identification. Although numerous studies that focus on network traffic detection have achieved excellent performance in closed-world scenarios, particularly those methods based on deep learning, they may exhibit significant performance degradation due to changes in network environment. To mitigate this issue, a few studies have attempted to use methods based on active probing to detect VPN servers. However, these methods still have two limitations. They cannot handle situations without probing responses and are limited in applicability due to their focus on specific VPNs. In this work, we propose VPNChecker, which utilizes the graph-represented behaviors to detect VPN servers in real-world scenarios. VPNChecker outperforms existing methods in four offline datasets. The results from our datasets, containing multiple different VPNs, indicate that VPNChecker has better applicability. Furthermore, we deploy VPNChecker in an Internet Service Provider's (ISP) environment to evaluate its effectiveness. The results show that VPNChecker can improve the coverage of sophisticated detection engines and serve as a complement to existing methods.
Chenxu Wang 0006, Jiangyi Yin, Zhao Li 0010, Qingyun Liu 0001
WWW2
2023 Detecting Fake-Normal Pornographic and Gambling Websites through one Multi-Attention HGNN
abstract
The rapid development of pornographic and gambling websites, fueled by the widespread abuse of information technology, has become a growing concern. They pose a serious threat to the physical and mental health of children and can also endanger personal property. Therefore, it is necessary to detect them. However, pornographic and gambling websites become more and more tricky, which shows fake-normal to evade censorship and challenges traditional content-based detection methods. Therefore, it is essential to rely on information about relationships between websites.We propose HMAN, one Multi-Attention Heterogeneous Graph Neural Network (HGNN) model to detect pornographic and gambling websites by integrating content features and structural information, even if they present fake-normal. By one multi-attention mechanism consisting of explicit weight, self-attention and attention mechanism, content features can be selectively utilized with the assistance of structural information. The experimental results show that our method achieves the best 95.1% Macro-Avg-F1 and outperforms all baselines. We also illustrate that all extracted metapaths do contribute to the detection, where the hyperlink, title/meta terms and IP address are relatively important.
Xiaoqing Ma, Chao Zheng 0001, Zhao Li 0010, Jiangyi Yin, Qingyun Liu 0001, Xunxun Chen
CSCWD4
2023 IDTracker: Discovering Illicit Website Communities via Third-party Service IDs
abstract
Illicit websites are restricted by governments and application marketplaces due to their detrimental impact on society. Third-party web services play a crucial role in enabling illicit webmasters to establish websites rapidly and evade detection. In this paper, we discover that third-party services usually assign unique credentials to website developers as their identifications (IDs). Websites using the same services with identical IDs are likely to be hosted on shared infrastructures and have textually similar domain names. This observation sparks the idea of building a community of illicit websites by leveraging third-party service IDs. Therefore, we design IDTracker, a novel system for detecting illicit website communities based on domain name semantic and infrastructure relationship features, which empower classification algorithms to achieve a high F1 score of 0.8968. Furthermore, we deploy IDTracker on an Internet Service Provider's (ISP) environment for three months and identify 6,830 illicit communities containing 165,378 illicit websites. Many of these illicit websites can not be identified by the most sophisticated engines, such as Symantec and Baidu, because of the cloaking tactics. In addition, we conduct a large-scale and long-term measurement on the network infrastructures and third-party services of illicit communities, revealing new phenomena. Our findings can help security communities to thwart illicit websites more effectively.
Chenxu Wang 0006, Zhao Li 0010, Jiangyi Yin, Zhenni Liu, Qingyun Liu 0001
DSN3
2022 Node-Imbalance Learning on Heterogeneous Graph for Pirated Video Website Detection
abstract
With the rapid development of video streaming, the problem of copyright infringement has become increasingly severe. Despite its explicit illegality in many countries, a large variety of pirated video websites are still active, causing huge damage to copyright holders and security risks to users. Traditional methods for detecting malicious websites, such as blacklists or feature-based classifiers, can be easily bypassed by evading approaches like Domain-Flux. Some researchers recently proposed sophisticated graph-based methods to utilize various relations between websites and convert the detection task into node representation learning. However, the node imbalance issue impairs their performance on real-world datasets. In this paper, given the limitations of the above methods, we propose a model named Heterogeneous Graph Node Re-weighting (HGNR) to detect pirated video websites. We construct a heterogeneous graph with diverse meta relations and design a weight adjustment mechanism to deal with node imbalance issue. The experiments with different imbalance ratios show that HGNR outperforms state-of-the-art graph-based methods. Furthermore, we analyze the best-performed meta relation and disclose how video pirates gain profits, which can help the security community thwart video piracy.
Jiangyi Yin, Zhao Li 0010, Rong Yang 0008, Meijie Du
CSCWD2
2022 Fighting Against Piracy: An Approach to Detect Pirated Video Websites Enhanced by Third-party Services
abstract
Along with the development of video streaming, the increasing number of pirated video websites has caused unprecedented damage to copyright holders and potential security risks to their users. Though many efforts have been made to take down pirated video websites, they are still emerging by utilizing evading approaches like Fast-Flux domains and Cybercrime-as-a-Service(CaaS) tools. In this paper, to detect pirated video websites, we propose a Third-party Enhanced Pirated Video Website Classification Network (TEP-Net), which integrates both semantic features and relationship information between websites and their third-party services. More specifically, we apply CNN-BiLSTM-Attention to explore both character-level and domain-level textual embedding and utilize relationship information by constructing statistical features in classification. The experiment shows that TEP-Net achieves a significant performance compared with existing methods. Furthermore, we perform an in-depth analysis of the CaaS behind pirated video websites. Our research can help the security community fight against video piracy more precisely and effectively.
Zhao Li 0010, Jiangyi Yin, Meijie Du, Qingyun Liu 0001
ISCC3
2022 A Lightweight Graph-based Method to Detect Pornographic and Gambling Websites with Imperfect Datasets
abstract
With the widespread abuse of information technology, pornographic and gambling websites develop rapidly. They affect the physical and mental health of children and endanger personal property. Therefore, it is necessary to detect them. However, the existing detection methods ignored that imperfect datasets are common in the scenario of pornographic and gambling websites which are hence adverse to the detection. Those imperfections specifically include sparse samples, mismatch and imbalanced datasets. In addition, over-reliance on visual features incurred high overhead.To overcome these shortcomings, we innovatively propose a lightweight graph-based method to detect pornographic and gambling websites through semi-supervised learning of textual content. The semi-supervised learning is to solve sparse samples and mismatch datasets, while the graph-based approach can combine the semi-supervised part with community discovery to deal with imbalanced datasets. Specifically, we perform the detection process with the utilization of modified TF-IDF and Louvain during the iteration and updating by the EM algorithm. The experimental results show that our method achieves the best 92.01% Macro-Avg-F1 with the shortest CPU time and outperforms all baselines. We also illustrate that the designed components in our model do contribute to the detection.
Xiaoqing Ma, Chao Zheng 0001, Zhao Li 0010, Jiangyi Yin, Qingyun Liu 0001, Xunxun Chen
TrustCom4