VLDB 2026 Research / reviewers in the wild / expert
Yang Li 0215
dblp:37/4190-215
· DBLP profile ↗
15ranked-venue papers
1as first author
15since 2021 · last 2026
0009-0005-9585-3472ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 5 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unreachable Features? Exposing the Security Risks of Invisible Interfaces in Embedded Web Services of IoT DevicesabstractIoT devices, now integral to our daily routines, offer unparalleled convenience but also face mounting security threats. Embedded web services, prevalent in public networks, pose a major risk to these devices. While research has focused on detecting vulnerabilities in IoT embedded web services, it has overlooked the presence of invisible interfaces, which have emerged as significant security threats. In this paper, we propose InvRadar, a novel framework for detecting vulnerabilities in invisible interfaces of embedded web services in IoT devices. Specifically, InvRadar identifies invisible interfaces by analyzing the differences between the front-end visible interface keywords and the back-end interface keywords through a correlation analysis method. Subsequently, InvRadar uses a static taint analysis method to detect the vulnerabilities that can be triggered by the invisible interfaces. To validate the performance of InvRadar, we conduct extensive experiments and compare InvRadar with the state-of-the-art methods. In testing 13 device firmware, InvRadar identifies 1,793 invisible interfaces and detects 124 vulnerabilities, including 53 newly discovered ones, with 34 receiving new CVE/CNVD IDs. Additionally, InvRadar outperforms the state-of-the-art methods in interface keyword extraction, border binary and data ingestion function identification. Yuanchao Chen, Yuwei Li 0002, Yi Shen 0012, Yu Chen 0053, Yang Li 0215, Taiyan Wang, Yuliang Lu, Zulie Pan, Shouling Ji |
IEEE Internet Things J. | 5 |
| 2026 | OwnerHunter: Multilingual Website Owner Identification Powered by Large Language ModelabstractAs cyberspace continues to expand, identifying the organization or individual behind a website has become increasingly vital in security incident response, phishing website detection, and other cybersecurity subfields. An existing solution for it involves analyzing webpage content and extracting owner names using named entity recognition techniques. However, since these techniques operate on a sentence-by-sentence basis, they struggle to identify the true owner when multiple individual or organizational names appear on a webpage. Moreover, they often perform poorly on non-English websites. To address these limitations, we propose OwnerHunter, a novel multilingual framework powered by large language models, which formulates website owner identification as a multilingual document-level information extraction task and utilizes global information from webpages to identify the owner. In OwnerHunter, we first craft prompts that fully leverage the capabilities of large language models to effectively recognize potential owners on webpages in different languages with minimal examples. To enhance the comprehensiveness and accuracy of recognition, we further design a multimodal augmentation strategy, an example pool strategy, and a self-verification strategy. Then, we devise a semantic and string similarity aggregation-based entity disambiguation technique to eliminate ambiguities among multiple potential owners recognized by large language models and a position-based hybrid ranking technique to exactly select the true owner. To evaluate OwnerHunter, we refine the publicly available English dataset ONER and construct the Chinese dataset WOI-cn with 16,036 real websites. Experimental results show that OwnerHunter achieves F1 scores of 0.9505 on ONER and 0.9621 on WOI-cn, setting new state-of-the-art performance on both datasets. Cheng Tu, Enhuan Dong, Zexiang Zhang, Min Zhang 0054, Yang Li 0215, Jiahai Yang 0001 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2026 | MetaRAG: Identifying Website Owner Using Meta-Path-Guided Dynamic Graph Retrieval-Augmented GenerationabstractWebsite owner identification aims to link websites to their real-world owners, which is crucial for credibility assessment and information provenance in information retrieval and vital for applications in cybersecurity, Internet governance, and digital regulation. Existing approaches for website owner identification primarily rely on querying infrastructure registration records or analyzing webpage content. However, these methods often fail due to incomplete or outdated registration records and sparse webpage content. We observe that inter-website relationships, derived from shared infrastructure data such as primary domains, IP blocks, and geolocations, can provide valuable but underutilized ownership cues. To exploit this insight, we propose MetaRAG, a meta-path-guided dynamic graph retrieval-augmented generation framework that performs reasoning using large language models over ownership-relevant paths in a website-centric knowledge graph. MetaRAG consists of three components: (1) a knowledge graph construction module that integrates infrastructure data and crawled webpage content into a unified representation; (2) a meta-path-guided dynamic reasoning module that constrains retrieval to ownership-relevant meta-paths and adaptively decides whether to retrieve more information or perform inference based on evidence completeness; and (3) a multi-path evidence refinement module that aggregates and scores retrieved paths to suppress noise and distill high-confidence ownership signals. We evaluate MetaRAG on two constructed real-world datasets, achieving up to 6.82% improvement over strong baselines. The results demonstrate the effectiveness of our approach in combining structured web knowledge with large language model-based reasoning for more accurate website owner identification. Cheng Tu, Yunshan Ma 0002, Bingyang Guo, Qianyu Li 0001, Yang Li 0215, Min Zhang 0054, Fan Shi 0003, Xiang Wang 0010 |
ACM Trans. Inf. Syst. | 5 |
| 2025 | Insvdf: Interface-State-Aware Virtual Device FuzzingabstractHypervisor is the core technology of virtualization for emulating independent hardware resources for each virtual machine. Virtual devices serve as the main interface of the hypervisor, making the security of virtual devices crucial, as any vulnerabilities can impact the entire virtualization environment and pose a threat to the host machine's security. Direct Memory Access (DMA) is the interface of virtual devices, enabling communication with the host machine. Recently, many efforts have focused on fuzzing against DMA to discover the hypervisor's vulnerabilities. However, the lack of sensitivity to the DMA state causes these efforts to be hindered in efficiency during fuzzing. Specifically, there are two main issues: the uncertain interaction moment and the unclear interaction depth. In this paper, we introduce InSVDF, a DMA interface stateaware fuzzing engine. InSVDF first models the intra-interface state of the DMA interface and incorporates an asynchronyaware state snapshot mechanism along with a depth-aware seed preservation mechanism. To validate our approach, we compare InSVDF with a state-of-the-art fuzzer. The results demonstrate that InSVDF significantly enhances vulnerability discovery speed, with improvements of up to 24.2 x in the best case. Furthermore, InSVDF has identified 2 new vulnerabilities, one of which has been assigned a CVE ID. Zexiang Zhang, Yiming Tao, Zulie Pan, Cheng Tu, Min Zhang 0054, Yang Li 0215, Yi Shen 0012, Chunming Wu 0001 |
ICSE | 8 |
| 2025 | DMut: Optimize Mutation Strategy in Directed Greybox Fuzzing by Multi-Population Genetic AlgorithmabstractDirected greybox fuzzing has become a crucial technique for discovering vulnerabilities in software. The seed mutation plays an important role in fuzzing by generating new inputs that explore diverse program states and find the target vulnerability. While seed mutation is critical to the effectiveness of fuzzing, most existing mutation strategies are designed for coverage-based fuzzing and lack the guidance required in directed scenarios. This limits the quality of generated testcases and reduces fuzzing efficiency in directed greybox fuzzing.In this paper, we propose DMut, a novel seed mutation strategy based on a multi-population genetic algorithm, designed to address these limitations. DMut models the seed mutation process using a genetic algorithm, optimizing the seed mutation probability distribution and iteratively evolving it to generate higher-quality testcases. The approach incorporates a well-designed fitness function and selection strategy that aligns with directed fuzzing scenarios to guide the evolution of the mutation strategy. Through comprehensive experiments on real-world CVEs, we demonstrate that DMut significantly improves the effectiveness of directed fuzzing. Compared to the widely adopted directed greybox fuzzing tool, AFLGo, DMut reduces the time to expose the target vulnerability by 41% on average. Additionally, DMut improves path exploration efficiency, covering more unique execution paths and speeding up the exploration process. In summary, the experimental results show that DMut provides a robust, efficient method for improving directed fuzzing performance, offering a significant advancement over existing approaches. Tingke Wen, Yuwei Li 0002, Huimin Ma 0004, Yang Li 0215, Zulie Pan |
SMC | 5 |
| 2025 | Website Owner Identification through Multi-level Contrastive Representation LearningabstractWebsite owner identification aims to recognize the organization or individual who owns a given website that is served on the web. It is a crucial step for cyberspace surveying and mapping, playing a significant role in cyberspace administration and governance. Existing widely employed solutions for website owner identification mainly fall into two paradigms: (1) querying the public information databases such as WHOIS, which store the Internet resource’s registered users or assignees; and (2) directly extracting the organization or individual name of the website owner from the webpage using the technique of named entity recognition. However, the former is less reliable due to the incomplete, encrypted, and outdated records in the public information databases. Meanwhile, the latter requires that the webpages explicitly and precisely present their owner names without ambiguity, which is often hard to guarantee in practice. To address these limitations, we propose to formulate website owner identification as a problem of webpage representation learning, thereby introducing a novel representation learning framework empowered by large language model-based text Rewriting and Multi-level contrastive learning, named ReMon. First, we devise a prompt to rewrite the webpages using large language models, which effectively filters out noise from the original webpages. Second, we model website–website, website–owner, and owner–owner interactions through multi-level contrastive learning, fully utilizing the self-supervision signals on long-tail items to learn the multi-level constraints. Third, we design a retrieval-based prediction framework and a clustering-based framework to apply websites’ and owners’ representations for different scenarios of the website owner identification task. To evaluate ReMon under our formulation, we construct two datasets based on real-world data. Compared to existing approaches, our ReMon can address the challenging scenarios when valid information cannot be found in public information databases and the owner’s name does not appear on the webpage. Meanwhile, the experimental results show that ReMon outperforms all representation learning-based baselines and significantly enhances training efficiency. The code is available at https://github.com/tuchen9/ReMon . Cheng Tu, Yunshan Ma 0002, Yang Li 0215, Min Zhang 0054, Fan Shi 0003, Xiang Wang 0010 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2024 | DynPen: Automated Penetration Testing in Dynamic Network Scenarios Using Deep Reinforcement LearningabstractPenetration testing, a crucial industrial practice for securing networked systems and infrastructures, has traditionally depended on the extensive expertise of human professionals. Addressing the scarcity of human experts, the development of automated penetration testing tools emerges as a promising avenue. Against the backdrop of rapid advancements in artificial intelligence technologies, reinforcement learning has demonstrated considerable potential for realizing automated penetration testing. However, existing research predominantly concentrates on reinforcement learning-based automated penetration testing tools within static scenarios, with limited exploration in dynamic network environments. This paper addresses a noteworthy challenge in developing autonomous agents for real-world applications, particularly focusing on scenarios marked by environmental changes. Such alterations necessitate autonomous agents to continuously monitor environmental characteristics, and adapt, and adjust learned actions to ensure the system’s effective operation. Consequently, the paper proposes an automated reinforcement learning-based penetration testing scheme tailored for dynamic network scenarios, named DynPen. DynPen captures observed changes in the scenario, aiding the penetration testing agent in decision-making based on historical experiences. Simulation results demonstrate the proposed scheme’s efficacy in significantly expediting the convergence speed of the penetration testing agent using reinforcement learning algorithms. Furthermore, the scheme successfully maintains the learning agility and adaptability of the agent in dynamic network scenarios. Qianyu Li 0001, Dong Li 0054, Fan Shi 0003, Min Zhang 0054, Anupam Chattopadhyay, Yi Shen 0012, Yang Li 0215 |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2023 | An Efficient Authentication and Key Agreement Scheme for CAV Internal Applications
Yang Li 0215, Qingyang Zhang 0001, Wenwen Cao, Jie Cui 0004, Hong Zhong 0001 |
CollaborateCom (2) | 1 |
| 2023 | SLCSA: Scalable Layered Cooperative Service Attestation Scheme in Cloud-Edge-End Cooperation EnvironmentsabstractIn a cloud-edge-end cooperation environment, edge and core cloud services are complementary and synergistic, jointly processing a large amount of private data uploaded by users. To prevent the leakage of private data, users must ensure that services are secure and trusted through remote attestation. Traditional one-to-one remote attestation schemes are typically used to test the cloud services. However, as the cloud platform scales and the number of edge and core cloud services grows rapidly, the traditional attestation method has problems, such as poor scalability and low attestation efficiency. Thus far, there has been a lack of feasible methods for users to verify multiple related services in a cloud-edge-end cooperation environment quickly. This paper presents a scalable layered cooperative service attestation (SLCSA) scheme, the first secure and scalable protocol for the efficient attestation of multiple cooperative services. The SLCSA scheme is based on a Boneh–Lynn–Shacham (BLS) multi-signature to improve the scalability of the scheme while enabling users to conduct the batch verification of services. We also analyze the security of the proposed scheme. To evaluate the proposed scheme, we implement it using Intel SGX, which can provide basic hardware-assisted attestation and a trusted execution environment for services. The experimental results show that the SLCSA scheme is practical and efficient in a cloud-edge-end cooperative environment. Jie Cui 0004, Qipeng Chen, Yang Li 0215, Qingyang Zhang 0001, Lu Liu 0001, Hong Zhong 0001 |
ICPADS | 4 |
| 2023 | AlphaEXP: An Expert System for Identifying Security-Sensitive Kernel Objects
Kaixiang Chen, Chao Zhang 0008, Zulie Pan, Qianyu Li 0001, Siliang Qin, Shenglin Xu, Min Zhang 0054, Yang Li 0215 |
USENIX Security Symposium | 9 |
| 2023 | INNES: An intelligent network penetration testing model based on deep reinforcement learning
Qianyu Li 0001, Min Zhang 0054, Yang Li 0215 |
Appl. Intell. | 5 |
| 2023 | A hierarchical deep reinforcement learning model with expert prior knowledge for intelligent penetration testing
Qianyu Li 0001, Min Zhang 0054, Yi Shen 0012, Yang Li 0215 |
Comput. Secur. | 6 |
| 2023 | Tunter: Assessing Exploitability of Vulnerabilities with Taint-Guided Exploitable States Exploration
Kaixiang Chen, Zulie Pan, Yuwei Li 0002, Qianyu Li 0001, Yang Li 0215, Min Zhang 0054, Chao Zhang 0008 |
Comput. Secur. | 6 |
| 2023 | Efficient Blockchain-Based Data Integrity Auditing for Multi-Copy in Decentralized StorageabstractAs the disruptor of cloud storage, decentralized storage could lead to a major shift in how organizations store data in the future. To ensure data availability, users generally encrypt the data and distribute it to multiple storage service providers. It is necessary to study data integrity verification in decentralized storage. Although some recent studies have proposed the using blockchain technology to assist auditing work in decentralized storage networks, the on-chain overhead still increases linearly with an increase in audit requests. Blockchain networks will inevitably be overloaded. In this study, we propose an efficient data integrity auditing scheme for multiple copies in decentralized storage. Particularly, using different polynomial commitment schemes, we first propose a basic scheme for verifying multiple copies of a single file, and then we propose an efficient batch auditing scheme for multiple copies of multiple files. Our scheme can significantly reduce the computation overhead of storage service providers while keeping the on-chain storage overhead constant. Security analysis and performance analysis show that our scheme is efficient and practical. Qingyang Zhang 0001, Jie Cui 0004, Hong Zhong 0001, Yang Li 0215, Chengjie Gu, Debiao He |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | Mining Centralization of Internet Service Infrastructure in the WildabstractThe last decade has witnessed the rapid evolution of the Internet structure, one of which is centralization, that is, Internet core infrastructure has been constantly transferred into the hands of a few popular market participants. Researchers are trying to measure centrality and analyze its security impact from the perspective of traffic analysis. But the underlying distribution of service providers is still enveloped in mysterious veils. In order to address this problem and assess the security risk associated with such centralization. Firstly, we performed linear regression on the data of each kind of service provider in the Alexa Top 1M domains to study the current underlying distribution of various services for the first time. The results show that Zipf’s law is universal in various service providers’ market share, which proves that Internet service infrastructures are centralized. Secondly, we explored the security impacts of centralized infrastructures on the Internet. we conducted attack simulations on providers. Results show that intentional attacks on core providers can greatly downgrade the performance of the Internet. To make matters worse, the quantitative analysis of the provider’s infrastructures found that a considerable number of provider’s infrastructures have low diversity. In addition, we proposed an algorithm to calculate the dependencies between different types of service providers and carried out an evaluation of our datasets, and found the tendency for different services to depend on each other. Our results indicate that the Internet is facing huge security challenges, because the centralized infrastructure will impair service redundancy, and at the same time, it will also cause dependence between infrastructures, which in turn strengthens its centralization. Bingyang Guo, Fan Shi 0003, Chengxi Xu, Min Zhang 0054, Yang Li 0215 |
MSN | 5 |