Tao Xue 0003

dblp:23/1877-3 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0002-2279-1988ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 3 · 2 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 NLACP Dataset Expansion With AI: Leveraging ChatGPT for Policy Generation
abstract
Natural Language Access Control Policies (NLACPs) play a pivotal role in access control systems, yet the scarcity of high-quality public datasets severely hinders advancements in NLP/ML-driven policy automation. Existing datasets, constrained by labor-intensive manual curation, confidentiality barriers, and rigid synthesis methods, suffer from limited scale, semantic inconsistency, and poor adaptability to emerging domains such as Smart Healthcare and the Internet of Vehicles. To address this critical gap, we proposeGenPy, a novel framework leveraging large language models (LLMs) like ChatGPT and domain-specific knowledge graphs to systematically generate scalable, semantically rich NLACPs datasets.GenPypioneers the integration of LLMs into NLACPs synthesis, offering three core innovations: (1) A generation module combining knowledge graphs with prompt engineering to produce diverse, domain-aligned policies; (2) An annotation module employing few-shot learning to filter non-policy statements and ensure quality; (3) A policy management module resolving redundancies and conflicts while enabling conflict injection for optimization research. By embedding domain semantics (e.g., Internet of Vehicles compliance with automotive data laws),GenPyenriches traditional datasets and supports fine-grained policy requirements. Experiments demonstrate thatGenPygenerates 5,000+ high-quality NLACPs across multiple domains, achieving an F1 score of exceeding 93% in NLACPs classification tasks. Notably, augmenting the classification models in RAGent withGenPy-generated data improves NLACPs classification accuracy by 2-9%, validating its practical utility. Our work establishes the first LLM-based pipeline that integrates knowledge graphs with customizable prompts to generate semantically rich NLACPs, supporting fine-grained access control models, NLACPs identification, and policy management across diverse security requirements. We providing an open-source repository of policies, code, and technical frameworks to catalyze future research in access control automation. The dataset and implementation are publicly available at https://github.com/fe1w0/NLACP-LLM.
Tao Xue 0003
IEEE Internet Things J.3
2025 Edge computing for IoT: Novel insights from a comparative analysis of access control models
Tao Xue 0003, Shuailou Li
Comput. Networks1
2024 Interpretable Risk-aware Access Control for Spark: Blocking Attack Purpose Behind Actions
abstract
The big data platform supports powerful data re-trieval and mining analysis, providing users with seamless access to extensive data for valuable insights. However, the increasing access to sensitive data raises privacy concerns. Existing studies utilize access control mechanisms to ensure secure data authorization. Nevertheless, previous approaches are deficient in facilitating risk control during real-time query processing and fail to elucidate the details of attack access. To address these limitations, we propose a novel Interpretable Risk-aware Access Control (IRAAC) for Spark - the advanced distributed engine for large-scale data computing in big data ecosystems. IRAAC utilizes the sequence representation techniques and contrastive learning idea from Natural Language Processing to learn patterns of attack queries for extracting critical attack subqueries. In terms of attack investigation, IRAAC designs specific templates to encourage large language models (LLMs) to provide a comprehensive delineation of potential query access risks.
Tao Xue 0003, Shuailou Li, Yu Wen 0001
ICCD2
2023 PRISPARK: Differential Privacy Enforcement for Big Data Computing in Apache Spark
abstract
Differential privacy has emerged as a gold standard privacy definition due to its persuasive mathematical guarantee. While various data protection mechanisms provide differential privacy for SQL queries of RDBMSs, enforcing differential privacy for big data platforms needs to be further researched. This work presents Prispark, which enforces differential privacy for Spark - the advanced distributed engine for large-scale data computing in big data ecosystems where sensitive data is often processed. Prispark targets to support various data processing (i.e., relational and unstructured queries) on Spark. In particular, to calculate a tighter sensitivity bound and improve the utility of results, we design the overall statistics estimation algorithm for estimating the upper bound of statistics with the filter condition, and propose a novel fine-grained operation-oriented rules set for calculating sensitivity of various relational and unstructured queries. Moreover, we propose a general differential privacy mechanism, Prispark, a suite including Prisparksql and Prisparkdag. We enforce Prisparksql at the Catalyst optimization layer for relational queries in Spark SQL and Prisparkdag at the RDD execution layer for unstructured queries in Spark core. Finally, we experimentally evaluate Prispark on TPC-H, TPC-DS, PigMix benchmarks, and real-world dataset LANL. The experimental results suggest that Prispark supports various applications/queries while improving the utility of all query results by orders of magnitude with negligible performance overhead.
Shuailou Li, Yu Wen 0001, Tao Xue 0003, Yanna Wu, Dan Meng 0002
SRDS3
2023 SparkAC: Fine-Grained Access Control in Spark for Secure Data Sharing and Analytics
abstract
With the development of computing and communication technologies, an extremely large amount of data has been collected, stored, utilized, and shared, while new security and privacy challenges arise. Existing access control mechanisms provided by big data platforms have limitations in granularity and expressiveness. In this article, we present SparkAC, a novel access control mechanism for secure data sharing and analysis in Spark. In particular, we first propose apurpose-aware access control(PAAC) model, which introduces new concepts ofdata processing purposeanddata operation purposeand an automatic purpose analysis algorithm that identifies purposes from data analytics operations and queries. Moreover, we develop a unified access control mechanism that implements PAAC model in two modules. GuardSpark++ supports structured data access control in Spark Catalyst and GuardDAG supports unstructured data access control in Spark core. Finally, we evaluate GuardSpark++ and GuardDAG with multiple data sources, applications, and data analytics engines. Experimental results show that SparkAC provides effective access control functionalities with very small (GuardSpark++) or medium (GuardDAG) performance overhead.
Tao Xue 0003, Yu Wen 0001, Bo Luo, Gang Li 0009, Yingjiu Li, Yanfei Hu, Dan Meng 0002
IEEE Trans. Dependable Secur. Comput.1
2020 GuardSpark++: Fine-Grained Purpose-Aware Access Control for Secure Data Sharing and Analysis in Spark
abstract
With the development of computing and communication technologies, extremely large amount of data has been collected, stored, utilized, and shared, while new security and privacy challenges arise. Existing platforms do not provide flexible and practical access control mechanisms for big data analytics applications. In this paper, we present GuardSpark++, a fine-grained access control mechanism for secure data sharing and analysis in Spark. In particular, we first propose a purpose-aware access control (PAAC) model, which introduces new concepts of data processing/operation purposes to conventional purpose-based access control. An automatic purpose analysis algorithm is developed to identify purposes from data analytics operations and queries, so that access control could be enforced accordingly. Moreover, we develop an access control mechanism in Spark Catalyst, which provides unified PAAC enforcement for heterogeneous data sources and upper-layer applications. We evaluate GuardSpark++ with five data sources and four structured data analytics engines in Spark. The experimental results show that GuardSpark++ provides effective access control functionalities with a very small performance overhead (average 3.97%).
Tao Xue 0003, Yu Wen 0001, Bo Luo, Yanfei Hu, Yingjiu Li, Gang Li 0009, Dan Meng 0002
ACSAC1