Shuzhan Ye

dblp:331/3492 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0001-3110-617XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 PrivAGS: Differentially Private Attributed Graph Synthesis
abstract
Attributed graphs are extensively utilized in marketing, friend recommendations, disease prediction, etc. In attributed graphs, nodes are associated with attributes to enrich the graph representation, while edges indicate relationships between nodes. However, ensuring data privacy when publishing attributed graphs is a significant challenge due to the sensitive nature of both attributes and relationships. Existing methods fail to preserve graph structures effectively and neglect correlations among node attributes, leading to diminished utility for published synthetic graphs. To address these issues, we propose PrivAGS, a framework for publishing attributed graphs with Rényi Differential Privacy (RDP) guarantees. PrivAGS reconstructs graph structures and attributes based on community structures to capture tightly connected features. We propose a bounded Gaussian threshold mechanism to preserve attribute correlations and utilize probabilistic graph models with optimized inference structures to infer distributions and release node attributes. Additionally, PrivAGS introduces a new structural model, MCEG, to capture clustering structures and enable efficient graph reconstruction. Extensive experiments on five real-world datasets show that PrivAGS generates privacy-preserving, high-utility synthetic data.
Shuzhan Ye, Lu Chen 0001, Zhikun Zhang 0001, Yunjun Gao, Yuxiang Wang 0001, Xiaoliang Xu 0001
Proc. ACM Manag. Data1
2024 Scalable Community Search with Accuracy Guarantee on Attributed Graphs
abstract
Given an attributed graph$G$and a query node$q$, Community Search over Attributed Graphs (CS-AG) aims to find a structure- and attribute-cohesive subgraph from$G$that contains$q$. Although CS-AG has been widely studied, they still face three challenges. (1) Exact methods based on graph traversal are time-consuming, especially for large graphs. Some tailored indices can improve efficiency, but introduce nonnegligible storage and maintenance overhead. (2) Approximate methods with a loose approximation ratio only provide a coarse-grained evaluation of a community's quality, rather than a reliable evaluation with an accuracy guarantee in runtime. (3) Attribute cohesiveness metrics often ignores the important correlation with the query node$q$. We formally define our CS-AG problem atop a$q- \mathbf{centric}$attribute cohesiveness metric considering both textual and numerical attributes, for$k-\mathbf{core}$model on homogeneous graphs. We show the problem is NP-hard. To solve it, we first propose an exact baseline with three pruning strategies. Then, we propose an index-free sampling-estimation-based method to quickly return an approximate community with an accuracy guarantee, in the form of a confidence interval. Once a good result satisfying a user-desired error bound is reached, we terminate it early. We extend it to heterogeneous graphs,$k-\mathbf{truss}$model, and size-bounded CS. Comprehensive experimental studies on ten real-world datasets show its superiority, e.g., at least$1.54\times (41.1\times$on average) faster in response time and a reliable relative error (within a user-specific error bound) of attribute cohesiveness is achieved.
Yuxiang Wang 0001, Shuzhan Ye, Yuxia Geng, Zhenghe Zhao, Xiangyu Ke, Tianxing Wu 0001
ICDE2
2022 Approximate and Interactive Processing of Aggregate Queries on Knowledge Graphs: A Demonstration
abstract
This paper demonstrates AGQ [26] - our system for approximate and interactive processing of aggregate queries on knowledge graphs (KGs), e.g., "what is the average price of cars produced in Germany?" One can support aggregate queries based on factoid queries, e.g., "find all cars produced in Germany", by applying an aggregate operation on factoid queries' answers. However, this straightforward method is problematic since both the accuracy and efficiency of factoid query processing would impact the performance of aggregate queries. Moreover, returning a one-time, exact result might add computation overhead and hinder users' engagement and interactivity.
Yuxiang Wang 0001, Arijit Khan 0001, Shuzhan Ye, Shihuang Pan, Yuhan Zhou 0001
CIKM4