VLDB 2026 Research / reviewers in the wild / expert
Jiajie He 0003
dblp:03/8106-3
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0009-7956-8355ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Auditing Approximate Machine Unlearning for Differentially Private ModelsabstractApproximate machine unlearning aims to remove the effect of specific data from trained models to ensure individuals' privacy. Existing methods focus on the removed records and assume the retained ones are unaffected. However, recent studies on the privacy onion effect indicate this assumption might be incorrect. Especially when the model is differentially private, no study has explored whether the retained ones still meet the differential privacy (DP) criteria under existing machine unlearning methods. This paper takes a holistic approach to auditing both unlearned and retained samples' privacy risks after applying approximate unlearning algorithms. We propose the privacy criteria for unlearned and retained samples, respectively, based on the perspectives of DP and membership inference attacks (MIAs). To make the auditing process more practical, we also develop an efficient MIA, A-LiRA, utilizing data augmentation to reduce the cost of shadow model training. Our experimental findings indicate that existing approximate machine unlearning algorithms may inadvertently compromise the privacy of retained samples for differentially private models, and that we need differentially private unlearning algorithms. For reproducibility, we have published our code: https://anonymous.4open.science/r/Auditing-machine-unlearning-CB10/README.md Yuechun Gu, Jiajie He 0003, Keke Chen |
ICDM | 2 |
| 2025 | Adaptive Domain Inference Attack with Concept HierarchyabstractWith increasingly deployed deep neural networks in sensitive application domains, such as healthcare and security, it's essential to understand what kind of sensitive information can be inferred from these models. Most known model-targeted attacks assume attackers have learned the application domain or training data distribution to ensure successful attacks. Can removing the domain information from model APIs protect models from these attacks? This paper studies this critical problem. Unfortunately, even with minimal knowledge, i.e., accessing the model as an unnamed function without leaking the meaning of input and output, the proposed adaptive domain inference attack (ADI) can still successfully estimate relevant subsets of training data. We show that the extracted relevant data can significantly improve, for instance, the performance of model-inversion attacks. Specifically, the ADI method utilizes the concept hierarchy extracted from the public and private datasets that the attacker can access and applies a novel algorithm to adaptively tune the likelihood of leaf concepts in the hierarchy showing up in the unseen training data. For comparison, we also designed a straightforward hypothesis-testing-based attack -- LDI. The ADI attack not only extracts partial training data at the concept level but also converges fastest and requires the fewest target-model accesses among all candidate methods. Our code is available at https://anonymous.4open.science/r/KDD-362D. Yuechun Gu, Jiajie He 0003, Keke Chen |
KDD (1) | 2 |
| 2025 | RecPS: Privacy Risk Scoring for Recommender SystemsabstractRecSys '25: Nineteenth ACM Conference on Recommender Systems Prague Czech Republic September 22 - 26, 2025 Jiajie He 0003, Yuechun Gu, Keke Chen |
RecSys | 1 |
| 2024 | Demo: FT-PrivacyScore: Personalized Privacy Scoring Service for Machine Learning ParticipationabstractData privacy has been a top concern in the AI era. Despite the recent development of differentially private learning methods, controlled data access remains a mainstream method for protecting data privacy in many industrial and research environments. In controlled data access, authorized model builders work in a restricted environment to access sensitive data, which can fully preserve data utility with reduced risk of data leak. However, unlike differential privacy, there is no quantitative measure for individual data contributors to tell their privacy risk before participating in a machine learning task. We developed the demo prototype FT-PrivacyScore to show that it's possible to efficiently and quantitatively estimate the privacy risk of participating in a model fine-tuning task. The demo source code will be available at https://github.com/RhincodonE/demo_privacy_scoring. Yuechun Gu, Jiajie He 0003, Keke Chen |
CCS | 2 |