VLDB 2026 Research / reviewers in the wild / expert
Fang Liu 0006
dblp:67/5807-6
· DBLP profile ↗
14ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0003-3028-5927ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Differentially Private Weighted Empirical Risk Minimization Procedure and Its Application to Outcome Weighted LearningabstractData used to train predictive models via empirical risk minimization (ERM) often contain sensitive personal information. While differential privacy (DP) provides mathematically provable bounds to protect such data, previous work has focused almost exclusively on unweighted ERM. We consider weighted ERM (wERM) -- an important generalization where individual contributions to the objective function vary. We propose the first DP algorithm for general wERM with formal privacy guarantees and derive both its empirical and population utility bounds. Crucially, this general wERM framework provides a pathway for deriving privacy-preserving learning methods for individualized treatment rules, including the popular outcome-weighted learning (OWL) approach. We evaluate DP-wERM applied to OWL in simulated and real data experiments. Our empirical results demonstrate that training OWL models via wERM provides strong DP guarantees while maintaining robust performance, proving the method is practical for sensitive, real-world data. Spencer Giddens, Yiwang Zhou, Kevin R. Krull, Tara M. Brinkman, Peter X. K. Song, Fang Liu 0006 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2026 | Efficient Approximation of Earth Mover's Distance Based on Nearest Neighbor Search
Guangyu Meng, Ruyu Zhou, Liu Liu 0023, Peixian Liang, Fang Liu 0006, Danny Ziyi Chen, Michael T. Niemier, Xiaobo Sharon Hu |
IEEE Trans. Multim. | 5 |
| 2023 | Disclosure Risk From Homogeneity Attack in Differentially Privately Sanitized Frequency DistributionabstractDifferential privacy (DP) provides a robust model to achieve privacy guarantees for released information. We examine the protection potency of sanitized multi-dimensional frequency distributions (FDs) via DP mechanisms against homogeneity attack (HA). Adversaries can obtain the exact values on sensitive attributes of their targets through HA without having to identify them from released data. We propose measures for disclosure risk (DR) from HA and derive closed-form relations between the privacy loss parameters and DR from HA. The availability of the closed-form relations will assist practitioners in understanding the abstract concepts of DP and privacy loss parameters by putting them in the context of a concrete privacy attack and offer a perspective for choosing privacy loss parameters when employing DP mechanisms. We apply the derived mathematical relations in real data to demonstrate the assessment of DR from HA on differentially privately sanitized FDs at various privacy loss parameters. The results suggest that relations between DR from HA and privacy loss are S-shaped; the former may not disappear even when privacy loss approaches 0. Fang Liu 0006, Xingyuan Zhao |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2023 | Privacy-Preserving Travel Time Prediction With Uncertainty Using GPS Trace DataabstractThe rapid growth of GPS technology and mobile devices has led to a massive accumulation of location data, bringing considerable benefits to individuals and society. One of the major usages of such data is travel time prediction, a typical service provided by GPS navigation devices and apps. Meanwhile, the constant collection and analysis of the individual location data also pose unprecedented privacy threats. We leverage the notion of geo-indistinguishability, an extension of differential privacy to the location privacy setting, and propose a procedure for privacy-preserving travel time prediction without collecting actual individual GPS trace data. We propose new concepts to examine the impact of geo-indistinguishability-based sanitization on the usefulness of GPS traces and provide analytical and experimental utility analysis for privacy-preserving travel time prediction. We also propose new metrics to measure the adversary error in learning individual GPS traces from the collected sanitized data. Our experiment results suggest that the proposed procedure provides travel time prediction with satisfactory accuracy at reasonably small privacy costs. Fang Liu 0006, Dong Wang 0019, Zhengquan Xu |
IEEE Trans. Mob. Comput. | 1 |
| 2022 | Disclosure Risk from Homogeneity Attack in Differentially Private Release of Frequency DistributionabstractDifferential privacy (DP) provides a robust model to achieve privacy guarantees in released information. We examine the robustness of the protection against homogeneity attack (HA) in multi-dimensional frequency distributions sanitized via DP randomization mechanisms. We propose measures for disclosure risk from HA and derive closed-form relationships between privacy loss parameters in DP and disclosure risk from HA. We also provide a lower bound to the disclosure risk on a sensitive attribute when all the cells formed by quasi-identifiers are homogeneous for the sensitive attribute. The availability of the closed-form relationships helps understand the abstract concepts of DP and privacy loss parameters by putting them in the context of a concrete privacy attack and offers a perspective for choosing privacy loss parameters when employing DP mechanisms to release information in practice. We apply the closed-form mathematical relationships on real-life datasets to assess disclosure risk due to HA in differentially private sanitized frequency distributions at various privacy loss parameters. Fang Liu 0006, Xingyuan Zhao |
CODASPY | 1 |
| 2022 | A New Bound for Privacy Loss from Bayesian Posterior SamplingabstractDifferential privacy (DP) is a state-of-the-art concept that formalizes privacy guarantees. We derive a new bound for the privacy loss from releasing Bayesian posterior samples in the setting of DP. The new bound is tighter than the existing bounds for common Bayesian models and is also consistent with the likelihood principle. We apply the privacy loss quantified by the new bound to release differentially private synthetic data from Bayesian models in several experiments and show the improved utility of the synthetic data compared to those generated from explicitly designed randomization mechanisms that privatize posterior distributions. Xingyuan Zhao, Fang Liu 0006 |
CODASPY | 2 |
| 2022 | Adaptive Noisy Data Augmentation for Regularized Estimation and Inference of Generalized Linear ModelsabstractWe propose the AdaPtive Noise Augmentation (PANDA) procedure to regularize the estimation and inference of generalized linear models (GLMs). PANDA iteratively optimizes the objective function given noise-augmented data to obtain regularized model estimates. The augmented noises are designed to achieve various regularization effects, including$l_{0}$, bridge (lasso and ridge included), elastic net, adaptive lasso, and SCAD, as well as group lasso and fused ridge. We examine the tail bound of the noise-augmented loss function and establish the almost sure convergence of the noise-augmented loss function and its minimizer to the expected penalized loss function and its minimizer, respectively. PANDA exhibits ensemble learning behaviors that help further decrease the generalization error of trained GLMs. We also derive the asymptotic distributions for PANDA-regularized parameters, based on which, inferences can be obtained for GLM parameters. Computationally, PANDA is easy to code and can leverage existing software for implementing unregularized GLMs. We demonstrate the superior or similar performance of PANDA against existing approaches that offer the same type of regularizers in simulated and real-life data. We show that inferences through PANDA achieve nominal or near-nominal coverage and are far more efficient compared to a popular existing post-selection procedure. Fang Liu 0006 |
COMPSAC | 2 |
| 2022 | Efficient Reinforcement Learning from Demonstration Using Local Ensemble and Reparameterization with Split and Merge of Expert PoliciesabstractThe current work on reinforcement learning (RL) from demonstrations often assumes the demonstrations are sam-ples from an optimal policy, an unrealistic assumption in practice. When demonstrations are generated by sub-optimal policies or have sparse state-action pairs, policy learned from sub-optimal demonstrations may mislead an agent with incorrect or non-local action decisions. We propose a new method called Local Ensemble and Reparameterization with Split and Merge of expert policies (LEARN-SAM) to improve efficiency and make better use of the sub-optimal demonstrations. First, LEARN-SAM employs a new concept, the A-function, based on a discrepancy measure between the current state to demonstrated states to “localize” the weights of the expert policies during learning. Second, LEARN-SAM employs a split-and-merge (SAM) mechanism by separating the helpful parts in each expert demonstration and regrouping them into new expert policies to use the demonstrations selectively. Both the A-function and SAM mechanism help boost the learning speed. Theoretically, we prove the invariant property of reparameterized policy before and after the SAM mechanism, providing theoretical guarantees for the convergence of the employed policy gradient method. We demonstrate the superiority of the LEARN-SAM method and its robustness with varying demonstration quality and sparsity in six experiments on complex continuous control problems of low to high dimensions, compared to existing methods on RL from demonstration. Fang Liu 0006 |
COMPSAC | 2 |
| 2021 | Adaptive Noisy Data Augmentation for Regularized Construction of Undirected Graphical ModelsabstractWe develop the AdaPtive Noise Augmentation (PANDA) technique to regularize the estimation of undirected graphical models. PANDA iteratively optimizes the objective function given adaptively augmented data to achieve regularization on model parameters. The augmented noisy data is designed to deliver various regularization effects on single graph estimation as well as simultaneous construction of multiple graphs, including but not limited to$l_{\gamma}$for$\gamma\in[0,2]$, elastic net, SCAD, group lasso, and adaptive lasso in single graph estimations; and the joint group lasso and the joint fused ridge regularizations for multiple graph estimation. PANDA can be seamlessly implemented in practice in software that implements generalized linear models and users do not have to employ ad-hoc optimizers to minimize regularized loss functions for graph construction. We show the non-inferiority of PANDA in various types of graph estimation in simulated data, benchmarked against some common graph estimation methods. We also apply PANDA to an autism spectrum disorder dataset to construct a graph with mixed node types and to a lung cancer microarray data set to simultaneously construct four protein networks, demonstrating the effectiveness of PANDA in constructing practically interpretable and meaningful graphical models. Fang Liu 0006 |
DSAA | 2 |
| 2021 | Continuous-Time Markov-Switching GARCH Process with Robust State Path Identification and Volatility Estimation
Fang Liu 0006 |
ECML/PKDD (1) | 2 |
| 2020 | Differentially Private Generation of Social Networks via Exponential Random Graph ModelsabstractMany social networks contain sensitive relational information. One approach to protect the sensitive relational information while offering flexibility for social network research and analysis is to release synthetic social networks at a pre-specified privacy risk level, given the original observed network. We propose the DP-ERGM procedure that synthesizes networks that satisfy the differential privacy (DP) via the exponential random graph model (EGRM). We apply DP-ERGM to a college student friendship network and compare its original network information preservation in the generated private networks with two other approaches: differentially private DyadWise Randomized Response (DWRR) and Sanitization of the Conditional probability of Edge given Attribute classes (SCEA). The results suggest that DP-EGRM preserves the original information significantly better than DWRR and SCEA in both network statistics and inferences from ERGMs and latent space models. In addition, DP-ERGM satisfies the node DP, a stronger notion of privacy than the edge DP that DWRR and SCEA satisfy. Fang Liu 0006, Evercita C. Eugenio, Ick-Hoon Jin, Claire McKay Bowen |
COMPSAC | 1 |
| 2020 | Utility Analysis of Horizontally Merged Multi-Party Synthetic Data with Differential PrivacyabstractA large amount of data is often needed to train machine learning algorithms with confidence. One way to achieve the necessary data volume is to share and combine data from multiple parties. On the other hand, how to protect sensitive personal information during data sharing is always a challenge. We focus on data sharing when parties have overlapping attributes but non-overlapping individuals. One approach to achieve privacy protection is through sharing differentially private synthetic data. Each party generates synthetic data at its own preferred privacy budget, which are then released and horizontally merged across the parties. The total privacy cost for this approach is capped at the maximum individual budget employed by a party. We derive the mean squared error bounds for the parameter estimation in common regression analysis based on the merged sanitized data across parties. We identify through theoretical analysis the conditions under which the utility of sharing and merging sanitized data outweighs the perturbation introduced for satisfying differential privacy and surpasses that based on individual party data. The experiments suggest that sanitized HOMM data obtained at a practically reasonable small privacy cost can lead to smaller prediction and estimation errors than individual parties, demonstrating benefits of data sharing while protecting privacy. Bingyue Su, Fang Liu 0006 |
ISNCC | 2 |
| 2020 | Adaptive Gaussian Noise Injection Regularization for Neural Networks
Fang Liu 0006 |
ISNN | 2 |
| 2019 | Generalized Gaussian Mechanism for Differential PrivacyabstractAssessment of disclosure risk is of paramount importance in data privacy research and applications. The concept of differential privacy (DP) formalizes privacy in probabilistic terms and provides a robust concept for privacy protection. Practical applications of DP involve development of DP mechanisms to release data at a pre-specified privacy budget. In this paper, we generalize the widely used Laplace mechanism to the family of generalized Gaussian (GG) mechanism based on the$l_p$global sensitivity of statistical queries. We explore the theoretical requirement for the GG mechanism to reach DP at prespecified privacy parameters, and investigate the connections and differences between the GG mechanism and the Exponential mechanism based on the GG distribution. We also present a lower bound on the scale parameter of the Gaussian mechanism of$(\epsilon, \delta)$-probabilistic DP as a special case of the GG mechanism, and compare the utility of sanitized results in the tail probability and dispersion between the Gaussian and Laplace mechanisms. Lastly, we apply the GG mechanism in three experiments and compare the accuracy of sanitized results in the$l_1$distance and Kullback-Leibler divergence, and examine the prediction power of a SVM classifier constructed with the sanitized data relative to the original results. Fang Liu 0006 |
IEEE Trans. Knowl. Data Eng. | 1 |