VLDB 2026 Research / reviewers in the wild / expert
Chen Chen 0056
dblp:65/4423-56
· DBLP profile ↗
6ranked-venue papers
1as first author
2since 2021 · last 2023
0000-0003-2104-534XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorArtificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Preserving Privacy for Distributed Genome-Wide Analysis Against Identity Tracing AttacksabstractGenome-wide analysis has demonstrated both health and social benefits. However, large scale sharing of such data may reveal sensitive information about individuals. One of the emerging challenges is identity tracing attack that exploits correlations among genomic data to reveal the identity of DNA samples. In this paper, we first demonstrate that the adversary can narrow down the sample's identity by detecting his/her genetic relatives and quantify such privacy threat by employing a Shannon entropy-based measurement. For example, we exemplify that when the dataset size reaches 30% of the population, for any target from that population, the uncertainty of the target's identity is reduced to merely 2.3 bits of entropy (i.e., the identity is pinned down within 5 people). Direct application of existing approaches such as differential privacy (DP), secure multiparty computation (MPC) and homomorphic encryption (HE) may not be applicable to this challenge in genome-wide analysis because of the compromise on utility (i.e., accuracy or efficiency). Towards addressing this challenge, this paper proposes a framework named$\upsilon$Fragto facilitate privacy-preserving data sharing and computation in genome-wide analysis.$\upsilon$Fragmitigates privacy risks by using a vertical fragmentation to disrupt the genetic architecture on which the adversary relies for identity tracing without sacrificing the capability of genome-wide analysis. We theoretically prove that it preserves the correctness of the primitive functionalities and algorithms ranging from basic summary statistics to advanced neural networks. Our experiments demonstrate that$\upsilon$Fragoutperforms secure multiparty computation (MPC) and homomorphic encryption (HE) protocols, with a speedup of more than 221x for training neural networks, and also traditional non-private algorithms and a state-of-the-art noise-based differential privacy (DP) solution in most settings. Yanjun Zhang 0002, Guangdong Bai, Xue Li 0001, Surya Nepal, Marthie Grobler, Chen Chen 0056, Ryan Kok Leong Ko |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2021 | Privacy-Preserving Gradient Descent for Distributed Genome-Wide Analysis
Yanjun Zhang 0002, Guangdong Bai, Xue Li 0001, Caitlin Curtis, Chen Chen 0056, Ryan Kok Leong Ko |
ESORICS (2) | 5 |
| 2020 | PrivColl: Practical Privacy-Preserving Collaborative Machine Learning
Yanjun Zhang 0002, Guangdong Bai, Xue Li 0001, Caitlin Curtis, Chen Chen 0056, Ryan Kok Leong Ko |
ESORICS (1) | 5 |
| 2019 | Enabling Privacy-Preserving Sharing of Genomic Data for GWASs in Decentralized NetworksabstractThe human genome can reveal sensitive information and is potentially re-identifiable, which raises privacy and security concerns about sharing such data on wide scales. In this work, we propose a preventive approach for privacy-preserving sharing of genomic data in decentralized networks for Genome-wide association studies (GWASs), which have been widely used in discovering the association between genotypes and phenotypes. The key components of this work are: a decentralized secure network, with a privacy- preserving sharing protocol, and a gene fragmentation framework that is trainable in an end-to-end manner. Our experiments on real datasets show the effectiveness of our privacy-preserving approaches as well as significant improvements in efficiency when compared with recent, related algorithms. Yanjun Zhang 0002, Xin Zhao 0013, Xue Li 0001, Mingyang Zhong, Caitlin Curtis, Chen Chen 0056 |
WSDM | 6 |
| 2013 | TeRec: A Temporal Recommender System Over Tweet StreamabstractAs social media further integrates into our daily lives, people are increasingly immersed in real-time social streams via services such as Twitter and Weibo. One important observation in these online social platforms is that users' interests and the popularity of topics shift very fast, which poses great challenges on existing recommender systems to provide the right topics at the right time. In this paper, we extend the online ranking technique and propose a temporal recommender system - TeRec. In TeRec, when posting tweets, users can get recommendations of topics (hashtags) according to their real-time interests, they can also generate fast feedbacks according to the recommendations. TeRec provides the browser-based client interface which enables the users to access the real time topic recommendations, and the server side processes and stores the real-time stream data. The experimental study demonstrates the superiority of TeRec in terms of temporal recommendation accuracy. Chen Chen 0056, Hongzhi Yin, Bin Cui 0001 |
Proc. VLDB Endow. | 1 |
| 2012 | Challenging the Long Tail RecommendationabstractThe success of "infinite-inventory" retailers such as Amazon.com and Netflix has been largely attributed to a "long tail" phenomenon. Although the majority of their inventory is not in high demand, these niche products, unavailable at limited-inventory competitors, generate a significant fraction of total revenue in aggregate. In addition, tail product availability can boost head sales by offering consumers the convenience of "one-stop shopping" for both their mainstream and niche tastes. However, most of existing recommender systems, especially collaborative filter based methods, can not recommend tail products due to the data sparsity issue. It has been widely acknowledged that to recommend popular products is easier yet more trivial while to recommend long tail products adds more novelty yet it is also a more challenging task. In this paper, we propose a novel suite of graph-based algorithms for the long tail recommendation. We first represent user-item information with undirected edge-weighted graph and investigate the theoretical foundation of applying Hitting Time algorithm for long tail item recommendation. To improve recommendation diversity and accuracy, we extend Hitting Time and propose efficient Absorbing Time algorithm to help users find their favorite long tail items. Finally, we refine the Absorbing Time algorithm and propose two entropy-biased Absorbing Cost algorithms to distinguish the variation on different user-item rating pairs, which further enhances the effectiveness of long tail recommendation. Empirical experiments on two real life datasets show that our proposed algorithms are effective to recommend long tail items and outperform state-of-the-art recommendation techniques. Hongzhi Yin, Bin Cui 0001, Jing Li 0021, Chen Chen 0056 |
Proc. VLDB Endow. | 5 |