VLDB 2026 Research / reviewers in the wild / expert
Zitai Chen
dblp:239/4477
· DBLP profile ↗
7ranked-venue papers
6as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLMs Leak Training Data Beyond Verbatim Memorization: Extraction via Membership DecodingabstractExtracting training data from large language models (LLMs) is a serious privacy breach that exposes (potentially private) data without data owners' consent. Existing extractions follow the generation-then-audit paradigm, where the greedy decoding method in generation limits the extraction scope and only verbatim memorized data is under audits. A majority of partially memorized member data (around 90%) remains unexplored, of which LLMs could memorize almost all tokens but fail to rank the training token to top-1 at some positions. To measure the degree of such a partial memorization, we introduce a new notion of memorization, k/n-correction, by the number of non-top-1 tokens k in a suffix of length n. This notion quantifies the memorization in a fine-grained manner. Experiments show that models memorize more with smaller average k values as the model size increases. To extract the partial memorization, we propose a new decoding method, named Membership Decoding, by introducing membership information in the generation process. The Membership Decoding method is a plug-and-play replacement for standard decoding that requires only black-box token probabilities. We formalize the data extraction problem as a next member token prediction problem. Accordingly, we propose a new token-level membership inference method by leveraging likelihood from reference models, shifting the generation from the original token distribution to the member token distribution. Extensive experiments show that partial memorization is much more prevalent than verbatim memorization, and membership decoding can extract previously unextractible partially-memorized sequences, succeeding in correcting sequences with up to k=2 non-member tokens in the suffix of length n=10. The proposed attack demonstrates the potential privacy risks in partial memorization. Zitai Chen, Reza Shokri |
Proc. Priv. Enhancing Technol. | 1 |
| 2022 | MetaEmu: An Architecture Agnostic Rehosting Framework for Automotive FirmwareabstractIn this paper we present MetaEmu, an architecture-agnostic framework geared towards rehosting and security analysis of automotive firmware. MetaEmu improves over existing rehosting environments in two ways: Firstly, it solves the hitherto open-problem of a lack of generic Virtual Execution Environments (VXEs) by synthesizing processor simulators from Ghidra's language definitions. Secondly, MetaEmu can rehost and analyze multiple targets, each of different architecture, simultaneously, and share analysis facts between each target's analysis environment, a technique we call inter-device analysis. Zitai Chen, Sam L. Thomas, Flavio D. Garcia |
CCS | 1 |
| 2021 | VoltPillager: Hardware-based fault injection attacks against Intel SGX Enclaves using the SVID voltage scaling interface
Zitai Chen, Georgios Vasilakis, Kit Murdock, Edward Dean, David F. Oswald, Flavio D. Garcia |
USENIX Security Symposium | 1 |
| 2021 | Auto-weighted robust low-rank tensor completion via tensor-train
Chuan Chen 0001, Zhebin Wu, Zitai Chen, Zibin Zheng, Xiongjun Zhang |
Inf. Sci. | 3 |
| 2019 | Tensor Decomposition for Multilayer Networks ClusteringabstractClustering on multilayer networks has been shown to be a promising approach to enhance the accuracy. Various multilayer networks clustering algorithms assume all networks derive from a latent clustering structure, and jointly learn the compatible and complementary information from different networks to excavate one shared underlying structure. However, such an assumption is in conflict with many emerging real-life applications due to the existence of noisy/irrelevant networks. To address this issue, we propose Centroid-based Multilayer Network Clustering (CMNC), a novel approach which can divide irrelevant relationships into different network groups and uncover the cluster structure in each group simultaneously. The multilayer networks is represented within a unified tensor framework for simultaneously capturing multiple types of relationships between a set of entities. By imposing the rank-(Lr,Lr,1) block term decomposition with nonnegativity, we are able to have well interpretations on the multiple clustering results based on graph cut theory. Numerically, we transform this tensor decomposition problem to an unconstrained optimization, thus can solve it efficiently under the nonlinear least squares (NLS) framework. Extensive experimental results on synthetic and real-world datasets show the effectiveness and robustness of our method against noise and irrelevant data. Zitai Chen, Chuan Chen 0001, Zibin Zheng |
AAAI | 1 |
| 2019 | SINE: Side Information Network Embedding
Zitai Chen, Tongzhao Cai, Chuan Chen 0001, Zibin Zheng, Guohui Ling |
DASFAA (1) | 1 |
| 2019 | Variational Graph Embedding and Clustering with Laplacian EigenmapsabstractAs a fundamental machine learning problem, graph clustering has facilitated various real-world applications, and tremendous efforts had been devoted to it in the past few decades. However, most of the existing methods like spectral clustering suffer from the sparsity, scalability, robustness and handling high dimensional raw information in clustering. To address this issue, we propose a deep probabilistic model, called Variational Graph Embedding and Clustering with Laplacian Eigenmaps (VGECLE), which learns node embeddings and assigns node clusters simultaneously. It represents each node as a Gaussian distribution to disentangle the true embedding position and the uncertainty from the graph. With a Mixture of Gaussian (MoG) prior, VGECLE is capable of learning an interpretable clustering by the variational inference and generative process. In order to learn the pairwise relationships better, we propose a Teacher-Student mechanism encouraging node to learn a better Gaussian from its instant neighbors in the stochastic gradient descent (SGD) training fashion. By optimizing the graph embedding and the graph clustering problem as a whole, our model can fully take the advantages in their correlation. To our best knowledge, we are the first to tackle graph clustering in a deep probabilistic viewpoint. We perform extensive experiments on both synthetic and real-world networks to corroborate the effectiveness and efficiency of the proposed framework. Zitai Chen, Chuan Chen 0001, Zong Zhang, Zibin Zheng, Qingsong Zou |
IJCAI | 1 |