VLDB 2026 Research / reviewers in the wild / expert
Zerui Cai
dblp:297/8977
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | PCVAE: Generating Prior Context for Dialogue Response GenerationabstractConditional Variational AutoEncoder (CVAE) is promising for modeling one-to-many relationships in dialogue generation, as it can naturally generate many responses from a given context. However, the conventional used continual latent variables in CVAE are more likely to generate generic rather than distinct and specific responses. To resolve this problem, we introduce a novel discrete variable called prior context which enables the generation of favorable responses. Specifically, we present Prior Context VAE (PCVAE), a hierarchical VAE that learns prior context from data automatically for dialogue generation. Meanwhile, we design Active Codeword Transport (ACT) to help the model actively discover potential prior context. Moreover, we propose Autoregressive Compatible Arrangement (ACA) that enables modeling prior context in autoregressive style, which is crucial for selecting appropriate prior context according to a given context. Extensive experiments demonstrate that PCVAE can generate distinct responses and significantly outperforms strong baselines. Zefeng Cai, Zerui Cai |
IJCAI | 2 |
| 2022 | PEVAE: A Hierarchical VAE for Personalized Explainable RecommendationabstractVariational autoencoders (VAEs) have been widely applied in recommendations. One reason is that their amortized inferences are beneficial for overcoming the data sparsity. However, in explainable recommendation that generates natural language explanations, they are still rarely explored. Thus, we aim to extend VAE to explainable recommendation. In this task, we find that VAE can generate acceptable explanations for users with few relevant training samples, however, it tends to generate less personalized explanations for users with relatively sufficient samples than autoencoders (AEs). We conjecture that information shared by different users in VAE disturbs the information for a specific user. To deal with this problem, we present PErsonalized VAE (PEVAE) that generates personalized natural language explanations for explainable recommendation. Moreover, we propose two novel mechanisms to aid our model in generating more personalized explanations, including 1) Self-Adaption Fusion (SAF) manipulates the latent space in a self-adaption manner for controlling the influence of shared information. In this way, our model can enjoy the advantage of overcoming the sparsity of data while generating more personalized explanations for a user with relatively sufficient training samples. 2) DEpendence Maximization (DEM) strengthens dependence between recommendations and explanations by maximizing the mutual information. It makes the explanation more specific to the input user-item pair and thus improves the personalization of the generated explanations. Extensive experiments show PEVAE can generate more personalized explanations and further analyses demonstrate the practical effect of our proposed methods. Zefeng Cai, Zerui Cai |
SIGIR | 2 |
| 2021 | SMedBERT: A Knowledge-Enhanced Pre-trained Language Model with Structured Semantics for Medical Text MiningabstractTaolin Zhang, Zerui Cai, Chengyu Wang, Minghui Qiu, Bite Yang, Xiaofeng He. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Taolin Zhang 0001, Zerui Cai, Chengyu Wang 0001, Minghui Qiu, Bite Yang |
ACL/IJCNLP (1) | 2 |
| 2021 | HORNET: Enriching Pre-trained Language Representations with Heterogeneous Knowledge SourcesabstractKnowledge-Enhanced Pre-trained Language Models (KEPLMs) improve the language understanding abilities of deep language models by leveraging the rich semantic knowledge from knowledge graphs, other than plain pre-training texts. However, previous efforts mostly use homogeneous knowledge (especially structured relation triples in knowledge graphs) to enhance the context-aware representations of entity mentions, whose performance may be limited by the coverage of knowledge graphs. Also, it is unclear whether these KEPLMs truly understand the injected semantic knowledge due to the "black-box'' training mechanism. In this paper, we propose a novel KEPLM named HORNET, which integrates Heterogeneous knowledge from various structured and unstructured sources into the Roberta NETwork and hence takes full advantage of both linguistic and factual knowledge simultaneously. Specifically, we design a hybrid attention heterogeneous graph convolution network (HaHGCN) to learn heterogeneous knowledge representations based on the structured relation triplets from knowledge graphs and the unstructured entity description texts. Meanwhile, we propose the explicit dual knowledge understanding tasks to help induce a more effective infusion of the heterogeneous knowledge, promoting our model for learning the complicated mappings from the knowledge graph embedding space to the deep context-aware embedding space and vice versa. Experiments show that our HORNET model outperforms various KEPLM baselines on knowledge-aware tasks including knowledge probing, entity typing and relation extraction. Our model also achieves substantial improvement over several GLUE benchmark datasets, compared to other KEPLMs. Taolin Zhang 0001, Zerui Cai, Chengyu Wang 0001, Peng Li 0056, Yang Li 0218, Minghui Qiu, Chengguang Tang, Jun Huang 0007 |
CIKM | 2 |
| 2021 | Generating Explanations for Recommendation Systems via Injective VAEabstractGenerating explanations for recommendation systems is essential for improving its transparency since informative explanations such as generated reviews can help users comprehend the reason for receiving a specified recommendation. The generated reviews should be specific for the given user, item, and rating, however, recent works only focus on designing more and more powerful decoder, merely treating this task as a plain natural language generation process. We argue that there may exist the risk that the powerful decoder neglects the input embeddings and suffers from the biases that exist in data. In this paper, we propose a novel Injective Variational Autoencoders (InVAE) for generating high-quality reviews. Specifically, we employ a Collaborative Kullback-Leibler divergences (CKL) mechanism to building a better latent space that captures meaningful information. Base on this, the Spectral Regularization on Flow-based transformation (SRF) method is designed to backward transfer the priorities of generated latent variables to the input embeddings. Therefore, our method can construct more informative input embeddings and provides more specific explanations for different inputs. Extensive empirical experiments demonstrate that our model can construct much more meaningful feature embeddings and generate personalized reviews in high quality. Zerui Cai |
ICDM | 1 |