VLDB 2026 Research / reviewers in the wild / expert
Zefeng Cai
dblp:323/9645
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2024
0000-0001-7585-034XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Recommender systems · 65% Information retrieval · 35% | |
| Artificial intelligence
3 papers |
Generative modeling · 56% Language models and text generation · 24% Question answering and dialogue systems · 21% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Recommender systems
explainable recommendation |
1.2 | 2 | 2023 | Disentangled CVAEs with Contrastive Learning for Explainable Recommendation · AAAI 2023 PEVAE: A Hierarchical VAE for Personalized Explainable Recommendation · SIGIR 2022 |
Machine learning › Generative modeling
variational autoencoder |
0.8 | 2 | 2023 | PCVAE: Generating Prior Context for Dialogue Response Generation · IJCAI 2022 Disentangled CVAEs with Contrastive Learning for Explainable Recommendation · AAAI 2023 |
Natural language and speech › Language models and text generation
prompt tuning |
0.7 | 1 | 2023 | HypeR: Multitask Hyper-Prompted Training Enables Large-Scale Retrieval Generalization · ICLR 2023 |
Information retrieval › retrieval models › neural retrieval
dense retrieval |
0.7 | 1 | 2023 | HypeR: Multitask Hyper-Prompted Training Enables Large-Scale Retrieval Generalization · ICLR 2023 |
Recommender systems › explainable recommendation
explanation generation |
0.7 | 1 | 2023 | Disentangled CVAEs with Contrastive Learning for Explainable Recommendation · AAAI 2023 |
Information retrieval › retrieval models
multi-task retrieval |
0.7 | 1 | 2023 | HypeR: Multitask Hyper-Prompted Training Enables Large-Scale Retrieval Generalization · ICLR 2023 |
Natural language and speech › Question answering and dialogue systems › dialogue generation
dialogue response generation |
0.6 | 1 | 2022 | PCVAE: Generating Prior Context for Dialogue Response Generation · IJCAI 2022 |
Machine learning › Generative modeling › variational autoencoder
hierarchical VAE |
0.6 | 1 | 2022 | PCVAE: Generating Prior Context for Dialogue Response Generation · IJCAI 2022 |
Recommender systems › generative recommendation
variational autoencoder-based recommendation |
0.6 | 1 | 2022 | PEVAE: A Hierarchical VAE for Personalized Explainable Recommendation · SIGIR 2022 |
Machine learning › Generative modeling › variational autoencoder
conditional variational autoencoder |
0.2 | 1 | 2023 | Disentangled CVAEs with Contrastive Learning for Explainable Recommendation · AAAI 2023 |
Methods — techniques the papers use, named apart from their topics
disentangled representation learning · 1.3contrastive learning · 1.3variational autoencoder · 0.6mutual information maximization · 0.6conditional variational autoencoder · 0.6autoregressive compatible arrangement · 0.6active codeword transport · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | TP-Link: Fine-grained Pre-Training for Text-to-SQL Parsing with Linking InformationabstractIn this paper, we introduce an innovative pre-training framework TP-Link, which aims to improve context-dependent Text-to-SQL Parsing by leveraging Linking information. This enhancement is achieved through better representation of both natural language utterances and the database schema, ultimately facilitating more effective text-to-SQL conversations. We present two novel pre-training objectives: (i) utterance linking prediction (ULP) task that models intricate syntactic relationships among natural language utterances in context-dependent text-to-SQL scenarios, and (ii) schema linking prediction (SLP) task that focuses on capturing fine-grained schema linking relationships between the utterances and the database schema. Extensive experiments demonstrate that our proposed TP-Link achieves state-of-the-art performance on two leading downstream benchmarks (i.e., SParC and CoSQL). Shujie Li 0001, Zefeng Cai, Yunshui Li, Chengming Li 0004, Xiping Hu, Ruifeng Xu 0001, Min Yang 0007 |
LREC/COLING | 3 |
| 2023 | Disentangled CVAEs with Contrastive Learning for Explainable RecommendationabstractModern recommender systems are increasingly expected to provide informative explanations that enable users to understand the reason for particular recommendations. However, previous methods struggle to interpret the input IDs of user--item pairs in real-world datasets, failing to extract adequate characteristics for controllable generation. To address this issue, we propose disentangled conditional variational autoencoders (CVAEs) for explainable recommendation, which leverage disentangled latent preference factors and guide the explanation generation with the refined condition of CVAEs via a self-regularization contrastive learning loss. Extensive experiments demonstrate that our method generates high-quality explanations and achieves new state-of-the-art results in diverse domains. Zefeng Cai, Gerard de Melo, Zhu Cao, Liang He 0001 |
AAAI | 2 |
| 2023 | HypeR: Multitask Hyper-Prompted Training Enables Large-Scale Retrieval Generalization
Zefeng Cai, Chongyang Tao, Tao Shen 0001, Can Xu 0002, Xiubo Geng, Xin Lin 0001, Liang He 0001, Daxin Jiang |
ICLR | 1 |
| 2022 | PCVAE: Generating Prior Context for Dialogue Response GenerationabstractConditional Variational AutoEncoder (CVAE) is promising for modeling one-to-many relationships in dialogue generation, as it can naturally generate many responses from a given context. However, the conventional used continual latent variables in CVAE are more likely to generate generic rather than distinct and specific responses. To resolve this problem, we introduce a novel discrete variable called prior context which enables the generation of favorable responses. Specifically, we present Prior Context VAE (PCVAE), a hierarchical VAE that learns prior context from data automatically for dialogue generation. Meanwhile, we design Active Codeword Transport (ACT) to help the model actively discover potential prior context. Moreover, we propose Autoregressive Compatible Arrangement (ACA) that enables modeling prior context in autoregressive style, which is crucial for selecting appropriate prior context according to a given context. Extensive experiments demonstrate that PCVAE can generate distinct responses and significantly outperforms strong baselines. Zefeng Cai, Zerui Cai |
IJCAI | 1 |
| 2022 | PEVAE: A Hierarchical VAE for Personalized Explainable RecommendationabstractVariational autoencoders (VAEs) have been widely applied in recommendations. One reason is that their amortized inferences are beneficial for overcoming the data sparsity. However, in explainable recommendation that generates natural language explanations, they are still rarely explored. Thus, we aim to extend VAE to explainable recommendation. In this task, we find that VAE can generate acceptable explanations for users with few relevant training samples, however, it tends to generate less personalized explanations for users with relatively sufficient samples than autoencoders (AEs). We conjecture that information shared by different users in VAE disturbs the information for a specific user. To deal with this problem, we present PErsonalized VAE (PEVAE) that generates personalized natural language explanations for explainable recommendation. Moreover, we propose two novel mechanisms to aid our model in generating more personalized explanations, including 1) Self-Adaption Fusion (SAF) manipulates the latent space in a self-adaption manner for controlling the influence of shared information. In this way, our model can enjoy the advantage of overcoming the sparsity of data while generating more personalized explanations for a user with relatively sufficient training samples. 2) DEpendence Maximization (DEM) strengthens dependence between recommendations and explanations by maximizing the mutual information. It makes the explanation more specific to the input user-item pair and thus improves the personalization of the generated explanations. Extensive experiments show PEVAE can generate more personalized explanations and further analyses demonstrate the practical effect of our proposed methods. Zefeng Cai, Zerui Cai |
SIGIR | 1 |