VLDB 2026 Research / reviewers in the wild / expert
Jingwei Yi
dblp:290/2312
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0001-2786-6395ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Measuring Human Contribution in AI-Assisted Content GenerationabstractYueqi Xie, Tao Qi, Jingwei Yi, Xiyuan Yang, Ryan Whalen, Junming Huang, Qian Ding, Yu Xie, Xing Xie, Fangzhao Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yueqi Xie, Tao Qi 0001, Jingwei Yi, Xiyuan Yang, Ryan Whalen, Junming Huang 0001, Xing Xie 0001, Fangzhao Wu |
ACL (1) | 3 |
| 2025 | Benchmarking and Defending against Indirect Prompt Injection Attacks on Large Language ModelsabstractThe integration of large language models (LLMs) with external content has enabled applications such as Microsoft Copilot but also introduced vulnerabilities to indirect prompt injection attacks. In these attacks, malicious instructions embedded within external content can manipulate LLM outputs, causing deviations from user expectations. To address this critical yet under-explored issue, we introduce the first benchmark for bindirect prompt injection attacks, named BIPIA, to assess the risk of such vulnerabilities. Using BIPIA, we evaluate existing LLMs and find them universally vulnerable. Our analysis identifies two key factors contributing to their success: LLMs' inability to distinguish between informational context and actionable instructions, and their lack of awareness in avoiding the execution of instructions within external content. Based on these findings, we propose two novel defense mechanisms -- boundary awareness and explicit reminder -- to address these vulnerabilities in both black-box and white-box settings. Extensive experiments demonstrate that our black-box defense provides substantial mitigation, while our white-box defense reduces the attack success rate to near-zero levels, all while preserving the output quality of LLMs. We hope this work inspires further research into securing LLM applications and fostering their safe and reliable use. Our code is available at https://github.com/microsoft/BIPIA. Jingwei Yi, Yueqi Xie, Bin B. Zhu, Emre Kiciman, Guangzhong Sun, Xing Xie 0001, Fangzhao Wu |
KDD (1) | 1 |
| 2023 | Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor WatermarkabstractWenjun Peng, Jingwei Yi, Fangzhao Wu, Shangxi Wu, Bin Bin Zhu, Lingjuan Lyu, Binxing Jiao, Tong Xu, Guangzhong Sun, Xing Xie. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Wenjun Peng 0001, Jingwei Yi, Fangzhao Wu, Shangxi Wu, Bin B. Zhu, Lingjuan Lyu, Binxing Jiao, Tong Xu 0001, Guangzhong Sun, Xing Xie 0001 |
ACL (1) | 2 |
| 2023 | Non-IID always Bad? Semi-Supervised Heterogeneous Federated Learning with Local Knowledge EnhancementabstractFederated learning (FL) is important for privacy-preserving services by training models without collecting raw user data. Most FL algorithms assume all data is annotated, which is impractical due to the high cost of labeling data in real applications. To alleviate the reliance on labeled data, semi-supervised federated learning (SSFL) has been proposed to utilize unlabeled data on clients to improve model performance. However, most existing methods either have privacy issues which share models trained on other clients, or generate pseudo-labels for unlabeled local datasets with the global model, which is usually biased towards the global data distribution. The latter may lead to sub-optimal accuracy of pseudo-labels, due to the gap between the local data distribution and the global model, especially in non-IID settings. In this paper, we propose a semi-supervised heterogeneous federated learning method with local knowledge enhancement, called FedLoKe, which aims to train an accurate global model from both labeled and unlabeled local data with non-IID distributions. Specifically, in FedLoKe, the server maintains a global model to capture global data distribution, and each client learns a local model to capture local data distribution. Since the distribution captured by the local model is aligned with the local data distribution, we utilize it to generate high-accuracy pseudo-labels of the unlabeled dataset for global model training. To prevent the local model from severely overfitting the small number of local labeled data, we further use the exponential moving average and apply the global model to generate pseudo-labels for local modeling training. Experiments on four datasets show the effectiveness of FedLoKe. Our code is available at: https://github.com/zcfinal/FedLoKe. Chao Zhang 0096, Fangzhao Wu, Jingwei Yi, Derong Xu, Yang Yu 0038, Jindong Wang 0001, Yidong Wang 0003, Tong Xu 0001, Xing Xie 0001, Enhong Chen |
CIKM | 3 |
| 2023 | UA-FedRec: Untargeted Attack on Federated News RecommendationabstractNews recommendation is essential for personalized news distribution. Federated news recommendation, which enables collaborative model learning from multiple clients without sharing their raw data, is a promising approach for preserving users' privacy. However, the security of federated news recommendation is still unclear. In this paper, we study this problem by proposing an untargeted attack on federated news recommendation called UA-FedRec. By exploiting the prior knowledge of news recommendation and federated learning, UA-FedRec can effectively degrade the model performance with a small percentage of malicious clients. First, the effectiveness of news recommendation highly depends on user modeling and news modeling. We design a news similarity perturbation method to make representations of similar news farther and those of dissimilar news closer to interrupt news modeling, and propose a user model perturbation method to make malicious user updates in opposite directions of benign updates to interrupt user modeling. Second, updates from different clients are typically aggregated with a weighted average based on their sample sizes. We propose a quantity perturbation method to enlarge sample sizes of malicious clients in a reasonable range to amplify the impact of malicious updates. Extensive experiments on two real-world datasets show that UA-FedRec can effectively degrade the accuracy of existing federated news recommendation methods, even when defense is applied. Our study reveals a critical security issue in existing federated news recommendation systems and calls for research efforts to address the issue. Our code is available at https://github.com/yjw1029/UA-FedRec. Jingwei Yi, Fangzhao Wu, Bin B. Zhu, Jing Yao 0003, Zhulin Tao, Guangzhong Sun, Xing Xie 0001 |
KDD | 1 |
| 2022 | Effective and Efficient Query-aware Snippet Extraction for Web SearchabstractQuery-aware webpage snippet extraction is widely used in search engines to help users better understand the content of the returned webpages before clicking.Although important, it is very rarely studied.In this paper, we propose an effective query-aware webpage snippet extraction method named DeepQSE, aiming to select a few sentences which can best summarize the webpage content in the context of input query.DeepQSE first learns query-aware sentence representations for each sentence to capture the fine-grained relevance between query and sentence, and then learns document-aware query-sentence relevance representations for snippet extraction.Since the query and each sentence are jointly modeled in DeepQSE, its online inference may be slow.Thus, we further propose an efficient version of DeepQSE, named Efficient-DeepQSE, which can significantly improve the inference speed of Deep-QSE without affecting its performance.The core idea of Efficient-DeepQSE is to decompose the query-aware snippet extraction task into two stages, i.e., a coarse-grained candidate sentence selection stage where sentence representations can be cached, and a fine-grained relevance modeling stage.Experiments on two real-world datasets validate the effectiveness and efficiency of our methods. Jingwei Yi, Fangzhao Wu, Chuhan Wu, Binxing Jiao, Guangzhong Sun, Xing Xie 0001 |
EMNLP | 1 |
| 2022 | Tiny-NewsRec: Effective and Efficient PLM-based News RecommendationabstractNews recommendation is a widely adopted technique to provide personalized news feeds for the user.Recently, pre-trained language models (PLMs) have demonstrated the great capability of natural language understanding and benefited news recommendation via improving news modeling.However, most existing works simply finetune the PLM with the news recommendation task, which may suffer from the known domain shift problem between the pre-training corpus and downstream news texts.Moreover, PLMs usually contain a large volume of parameters and have high computational overhead, which imposes a great burden on low-latency online services.In this paper, we propose Tiny-NewsRec, which can improve both the effectiveness and the efficiency of PLM-based news recommendation.We first design a self-supervised domain-specific posttraining method to better adapt the general PLM to the news domain with a contrastive matching task between news titles and news bodies.We further propose a two-stage knowledge distillation method to improve the efficiency of the large PLM-based news recommendation model while maintaining its performance.Multiple teacher models originated from different time steps of our post-training procedure are used to transfer comprehensive knowledge to the student model in both its posttraining stage and finetuning stage.Extensive experiments on two real-world datasets validate the effectiveness and efficiency of our method. Yang Yu 0038, Fangzhao Wu, Chuhan Wu, Jingwei Yi, Qi Liu 0003 |
EMNLP | 4 |
| 2021 | Efficient-FedRec: Efficient Federated Learning Framework for Privacy-Preserving News RecommendationabstractNews recommendation is critical for personalized news access.Most existing news recommendation methods rely on centralized storage of users' historical news click behavior data, which may lead to privacy concerns and hazards.Federated Learning is a privacy-preserving framework for multiple clients to collaboratively train models without sharing their private data.However, the computation and communication cost of directly learning many existing news recommendation models in a federated way are unacceptable for user clients.In this paper, we propose an efficient federated learning framework for privacy-preserving news recommendation.Instead of training and communicating the whole model, we decompose the news recommendation model into a large news model maintained in the server and a light-weight user model shared on both server and clients, where news representations and user model are communicated between server and clients.More specifically, the clients request the user model and news representations from the server, and send their locally computed gradients to the server for aggregation.The server updates its global user model with the aggregated gradients, and further updates its news model to infer updated news representations.Since the local gradients may contain private information, we propose a secure aggregation method to aggregate gradients in a privacy-preserving way.Experiments on two real-world datasets show that our method can reduce the computation and communication cost on clients while keep promising model performance. Jingwei Yi, Fangzhao Wu, Chuhan Wu, Ruixuan Liu, Guangzhong Sun, Xing Xie 0001 |
EMNLP (1) | 1 |