VLDB 2026 Research / reviewers in the wild / expert
Yi Zhang 0050
dblp:64/6544-50
· DBLP profile ↗
16ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-9700-0693ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CardRewriter: Leveraging Knowledge Cards for Long-Tail Query Rewriting on Short-Video PlatformsabstractShort-video platforms have rapidly become a new generation of information retrieval systems, where users formulate queries to access desired videos. However, user queries, especially long-tail ones, often suffer from spelling errors, incomplete phrasing, and ambiguous intent, resulting in mismatches between user expectations and retrieved results. While large language models (LLMs) have shown success in long-tail query rewriting within e-commerce, they struggle on short-video platforms, where proprietary content such as short videos, live streams, micro dramas, and user social networks falls outside their training distribution. To address this challenge, we introduce CardRewriter, an LLM-based framework that incorporates domain-specific knowledge to enhance long-tail query rewriting. For each query, our method aggregates multi-source knowledge relevant to the query and summarizes it into an informative and query-relevant knowledge card. This card then guides the LLM to better capture user intent and produce more effective query rewrites. We optimize CardRewriter using a two-stage training pipeline: supervised fine-tuning followed by group relative policy optimization, with a tailored reward system balancing query relevance and retrieval effectiveness. Offline experiments show that CardRewriter substantially improves rewriting quality for queries targeting proprietary content. Online A/B testing further confirms significant gains in long-view rate (LVR) and click-through rate (CTR), along with a notable reduction in initiative query reformulation rate (IQRR). Since September 2025, CardRewriter has been deployed on Kuaishou, one of China's largest short-video platforms, serving hundreds of millions of users daily. Peiyuan Gong, Feiran Zhu, Yaqi Yin, Chenglei Dai, Wentian Bao, Jiaxin Mao, Yi Zhang 0050 |
WWW | 9 |
| 2025 | Improving Retrieval-Augmented Generation through Multi-Agent Reinforcement LearningabstractRetrieval-augmented generation (RAG) is widely utilized to incorporate external knowledge into large language models, thereby enhancing factuality and reducing hallucinations in question-answering (QA) tasks. A standard RAG pipeline consists of several components, such as query rewriting, document retrieval, document filtering, and answer generation. However, these components are typically optimized separately through supervised fine-tuning, which can lead to misalignments between the objectives of individual components and the overarching aim of generating accurate answers. Although recent efforts have explored using reinforcement learning (RL) to optimize specific RAG components, these approaches often focus on simple pipelines with only two components or do not adequately address the complex interdependencies and collaborative interactions among the modules. To overcome these limitations, we propose treating the complex RAG pipeline with multiple components as a multi-agent cooperative task, in which each component can be regarded as an RL agent. Specifically, we present MMOA-RAG\footnote{The code of MMOA-RAG is on \url{https://github.com/chenyiqun/MMOA-RAG}.}, \textbf{M}ulti-\textbf{M}odule joint \textbf{O}ptimization \textbf{A}lgorithm for \textbf{RAG}, which employs multi-agent reinforcement learning to harmonize all agents' goals toward a unified reward, such as the F1 score of the final answer. Experiments conducted on various QA benchmarks demonstrate that MMOA-RAG effectively boost the overall performance of the pipeline and outperforms existing baselines. Furthermore, comprehensive ablation studies validate the contributions of individual components and demonstrate MMOA-RAG can be adapted to different RAG pipelines and benchmarks. Yiqun Chen 0004, Lingyong Yan, Weiwei Sun 0001, Xinyu Ma 0001, Yi Zhang 0050, Shuaiqiang Wang, Dawei Yin 0001, Yiming Yang 0002, Jiaxin Mao |
NeurIPS | 5 |
| 2025 | TourRank: Utilizing Large Language Models for Documents Ranking with a Tournament-Inspired StrategyabstractLarge Language Models (LLMs) are increasingly employed in zero-shot documents ranking, yielding commendable results. However, several significant challenges still persist in LLMs for ranking: (1) LLMs are constrained by limited input length, precluding them from processing a large number of documents simultaneously; (2) The output document sequence is influenced by the input order of documents, resulting in inconsistent ranking outcomes; (3) Achieving a balance between cost and ranking performance is challenging. To tackle these issues, we introduce a novel documents ranking method called TourRank1. which is inspired by the sport tournaments, such as FIFA World Cup. Specifically, we 1) overcome the limitation in input length and reduce the ranking latency by incorporating a multi-stage grouping strategy similar to the parallel group stage of sport tournaments; 2) improve the ranking performance and robustness to input orders by using a points system to ensemble multiple ranking results. We test TourRank with different LLMs on the TREC DL datasets and the BEIR benchmark. The experimental results demonstrate that TourRank delivers state-of-the-art performance at a modest cost. Yiqun Chen 0004, Qi Liu 0071, Yi Zhang 0050, Weiwei Sun 0001, Xinyu Ma 0001, Wei Yang 0041, Daiting Shi, Jiaxin Mao, Dawei Yin 0001 |
WWW | 3 |
| 2025 | MA4DIV: Multi-Agent Reinforcement Learning for Search Result DiversificationabstractSearch result diversification (SRD), which aims to ensure that documents in a ranking list cover a broad range of subtopics, is a significant and widely studied problem in Information Retrieval and Web Search. Existing methods primarily utilize a paradigm of ''greedy selection'', i.e., selecting one document with the highest diversity score at a time or optimize an approximation of the objective function. These approaches tend to be inefficient and are easily trapped in a suboptimal state. To address these challenges, we introduce Multi-Agent reinforcement learning (MARL) for search result DIVersity, which called MA4DIV. In this approach, each document is an agent and the search result diversification is modeled as a cooperative task among multiple agents. By modeling the SRD ranking problem as a cooperative MARL problem, this approach allows for directly optimizing the diversity metrics, such as α-NDCG, while achieving high training efficiency. We conducted experiments on public TREC datasets and a larger scale dataset in the industrial setting. The experiemnts show that MA4DIV achieves substantial improvements in both effectiveness and efficiency than existing baselines, especially on the industrial dataset. Yiqun Chen 0004, Jiaxin Mao, Yi Zhang 0050, Dehong Ma, Daiting Shi, Zhicong Cheng, Simiu Gu, Dawei Yin 0001 |
WWW | 3 |
| 2022 | Alleviating the Knowledge-Language Inconsistency: A Study for Deep Commonsense KnowledgeabstractKnowledge facts are typically represented by relational triples, while we observe that some commonsense facts are represented by triples whose forms are inconsistent with the corresponding language expressions. For commonsense mining tasks, this inconsistency raises a challenge for the prevailing methods using pre-trained language models that learn the expression of language. However, there are few studies which focus on this inconsistency issue. To fill this empty, in this paper, we term the commonsense knowledge whose triple form is heavily inconsistent with the language expression asdeep commonsense knowledgeand first conduct extensive exploratory experiments to study deep commonsense knowledge. We show that deep commonsense knowledge occupies a significant part of commonsense knowledge, while the conventional methods based on pre-trained language models fail to capture it effectively. We further propose a novel method to mine the deep commonsense knowledge from raw text that is exactly language expression, alleviating the reliance of conventional methods on the triple representation form. Experiments demonstrate that our proposed method substantially improves the performance in mining deep commonsense knowledge. Yi Zhang 0050, Lei Li 0039, Yunfang Wu, Qi Su 0001, Xu Sun 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | A Global Past-Future Early Exit Method for Accelerating Inference of Pre-trained Language ModelsabstractKaiyuan Liao, Yi Zhang, Xuancheng Ren, Qi Su, Xu Sun, Bin He. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Kaiyuan Liao, Yi Zhang 0050, Xuancheng Ren, Qi Su 0001, Xu Sun 0001 |
NAACL-HLT | 2 |
| 2020 | Parallel Data Augmentation for Formality Style TransferabstractThe main barrier to progress in the task of Formality Style Transfer is the inadequacy of training data.In this paper, we study how to augment parallel data and propose novel and simple data augmentation methods for this task to obtain useful sentence pairs with easily accessible models and systems.Experiments demonstrate that our augmented parallel data largely helps improve formality style transfer when it is used to pre-train the model, leading to the state-of-the-art results in the GYAFC benchmark dataset 1 . Yi Zhang 0050, Tao Ge 0001, Xu Sun 0001 |
ACL | 1 |
| 2020 | Training Simplification and Model Simplification for Deep Learning : A Minimal Effort Back Propagation MethodabstractWe propose a simple yet effective technique to simplify the training and the resulting model of neural networks. In back propagation, only a small subset of the full gradient is computed to update the model parameters. The gradient vectors are sparsified in such a way that only the top-k elements (in terms of magnitude) are kept. As a result, only k rows or columns (depending on the layout) of the weight matrix are modified, leading to a linear reduction in the computational cost. Based on the sparsified gradients, we further simplify the model by eliminating the rows or columns that are seldom updated, which will reduce the computational cost both in the training and decoding, and potentially accelerate decoding in real-world applications. Surprisingly, experimental results demonstrate that most of the time we only need to update fewer than 5 percent of the weights at each back propagation pass. More interestingly, the accuracy of the resulting models is actually improved rather than degraded, and a detailed analysis is given. The model simplification results show that we could adaptively simplify the model which could often be reduced by around 9x, without any loss on accuracy or even with improved accuracy. Xu Sun 0001, Xuancheng Ren, Shuming Ma, Bingzhen Wei, Wei Li 0101, Jingjing Xu 0001, Houfeng Wang, Yi Zhang 0050 |
IEEE Trans. Knowl. Data Eng. | 8 |
| 2019 | Towards easier and faster sequence labeling for natural language processing: A search-based probabilistic online learning framework (SAPO)
Xu Sun 0001, Shuming Ma, Yi Zhang 0050, Xuancheng Ren |
Inf. Sci. | 3 |
| 2019 | Regularizing Output Distribution of Abstractive Chinese Social Media Text Summarization for Improved Semantic ConsistencyabstractAbstractive text summarization is a highly difficult problem, and the sequence-to-sequence model has shown success in improving the performance on the task. However, the generated summaries are often inconsistent with the source content in semantics. In such cases, when generating summaries, the model selects semantically unrelated words with respect to the source content as the most probable output. The problem can be attributed to heuristically constructed training data, where summaries can be unrelated to the source content, thus containing semantically unrelated words and spurious word correspondence. In this article, we propose a regularization approach for the sequence-to-sequence model and make use of what the model has learned to regularize the learning objective to alleviate the effect of the problem. In addition, we propose a practical human evaluation method to address the problem that the existing automatic evaluation method does not evaluate the semantic consistency with the source content properly. Experimental results demonstrate the effectiveness of the proposed approach, which outperforms almost all the existing models. Especially, the proposed approach improves the semantic consistency by 4% in terms of human evaluation. Bingzhen Wei, Xuancheng Ren, Yi Zhang 0050, Xiaoyan Cai, Qi Su 0001, Xu Sun 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2018 | Does Higher Order LSTM Have Better Accuracy for Segmenting and Labeling Sequence Data?abstractExisting neural models usually predict the tag of the current token independent of the neighboring tags. The popular LSTM-CRF model considers the tag dependencies between every two consecutive tags. However, it is hard for existing neural models to take longer distance dependencies between tags into consideration. The scalability is mainly limited by the complex model structures and the cost of dynamic programming during training. In our work, we first design a new model called “high order LSTM” to predict multiple tags for the current token which contains not only the current tag but also the previous several tags. We call the number of tags in one prediction as “order”. Then we propose a new method called Multi-Order BiLSTM (MO-BiLSTM) which combines low order and high order LSTMs together. MO-BiLSTM keeps the scalability to high order models with a pruning technique. We evaluate MO-BiLSTM on all-phrase chunking and NER datasets. Experiment results show that MO-BiLSTM achieves the state-of-the-art result in chunking and highly competitive results in two NER datasets. Yi Zhang 0050, Xu Sun 0001, Shuming Ma, Yang Yang 0125, Xuancheng Ren |
COLING | 1 |
| 2018 | A Skeleton-Based Model for Promoting Coherence Among Sentences in Narrative Story GenerationabstractNarrative story generation is a challenging problem because it demands the generated sentences with tight semantic connections, which has not been well studied by most existing generative models.To address this problem, we propose a skeleton-based model to promote the coherence of generated stories.Different from traditional models that generate a complete sentence at a stroke, the proposed model first generates the most critical phrases, called skeleton, and then expands the skeleton to a complete and fluent sentence.The skeleton is not manually defined, but learned by a reinforcement learning method.Compared to the state-of-the-art models, our skeleton-based model can generate significantly more coherent text according to human evaluation and automatic evaluation.The G-score is improved by 20.1% in human evaluation. 1 Jingjing Xu 0001, Xuancheng Ren, Yi Zhang 0050, Qi Zeng 0001, Xiaoyan Cai, Xu Sun 0001 |
EMNLP | 3 |
| 2018 | Learning Sentiment Memories for Sentiment Modification without Parallel DataabstractThe task of sentiment modification requires reversing the sentiment of the input and preserving the sentiment-independent content.However, aligned sentences with the same content but different sentiments are usually unavailable.Due to the lack of such parallel data, it is hard to extract sentiment independent content and reverse the sentiment in an unsupervised way.Previous work usually can not reconcile sentiment transformation and content preservation.In this paper, motivated by the fact the non-emotional context (e.g., "staff") provides strong cues for the occurrence of emotional words (e.g., "friendly"), we propose a novel method that automatically extracts appropriate sentiment information from the learned sentiment memories according to the specific context.Experiments show that our method substantially improves the content preservation degree and achieves the state-of-the-art performance.1 Yi Zhang 0050, Jingjing Xu 0001, Xu Sun 0001 |
EMNLP | 1 |
| 2018 | A Chinese Dataset with Negative Full Forms for General Abbreviation Prediction
Yi Zhang 0050, Xu Sun 0001 |
LREC | 1 |
| 2018 | Accelerating Graph-Based Dependency Parsing with Lock-Free Parallel Perceptron
Shuming Ma, Xu Sun 0001, Yi Zhang 0050, Bingzhen Wei |
NLPCC (1) | 3 |
| 2017 | Transfer Deep Learning for Low-Resource Chinese Word Segmentation with a Novel Neural Network
Jingjing Xu 0001, Shuming Ma, Yi Zhang 0050, Bingzhen Wei, Xiaoyan Cai, Xu Sun 0001 |
NLPCC | 3 |