VLDB 2026 Research / reviewers in the wild / expert
Fangwei Zhu
dblp:237/4830
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0006-6232-1610ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 58% Efficient and distributed learning · 39% Information extraction and text analysis · 3% | |
| Databases, data mining, and information retrieval
2 papers |
Knowledge graphs · 88% Information retrieval · 6% Web and social media mining · 6% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › federated learning
contribution evaluation |
1.0 | 1 | 2026 | RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection · AAAI 2026 |
Machine learning › Efficient and distributed learning
data selection |
1.0 | 1 | 2026 | RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection · AAAI 2026 |
Natural language and speech › Language models and text generation › instruction tuning
instruction data selection |
1.0 | 1 | 2026 | RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection · AAAI 2026 |
Natural language and speech › Language models and text generation
instruction tuning |
1.0 | 1 | 2026 | RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection · AAAI 2026 |
Knowledge graphs
entity linking |
0.8 | 1 | 2024 | XLORE 3: A Large-Scale Multilingual Knowledge Graph from Heterogeneous Wiki Knowledge Resources · ACM Trans. Inf. Syst. 2024 |
Knowledge graphs
link prediction |
0.8 | 1 | 2024 | XLORE 3: A Large-Scale Multilingual Knowledge Graph from Heterogeneous Wiki Knowledge Resources · ACM Trans. Inf. Syst. 2024 |
Natural language and speech › Language models and text generation › text summarization › scientific document summarization
abstract generation |
0.5 | 1 | 2021 | TWAG: A Topic-Guided Wikipedia Abstract Generator · ACL/IJCNLP (1) 2021 |
Natural language and speech › Language models and text generation
text generation |
0.5 | 1 | 2021 | TWAG: A Topic-Guided Wikipedia Abstract Generator · ACL/IJCNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis
topic model |
0.1 | 1 | 2021 | TWAG: A Topic-Guided Wikipedia Abstract Generator · ACL/IJCNLP (1) 2021 |
Information retrieval
search engines |
0.1 | 1 | 2021 | TWAG: A Topic-Guided Wikipedia Abstract Generator · ACL/IJCNLP (1) 2021 |
Web and social media mining › user-generated content
wikipedia |
0.1 | 1 | 2021 | TWAG: A Topic-Guided Wikipedia Abstract Generator · ACL/IJCNLP (1) 2021 |
Methods — techniques the papers use, named apart from their topics
topic-guided generation · 1.0neural abstractive summarization · 1.0in-context learning · 1.0gradient-free contribution measurement · 1.0pre-trained language model · 0.8knowledge completion · 0.8entity linking · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data SelectionabstractData selection for instruction tuning is crucial for improving the performance of large language models (LLMs) while reducing training costs. In this paper, we propose Refined Contribution Measurement with In-Context Learning (RICo), a novel gradient-free method that quantifies the fine-grained contribution of individual samples to both task-level and global-level model performance. RICo enables more accurate identification of high-contribution data, leading to better instruction tuning. We also introduce a lightweight selection paradigm trained on RICo scores, enabling scalable data selection with strictly linear inference complexity. Extensive experiments on 3 LLMs across 12 benchmarks and 5 pairwise evaluation sets demonstrate the effectiveness of RICo. Remarkably, on LLaMA3.1-8B, models trained in 15% of RICo-selected data outperform full datasets by 5.42 percentage points and exceed the best performance of widely used selection methods by 1.48 percentage points. We further analyze high-contribution samples selected by RICo, which show both diverse tasks and appropriate difficulty levels, rather than merely the most difficult cases. Qingxiu Dong, Linli Yao, Fangwei Zhu, Weilin Luo, Zhifang Sui |
AAAI | 4 |
| 2025 | LLMAEL: Large Language Models are Good Context Augmenters for Entity LinkingabstractSpecialized entity linking (EL) models are well-trained at mapping mentions to unique knowledge base (KB) entities according to a given context. However, specialized EL models struggle to disambiguate long-tail entities due to their limited training data. Meanwhile, extensively pre-trained large language models (LLMs) possess broader knowledge of uncommon entities. Yet, with a lack of specialized EL training, LLMs frequently fail to generate accurate KB entity names, limiting their standalone effectiveness in EL. With the observation that LLMs are more adept at context generation instead of EL execution, we introduce LLM-Augmented Entity Linking (LLMAEL), the first framework to enhance specialized EL models with LLM data augmentation. LLMAEL leverages off-the-shelf, tuning-free LLMs as context augmenters, generating entity descriptions to serve as additional input for specialized EL models. Experiments show that LLMAEL sets new state-of-the-art results across 6 widely adopted EL benchmarks: compared to prior methods that integrate tuning-free LLMs into EL, LLMAEL achieves an absolute 8.9% gain in EL accuracy. We release our code and datasets. Amy Xin, Yunjia Qi, Zijun Yao 0002, Fangwei Zhu, Kaisheng Zeng, Bin Xu 0001, Lei Hou 0001, Juan-Zi Li |
CIKM | 4 |
| 2025 | Language Models Encode the Value of Numbers LinearlyabstractLarge language models (LLMs) have exhibited impressive competence in various tasks, but their internal mechanisms on mathematical problems are still under-explored. In this paper, we study a fundamental question: how language models encode the value of numbers, a basic element in math. To study the question, we construct a synthetic dataset comprising addition problems and utilize linear probes to read out input numbers from the hidden states. Experimental results support the existence of encoded number values in LLMs on different layers, and these values can be extracted via linear probes. Further experiments show that LLMs store their calculation results in a similar manner, and we can intervene the output via simple vector additions, proving the causal connection between encoded numbers and language model outputs. Our research provides evidence that LLMs encode the value of numbers linearly, offering insights for better exploring, designing, and utilizing numeric information in LLMs. Fangwei Zhu, Damai Dai, Zhifang Sui |
COLING | 1 |
| 2024 | CoUDA: Coherence Evaluation via Unified Data AugmentationabstractDawei Zhu, Wenhao Wu, Yifan Song, Fangwei Zhu, Ziqiang Cao, Sujian Li. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yifan Song 0002, Fangwei Zhu, Ziqiang Cao, Sujian Li |
NAACL-HLT | 4 |
| 2024 | XLORE 3: A Large-Scale Multilingual Knowledge Graph from Heterogeneous Wiki Knowledge ResourcesabstractIn recent years, knowledge graph (KG) has attracted significant attention from academia and industry, resulting in the development of numerous technologies for KG construction, completion, and application. XLORE is one of the largest multilingual KGs built from Baidu Baike and Wikipedia via a series of knowledge modeling and acquisition methods. In this article, we utilize systematic methods to improve XLORE's data quality and present its latest version, XLORE 3, which enables the effective integration and management of heterogeneous knowledge from diverse resources. Compared with previous versions, XLORE 3 has three major advantages: (1) We design a comprehensive and reasonable schema, namely XLORE ontology, which can effectively organize and manage entities from various resources. (2) We merge equivalent entities in different languages to facilitate knowledge sharing. We provide a large-scale entity linking system to establish the associations between unstructured text and structured KG. (3) We design a multi-strategy knowledge completion framework, which leverages pre-trained language models and vast amounts of unstructured text to discover missing and new facts. The resulting KG contains 446 concepts, 2,608 properties, 66 million entities, and more than 2 billion facts. It is available and downloadable online at https://www.xlore.cn/ , providing a valuable resource for researchers and practitioners in various fields. Kaisheng Zeng, Hailong Jin, Fangwei Zhu, Lei Hou 0001, Yi Zhang 0163, Fan Pang, Dingxiao Liu, Juan-Zi Li |
ACM Trans. Inf. Syst. | 4 |
| 2022 | UPER: Boosting Multi-Document Summarization with an Unsupervised Prompt-based ExtractorabstractMulti-Document Summarization (MDS) commonly employs the 2-stage extract-then-abstract paradigm, which first extracts a relatively short meta-document, then feeds it into the deep neural networks to generate an abstract. Previous work usually takes the ROUGE score as the label for training a scoring model to evaluate source documents. However, the trained scoring model is prone to under-fitting for low-resource settings, as it relies on the training data. To extract documents effectively, we construct prompting templates that invoke the underlying knowledge in Pre-trained Language Model (PLM) to calculate the document and keyword’s perplexity, which can assess the document’s semantic salience. Our unsupervised approach can be applied as a plug-in to boost other metrics for evaluating a document’s salience, thus improving the subsequent abstract generation. We get positive results on 2 MDS datasets, 2 data settings, and 2 abstractive backbone models, showing our method’s effectiveness. Our code is available at https://github.com/THU-KEG/UPER Shangqing Tu, Jifan Yu, Fangwei Zhu, Juan-Zi Li, Lei Hou 0001, Jian-Yun Nie |
COLING | 3 |
| 2021 | TWAG: A Topic-Guided Wikipedia Abstract GeneratorabstractFangwei Zhu, Shangqing Tu, Jiaxin Shi, Juanzi Li, Lei Hou, Tong Cui. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Fangwei Zhu, Shangqing Tu, Jiaxin Shi, Juan-Zi Li, Lei Hou 0001, Tong Cui |
ACL/IJCNLP (1) | 1 |