Fangwei Zhu

dblp:237/4830 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0006-6232-1610ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 58% Efficient and distributed learning · 39% Information extraction and text analysis · 3%
Databases, data mining, and information retrieval
2 papers
Knowledge graphs · 88% Information retrieval · 6% Web and social media mining · 6%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › federated learning
contribution evaluation
1.012026
RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection · AAAI 2026
Machine learning › Efficient and distributed learning
data selection
1.012026
RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection · AAAI 2026
Natural language and speech › Language models and text generation › instruction tuning
instruction data selection
1.012026
RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection · AAAI 2026
Natural language and speech › Language models and text generation
instruction tuning
1.012026
RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection · AAAI 2026
Knowledge graphs
entity linking
0.812024
XLORE 3: A Large-Scale Multilingual Knowledge Graph from Heterogeneous Wiki Knowledge Resources · ACM Trans. Inf. Syst. 2024
Knowledge graphs
link prediction
0.812024
XLORE 3: A Large-Scale Multilingual Knowledge Graph from Heterogeneous Wiki Knowledge Resources · ACM Trans. Inf. Syst. 2024
Natural language and speech › Language models and text generation › text summarization › scientific document summarization
abstract generation
0.512021
TWAG: A Topic-Guided Wikipedia Abstract Generator · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation
text generation
0.512021
TWAG: A Topic-Guided Wikipedia Abstract Generator · ACL/IJCNLP (1) 2021
Natural language and speech › Information extraction and text analysis
topic model
0.112021
TWAG: A Topic-Guided Wikipedia Abstract Generator · ACL/IJCNLP (1) 2021
Information retrieval
search engines
0.112021
TWAG: A Topic-Guided Wikipedia Abstract Generator · ACL/IJCNLP (1) 2021
Web and social media mining › user-generated content
wikipedia
0.112021
TWAG: A Topic-Guided Wikipedia Abstract Generator · ACL/IJCNLP (1) 2021

Methods — techniques the papers use, named apart from their topics

topic-guided generation · 1.0neural abstractive summarization · 1.0in-context learning · 1.0gradient-free contribution measurement · 1.0pre-trained language model · 0.8knowledge completion · 0.8entity linking · 0.8
YearPublicationVenuePosition
2026 RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection
abstract
Data selection for instruction tuning is crucial for improving the performance of large language models (LLMs) while reducing training costs. In this paper, we propose Refined Contribution Measurement with In-Context Learning (RICo), a novel gradient-free method that quantifies the fine-grained contribution of individual samples to both task-level and global-level model performance. RICo enables more accurate identification of high-contribution data, leading to better instruction tuning. We also introduce a lightweight selection paradigm trained on RICo scores, enabling scalable data selection with strictly linear inference complexity. Extensive experiments on 3 LLMs across 12 benchmarks and 5 pairwise evaluation sets demonstrate the effectiveness of RICo. Remarkably, on LLaMA3.1-8B, models trained in 15% of RICo-selected data outperform full datasets by 5.42 percentage points and exceed the best performance of widely used selection methods by 1.48 percentage points. We further analyze high-contribution samples selected by RICo, which show both diverse tasks and appropriate difficulty levels, rather than merely the most difficult cases.
Qingxiu Dong, Linli Yao, Fangwei Zhu, Weilin Luo, Zhifang Sui
AAAI4
2025 LLMAEL: Large Language Models are Good Context Augmenters for Entity Linking
abstract
Specialized entity linking (EL) models are well-trained at mapping mentions to unique knowledge base (KB) entities according to a given context. However, specialized EL models struggle to disambiguate long-tail entities due to their limited training data. Meanwhile, extensively pre-trained large language models (LLMs) possess broader knowledge of uncommon entities. Yet, with a lack of specialized EL training, LLMs frequently fail to generate accurate KB entity names, limiting their standalone effectiveness in EL. With the observation that LLMs are more adept at context generation instead of EL execution, we introduce LLM-Augmented Entity Linking (LLMAEL), the first framework to enhance specialized EL models with LLM data augmentation. LLMAEL leverages off-the-shelf, tuning-free LLMs as context augmenters, generating entity descriptions to serve as additional input for specialized EL models. Experiments show that LLMAEL sets new state-of-the-art results across 6 widely adopted EL benchmarks: compared to prior methods that integrate tuning-free LLMs into EL, LLMAEL achieves an absolute 8.9% gain in EL accuracy. We release our code and datasets.
Amy Xin, Yunjia Qi, Zijun Yao 0002, Fangwei Zhu, Kaisheng Zeng, Bin Xu 0001, Lei Hou 0001, Juan-Zi Li
CIKM4
2025 Language Models Encode the Value of Numbers Linearly
abstract
Large language models (LLMs) have exhibited impressive competence in various tasks, but their internal mechanisms on mathematical problems are still under-explored. In this paper, we study a fundamental question: how language models encode the value of numbers, a basic element in math. To study the question, we construct a synthetic dataset comprising addition problems and utilize linear probes to read out input numbers from the hidden states. Experimental results support the existence of encoded number values in LLMs on different layers, and these values can be extracted via linear probes. Further experiments show that LLMs store their calculation results in a similar manner, and we can intervene the output via simple vector additions, proving the causal connection between encoded numbers and language model outputs. Our research provides evidence that LLMs encode the value of numbers linearly, offering insights for better exploring, designing, and utilizing numeric information in LLMs.
Fangwei Zhu, Damai Dai, Zhifang Sui
COLING1
2024 CoUDA: Coherence Evaluation via Unified Data Augmentation
abstract
Dawei Zhu, Wenhao Wu, Yifan Song, Fangwei Zhu, Ziqiang Cao, Sujian Li. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yifan Song 0002, Fangwei Zhu, Ziqiang Cao, Sujian Li
NAACL-HLT4
2024 XLORE 3: A Large-Scale Multilingual Knowledge Graph from Heterogeneous Wiki Knowledge Resources
abstract
In recent years, knowledge graph (KG) has attracted significant attention from academia and industry, resulting in the development of numerous technologies for KG construction, completion, and application. XLORE is one of the largest multilingual KGs built from Baidu Baike and Wikipedia via a series of knowledge modeling and acquisition methods. In this article, we utilize systematic methods to improve XLORE's data quality and present its latest version, XLORE 3, which enables the effective integration and management of heterogeneous knowledge from diverse resources. Compared with previous versions, XLORE 3 has three major advantages: (1) We design a comprehensive and reasonable schema, namely XLORE ontology, which can effectively organize and manage entities from various resources. (2) We merge equivalent entities in different languages to facilitate knowledge sharing. We provide a large-scale entity linking system to establish the associations between unstructured text and structured KG. (3) We design a multi-strategy knowledge completion framework, which leverages pre-trained language models and vast amounts of unstructured text to discover missing and new facts. The resulting KG contains 446 concepts, 2,608 properties, 66 million entities, and more than 2 billion facts. It is available and downloadable online at https://www.xlore.cn/ , providing a valuable resource for researchers and practitioners in various fields.
Kaisheng Zeng, Hailong Jin, Fangwei Zhu, Lei Hou 0001, Yi Zhang 0163, Fan Pang, Dingxiao Liu, Juan-Zi Li
ACM Trans. Inf. Syst.4
2022 UPER: Boosting Multi-Document Summarization with an Unsupervised Prompt-based Extractor
abstract
Multi-Document Summarization (MDS) commonly employs the 2-stage extract-then-abstract paradigm, which first extracts a relatively short meta-document, then feeds it into the deep neural networks to generate an abstract. Previous work usually takes the ROUGE score as the label for training a scoring model to evaluate source documents. However, the trained scoring model is prone to under-fitting for low-resource settings, as it relies on the training data. To extract documents effectively, we construct prompting templates that invoke the underlying knowledge in Pre-trained Language Model (PLM) to calculate the document and keyword’s perplexity, which can assess the document’s semantic salience. Our unsupervised approach can be applied as a plug-in to boost other metrics for evaluating a document’s salience, thus improving the subsequent abstract generation. We get positive results on 2 MDS datasets, 2 data settings, and 2 abstractive backbone models, showing our method’s effectiveness. Our code is available at https://github.com/THU-KEG/UPER
Shangqing Tu, Jifan Yu, Fangwei Zhu, Juan-Zi Li, Lei Hou 0001, Jian-Yun Nie
COLING3
2021 TWAG: A Topic-Guided Wikipedia Abstract Generator
abstract
Fangwei Zhu, Shangqing Tu, Jiaxin Shi, Juanzi Li, Lei Hou, Tong Cui. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Fangwei Zhu, Shangqing Tu, Jiaxin Shi, Juan-Zi Li, Lei Hou 0001, Tong Cui
ACL/IJCNLP (1)1