VLDB 2026 Research / reviewers in the wild / expert
Kewen Zhao
dblp:03/1469
· DBLP profile ↗
5ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Information extraction and text analysis · 33% Trustworthy machine learning · 30% Efficient and distributed learning · 15% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 77% Knowledge graphs · 23% |
Topics — the 8 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › document understanding › legal text analysis
patent approval prediction |
1.3 | 2 | 2024 | Beyond Scaling: Predicting Patent Approval with Domain-specific Fine-grained Claim Dependency Graph · ACL (1) 2024 Towards Comprehensive Patent Approval Predictions: Beyond Traditional Document Classification · ACL (1) 2022 |
Machine learning › Trustworthy machine learning › Data-centric AI
data valuation |
0.9 | 1 | 2025 | What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions · NeurIPS 2025 |
Machine learning › Learning paradigms › continual learning
gradient projection |
0.9 | 1 | 2025 | What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › interpretability › training data attribution
influence function |
0.9 | 1 | 2025 | What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › large-scale learning
scalable training |
0.9 | 1 | 2025 | What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions · NeurIPS 2025 |
Natural language and speech › Information extraction and text analysis
text classification |
0.6 | 1 | 2022 | Towards Comprehensive Patent Approval Predictions: Beyond Traditional Document Classification · ACL (1) 2022 |
Natural language and speech › Language models and text generation
large language model training |
0.3 | 1 | 2025 | What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions · NeurIPS 2025 |
Knowledge graphs
domain-specific knowledge graph |
0.2 | 1 | 2024 | Beyond Scaling: Predicting Patent Approval with Domain-specific Fine-grained Claim Dependency Graph · ACL (1) 2024 |
Methods — techniques the papers use, named apart from their topics
graph neural network · 1.5fine-grained graph modeling · 1.5influence functions · 0.9gradient projection · 0.9document classification · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence FunctionsabstractLarge language models (LLMs) are trained on a vast amount of human-written data, but data providers often remain uncredited. In response to this issue, data valuation (or data attribution), which quantifies the contribution or value of each data to the model output, has been discussed as a potential solution. Nevertheless, applying existing data valuation methods to recent LLMs and their vast training datasets has been largely limited by prohibitive compute and memory costs. In this work, we focus on influence functions, a popular gradient-based data valuation method, and significantly improve its scalability with an efficient gradient projection strategy called LoGra that leverages the gradient structure in backpropagation. We then provide a theoretical motivation of gradient projection approaches to influence functions to promote trust in the data valuation process. Lastly, we lower the barrier to implementing data valuation systems by introducing LogIX, a software package that can transform existing training code into data valuation code with minimal effort. In our data valuation experiments, LoGra achieves competitive accuracy against more expensive baselines while showing up to 6,500x improvement in throughput and 5x reduction in GPU memory usage when applied to Llama3-8B-Instruct and the 1B-token dataset. Sang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, Jeff G. Schneider, Eduard H. Hovy, Roger B. Grosse, Eric P. Xing |
NeurIPS | 4 |
| 2024 | Beyond Scaling: Predicting Patent Approval with Domain-specific Fine-grained Claim Dependency GraphabstractXiaochen Gao, Feng Yao, Kewen Zhao, Beilei He, Animesh Kumar, Vish Krishnan, Jingbo Shang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Xiaochen Gao, Kewen Zhao, Beilei He, Animesh Kumar, Vish Krishnan, Jingbo Shang |
ACL (1) | 3 |
| 2022 | Towards Comprehensive Patent Approval Predictions: Beyond Traditional Document ClassificationabstractXiaochen Gao, Zhaoyi Hou, Yifei Ning, Kewen Zhao, Beilei He, Jingbo Shang, Vish Krishnan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Xiaochen Gao, Zhaoyi Hou, Yifei Ning, Kewen Zhao, Beilei He, Jingbo Shang, Vish Krishnan |
ACL (1) | 4 |
| 2012 | Generalizing Sufficient Conditions and Traceable Graphs
Kewen Zhao |
ICIC (1) | 1 |
| 2009 | A sufficient condition for pancyclic graphs
Kewen Zhao, Ping Zhang 0004 |
Inf. Process. Lett. | 1 |