VLDB 2026 Research / reviewers in the wild / expert
Yanzhen Shen
dblp:350/0200
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0005-0113-3942ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Information extraction and text analysis · 41% Graph learning · 33% Efficient and distributed learning · 16% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 34% Recommender systems · 26% Knowledge graphs · 20% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning
heterogeneous graph learning |
0.9 | 1 | 2025 | GraphRouter: A Graph-based Router for LLM Selections · ICLR 2025 |
Machine learning › Graph learning
link prediction |
0.9 | 1 | 2025 | GraphRouter: A Graph-based Router for LLM Selections · ICLR 2025 |
Machine learning › Efficient and distributed learning › inference efficiency
LLM routing |
0.9 | 1 | 2025 | GraphRouter: A Graph-based Router for LLM Selections · ICLR 2025 |
Information retrieval › retrieval models › neural retrieval
dense retrieval |
0.9 | 1 | 2025 | LogiCoL: Logically-Informed Contrastive Learning for Set-based Dense Retrieval · EMNLP 2025 |
Recommender systems › user recommendation
reviewer assignment |
0.9 | 1 | 2025 | Chain-of-Factors Paper-Reviewer Matching · WWW 2025 |
Natural language and speech › Information extraction and text analysis
entity typing |
0.8 | 1 | 2024 | Seed-Guided Fine-Grained Entity Typing in Science and Engineering Domains · AAAI 2024 |
Natural language and speech › Information extraction and text analysis › entity typing
fine-grained entity typing |
0.8 | 1 | 2024 | Seed-Guided Fine-Grained Entity Typing in Science and Engineering Domains · AAAI 2024 |
Data mining › text mining › text classification
hierarchical text classification |
0.7 | 1 | 2023 | Weakly Supervised Multi-Label Classification of Full-Text Scientific Papers · KDD 2023 |
Knowledge graphs
taxonomy construction |
0.7 | 1 | 2023 | Weakly Supervised Multi-Label Classification of Full-Text Scientific Papers · KDD 2023 |
Machine learning › Learning theory
model selection |
0.3 | 1 | 2025 | GraphRouter: A Graph-based Router for LLM Selections · ICLR 2025 |
Information retrieval › search engines › semantic search
entity retrieval |
0.3 | 1 | 2025 | LogiCoL: Logically-Informed Contrastive Learning for Set-based Dense Retrieval · EMNLP 2025 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.2 | 1 | 2024 | Seed-Guided Fine-Grained Entity Typing in Science and Engineering Domains · AAAI 2024 |
Methods — techniques the papers use, named apart from their topics
weak supervision · 2.1network-aware contrastive fine-tuning · 1.3hierarchy-aware aggregation · 1.3supervised contrastive learning · 0.9instruction tuning · 0.9inductive learning · 0.9graph neural network · 0.9edge prediction · 0.9contrastive learning · 0.9contextualized language model · 0.9textual entailment · 0.8contextualized representations · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LogiCoL: Logically-Informed Contrastive Learning for Set-based Dense RetrievalabstractWhile significant progress has been made with dual-and bi-encoder dense retrievers, they often struggle on queries with logical connectives, a use case often overlooked yet important in downstream applications.In this paper, we introduce LOGICOL, a logically informed contrastive learning objective for dense retrievers.LOGICOL builds upon in-batch supervised contrastive learning and learns dense retrievers to respect the subset and mutually exclusive set relation between query results.We evaluated the effectiveness of LOGICOL in the entity retrieval task, where the model is expected to retrieve a set of Wikipedia entities that satisfy the implicit logical constraints of the query.We show that models trained with LOGICOL show improvements both in terms of retrieval performance and logical consistency in the results.We provide detailed analysis and insights to uncover why queries with logical connectives are challenging for dense retrievers and why LOGI-COL is effective.Our codes and data are available at https://github.com/yanzhen4/LogiCoL. A not B: Species of orchids inMalaysia but not Thailand. Yanzhen Shen, Xueqiang Xu, Yunyi Zhang 0001, Chaitanya Malaviya, Dan Roth 0001 |
EMNLP | 1 |
| 2025 | GraphRouter: A Graph-based Router for LLM SelectionsabstractThe rapidly growing number and variety of Large Language Models (LLMs)
present significant challenges in efficiently selecting the appropriate LLM for
a given query, especially considering the trade-offs between performance and
computational cost. Current LLM selection methods often struggle to generalize
across new LLMs and different tasks because of their limited ability to leverage
contextual interactions among tasks, queries, and LLMs, as well as their depen-
dence on a transductive learning framework. To address these shortcomings, we
introduce a novel inductive graph framework, named as GraphRouter, which
fully utilizes the contextual information among tasks, queries, and LLMs to en-
hance the LLM selection process. GraphRouter constructs a heterogeneous
graph comprising task, query, and LLM nodes, with interactions represented as
edges, which efficiently captures the contextual information between the query’s
requirements and the LLM’s capabilities. Through an innovative edge prediction
mechanism, GraphRouter is able to predict attributes (the effect and cost of
LLM response) of potential edges, allowing for optimized recommendations that
adapt to both existing and newly introduced LLMs without requiring retraining.
Comprehensive experiments across three distinct effect-cost weight scenarios have
shown that GraphRouter substantially surpasses existing routers, delivering a
minimum performance improvement of 12.3%. In addition, it achieves enhanced
generalization across new LLMs settings and supports diverse tasks with at least a
9.5% boost in effect and a significant reduction in computational demands. This
work endeavors to apply a graph-based approach for the contextual and adaptive
selection of LLMs, offering insights for real-world applications. Yanzhen Shen, Jiaxuan You |
ICLR | 2 |
| 2025 | Chain-of-Factors Paper-Reviewer MatchingabstractWith the rapid increase in paper submissions to academic conferences, the need for automated and accurate paper-reviewer matching is more critical than ever. Previous efforts in this area have considered various factors to assess the relevance of a reviewer's expertise to a paper, such as the semantic similarity, shared topics, and citation connections between the paper and the reviewer's previous works. However, most of these studies focus on only one factor, resulting in an incomplete evaluation of the paper-reviewer relevance. To address this issue, we propose a unified model for paper-reviewer matching that jointly considers semantic, topic, and citation factors. To be specific, during training, we instruction-tune a contextualized language model shared across all factors to capture their commonalities and characteristics; during inference, we chain the three factors to enable step-by-step, coarse-to-fine search for qualified reviewers given a submission. Experiments on four datasets (one of which is newly contributed by us) spanning various fields such as machine learning, computer vision, information retrieval, and data mining consistently demonstrate the effectiveness of our proposed Chain-of-Factors model in comparison with state-of-the-art paper-reviewer matching methods and scientific pre-trained language models. Yu Zhang 0044, Yanzhen Shen, Seongku Kang, Xiusi Chen, Bowen Jin, Jiawei Han 0001 |
WWW | 2 |
| 2024 | Seed-Guided Fine-Grained Entity Typing in Science and Engineering DomainsabstractAccurately typing entity mentions from text segments is a fundamental task for various natural language processing applications. Many previous approaches rely on massive human-annotated data to perform entity typing. Nevertheless, collecting such data in highly specialized science and engineering domains (e.g., software engineering and security) can be time-consuming and costly, without mentioning the domain gaps between training and inference data if the model needs to be applied to confidential datasets. In this paper, we study the task of seed-guided fine-grained entity typing in science and engineering domains, which takes the name and a few seed entities for each entity type as the only supervision and aims to classify new entity mentions into both seen and unseen types (i.e., those without seed entities). To solve this problem, we propose SEType which first enriches the weak supervision by finding more entities for each seen type from an unlabeled corpus using the contextualized representations of pre-trained language models. It then matches the enriched entities to unlabeled text to get pseudo-labeled samples and trains a textual entailment model that can make inferences for both seen and unseen types. Extensive experiments on two datasets covering four domains demonstrate the effectiveness of SEType in comparison with various baselines. Code and data are available at: https://github.com/yuzhimanhua/SEType. Yu Zhang 0044, Yunyi Zhang 0001, Yanzhen Shen, Yu Deng 0004, Lucian Popa 0001, Larisa Shwartz, ChengXiang Zhai, Jiawei Han 0001 |
AAAI | 3 |
| 2023 | Weakly Supervised Multi-Label Classification of Full-Text Scientific PapersabstractInstead of relying on human-annotated training samples to build a classifier, weakly supervised scientific paper classification aims to classify papers only using category descriptions (e.g., category names, category-indicative keywords). Existing studies on weakly supervised paper classification are less concerned with two challenges: (1) Papers should be classified into not only coarse-grained research topics but also fine-grained themes, and potentially into multiple themes, given a large and fine-grained label space; and (2) full text should be utilized to complement the paper title and abstract for classification. Moreover, instead of viewing the entire paper as a long linear sequence, one should exploit the structural information such as citation links across papers and the hierarchy of sections and paragraphs in each paper. To tackle these challenges, in this study, we propose FUTEX, a framework that uses the cross-paper network structure and the in-paper hierarchy structure to classify full-text scientific papers under weak supervision. A network-aware contrastive fine-tuning module and a hierarchy-aware aggregation module are designed to leverage the two types of structural signals, respectively. Experiments on two benchmark datasets demonstrate that FUTEX significantly outperforms competitive baselines and is on par with fully supervised classifiers that use 1,000 to 60,000 ground-truth training samples. Yu Zhang 0044, Bowen Jin, Xiusi Chen, Yanzhen Shen, Yunyi Zhang 0001, Yu Meng 0001, Jiawei Han 0001 |
KDD | 4 |