VLDB 2026 Research / reviewers in the wild / expert
Kaijian Zou
dblp:359/5930
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 80% Trustworthy machine learning · 15% Information extraction and text analysis · 5% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | On Many-Shot In-Context Learning for Long-Context Evaluation · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model evaluation › capability evaluation
long-context language model evaluation |
0.9 | 1 | 2025 | On Many-Shot In-Context Learning for Long-Context Evaluation · ACL (1) 2025 |
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context understanding |
0.9 | 1 | 2025 | SYNC: A Synthetic Long-Context Understanding Benchmark for Controlled Comparisons of Model Capabilities · EMNLP 2025 |
Natural language and speech › Language models and text generation › in-context learning
many-shot in-context learning |
0.9 | 1 | 2025 | On Many-Shot In-Context Learning for Long-Context Evaluation · ACL (1) 2025 |
Machine learning › Trustworthy machine learning › fairness › bias evaluation › bias detection
media bias detection |
0.7 | 1 | 2023 | All Things Considered: Detecting Partisan Events from News Media with Cross-Article Comparison · EMNLP 2023 |
Methods — techniques the papers use, named apart from their topics
synthetic task generation · 0.9retrieval · 0.9benchmark construction · 0.9latent variable model · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On Many-Shot In-Context Learning for Long-Context EvaluationabstractMany-shot in-context learning (ICL) has emerged as a unique setup to both utilize and test the ability of large language models to handle long context.This paper delves into long-context language model (LCLM) evaluation through many-shot ICL.We first ask: what types of ICL tasks benefit from additional demonstrations, and how effective are they in evaluating LCLMs?We find that classification and summarization tasks show performance improvements with additional demonstrations, while translation and reasoning tasks do not exhibit clear trends.Next, we investigate the extent to which different tasks necessitate retrieval versus global context understanding.We develop metrics to categorize ICL tasks into two groups: (i) similar-sample learning (SSL): tasks where retrieval of the most similar examples is sufficient for good performance, and (ii) all-sample learning (ASL): tasks that necessitate a deeper comprehension of all examples in the prompt.Lastly, we introduce a new many-shot ICL benchmark built on existing ICL tasks, MANYICLBENCH, to characterize model's ability on both fronts and benchmark 12 LCLMs using MANYICLBENCH.We find that while state-of-the-art models demonstrate good performance up to 64k tokens in SSL tasks, many models experience significant performance drops at only 16k tokens in ASL tasks. Kaijian Zou, Muhammad Khalifa, Lu Wang 0008 |
ACL (1) | 1 |
| 2025 | SYNC: A Synthetic Long-Context Understanding Benchmark for Controlled Comparisons of Model CapabilitiesabstractRecently, researchers have turned to synthetic tasks for evaluating long-context capabilities of large language models (LLMs) , as they offer more flexibility than realistic benchmarks in scaling both input length and dataset size.However, existing synthetic tasks typically target narrow skill sets such as retrieving information from massive input, limiting their ability to comprehensively assess model capabilities.Furthermore, existing benchmarks often pair each task with a different input context, creating confounding factors that prevent fair crosstask comparison.To address these limitations, we introduce SYNC, a new evaluation suite of synthetic tasks spanning domains including graph understanding and translation.Each domain includes three tasks designed to test a wide range of capabilities-from retrieval, to multi-hop tracking, and to global context understanding that that requires chain-of-thought (CoT) reasoning.Crucially, all tasks share the same context, enabling controlled comparisons of model performance.We evaluate 14 LLMs on SYNC and observe substantial performance drops on more challenging tasks, underscoring the benchmark's difficulty.Additional experiments highlight the necessity of CoT reasoning and demonstrate that SYNC poses a robust challenge for future models. Shuyang Cao, Kaijian Zou, Lu Wang 0008 |
EMNLP | 2 |
| 2023 | All Things Considered: Detecting Partisan Events from News Media with Cross-Article ComparisonabstractPublic opinion is shaped by the information news media provide, and that information in turn may be shaped by the ideological preferences of media outlets.But while much attention has been devoted to media bias via overt ideological language or topic selection, a more unobtrusive way in which the media shape opinion is via the strategic inclusion or omission of partisan events that may support one side or the other.We develop a latent variable-based framework to predict the ideology of news articles by comparing multiple articles on the same story and identifying partisan events whose inclusion or omission reveals ideology.Our experiments first validate the existence of partisan event selection, and then show that article alignment and cross-document comparison detect partisan events and article ideology better than competitive baselines.Our results reveal the high-level form of media bias, which is present even among mainstream media with strong norms of objectivity and nonpartisanship. Yujian Liu, Xinliang Frederick Zhang, Kaijian Zou, Ruihong Huang, Nick Beauchamp, Lu Wang 0008 |
EMNLP | 3 |