Kaijian Zou

dblp:359/5930 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 80% Trustworthy machine learning · 15% Information extraction and text analysis · 5%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
in-context learning
0.912025
On Many-Shot In-Context Learning for Long-Context Evaluation · ACL (1) 2025
Natural language and speech › Language models and text generation › large language model evaluation › capability evaluation
long-context language model evaluation
0.912025
On Many-Shot In-Context Learning for Long-Context Evaluation · ACL (1) 2025
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context understanding
0.912025
SYNC: A Synthetic Long-Context Understanding Benchmark for Controlled Comparisons of Model Capabilities · EMNLP 2025
Natural language and speech › Language models and text generation › in-context learning
many-shot in-context learning
0.912025
On Many-Shot In-Context Learning for Long-Context Evaluation · ACL (1) 2025
Machine learning › Trustworthy machine learning › fairness › bias evaluation › bias detection
media bias detection
0.712023
All Things Considered: Detecting Partisan Events from News Media with Cross-Article Comparison · EMNLP 2023

Methods — techniques the papers use, named apart from their topics

synthetic task generation · 0.9retrieval · 0.9benchmark construction · 0.9latent variable model · 0.7
YearPublicationVenuePosition
2025 On Many-Shot In-Context Learning for Long-Context Evaluation
abstract
Many-shot in-context learning (ICL) has emerged as a unique setup to both utilize and test the ability of large language models to handle long context.This paper delves into long-context language model (LCLM) evaluation through many-shot ICL.We first ask: what types of ICL tasks benefit from additional demonstrations, and how effective are they in evaluating LCLMs?We find that classification and summarization tasks show performance improvements with additional demonstrations, while translation and reasoning tasks do not exhibit clear trends.Next, we investigate the extent to which different tasks necessitate retrieval versus global context understanding.We develop metrics to categorize ICL tasks into two groups: (i) similar-sample learning (SSL): tasks where retrieval of the most similar examples is sufficient for good performance, and (ii) all-sample learning (ASL): tasks that necessitate a deeper comprehension of all examples in the prompt.Lastly, we introduce a new many-shot ICL benchmark built on existing ICL tasks, MANYICLBENCH, to characterize model's ability on both fronts and benchmark 12 LCLMs using MANYICLBENCH.We find that while state-of-the-art models demonstrate good performance up to 64k tokens in SSL tasks, many models experience significant performance drops at only 16k tokens in ASL tasks.
Kaijian Zou, Muhammad Khalifa, Lu Wang 0008
ACL (1)1
2025 SYNC: A Synthetic Long-Context Understanding Benchmark for Controlled Comparisons of Model Capabilities
abstract
Recently, researchers have turned to synthetic tasks for evaluating long-context capabilities of large language models (LLMs) , as they offer more flexibility than realistic benchmarks in scaling both input length and dataset size.However, existing synthetic tasks typically target narrow skill sets such as retrieving information from massive input, limiting their ability to comprehensively assess model capabilities.Furthermore, existing benchmarks often pair each task with a different input context, creating confounding factors that prevent fair crosstask comparison.To address these limitations, we introduce SYNC, a new evaluation suite of synthetic tasks spanning domains including graph understanding and translation.Each domain includes three tasks designed to test a wide range of capabilities-from retrieval, to multi-hop tracking, and to global context understanding that that requires chain-of-thought (CoT) reasoning.Crucially, all tasks share the same context, enabling controlled comparisons of model performance.We evaluate 14 LLMs on SYNC and observe substantial performance drops on more challenging tasks, underscoring the benchmark's difficulty.Additional experiments highlight the necessity of CoT reasoning and demonstrate that SYNC poses a robust challenge for future models.
Shuyang Cao, Kaijian Zou, Lu Wang 0008
EMNLP2
2023 All Things Considered: Detecting Partisan Events from News Media with Cross-Article Comparison
abstract
Public opinion is shaped by the information news media provide, and that information in turn may be shaped by the ideological preferences of media outlets.But while much attention has been devoted to media bias via overt ideological language or topic selection, a more unobtrusive way in which the media shape opinion is via the strategic inclusion or omission of partisan events that may support one side or the other.We develop a latent variable-based framework to predict the ideology of news articles by comparing multiple articles on the same story and identifying partisan events whose inclusion or omission reveals ideology.Our experiments first validate the existence of partisan event selection, and then show that article alignment and cross-document comparison detect partisan events and article ideology better than competitive baselines.Our results reveal the high-level form of media bias, which is present even among mainstream media with strong norms of objectivity and nonpartisanship.
Yujian Liu, Xinliang Frederick Zhang, Kaijian Zou, Ruihong Huang, Nick Beauchamp, Lu Wang 0008
EMNLP3