VLDB 2026 Research / reviewers in the wild / expert
Shuyang Cao
dblp:227/2764
· DBLP profile ↗
9ranked-venue papers
7as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 7 first-author · 9 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 79% Question answering and dialogue systems · 19% Deep learning architectures and training · 3% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems
question generation |
1.1 | 2 | 2022 | HIBRIDS: Attention with Hierarchical Biases for Structure-aware Long Document Summarization · ACL (1) 2022 Controllable Open-ended Question Generation with A New Question Type Ontology · ACL/IJCNLP (1) 2021 |
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context understanding |
0.9 | 1 | 2025 | SYNC: A Synthetic Long-Context Understanding Benchmark for Controlled Comparisons of Model Capabilities · EMNLP 2025 |
Natural language and speech › Language models and text generation › evaluation of language models
faithfulness evaluation |
0.7 | 1 | 2023 | BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness Metrics · ACL (1) 2023 |
Natural language and speech › Language models and text generation › large language model evaluation
meta-evaluation |
0.7 | 1 | 2023 | BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness Metrics · ACL (1) 2023 |
Natural language and speech › Language models and text generation
text generation evaluation |
0.7 | 1 | 2023 | BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness Metrics · ACL (1) 2023 |
Natural language and speech › Language models and text generation › text summarization
document summarization |
0.6 | 1 | 2022 | HIBRIDS: Attention with Hierarchical Biases for Structure-aware Long Document Summarization · ACL (1) 2022 |
Natural language and speech › Language models and text generation › text summarization
long document summarization |
0.6 | 1 | 2022 | HIBRIDS: Attention with Hierarchical Biases for Structure-aware Long Document Summarization · ACL (1) 2022 |
Natural language and speech › Language models and text generation › text summarization
abstractive summarization |
0.5 | 1 | 2021 | CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive Summarization · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation › text summarization › abstractive summarization
faithful summarization |
0.5 | 1 | 2021 | CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive Summarization · EMNLP (1) 2021 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.2 | 1 | 2022 | HIBRIDS: Attention with Hierarchical Biases for Structure-aware Long Document Summarization · ACL (1) 2022 |
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability
factuality |
0.1 | 1 | 2021 | CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive Summarization · EMNLP (1) 2021 |
Methods — techniques the papers use, named apart from their topics
synthetic task generation · 0.9minimal pair construction · 0.7hierarchical attention bias injection · 0.6negative sampling · 0.5controllable generation · 0.5contrastive learning · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SYNC: A Synthetic Long-Context Understanding Benchmark for Controlled Comparisons of Model CapabilitiesabstractRecently, researchers have turned to synthetic tasks for evaluating long-context capabilities of large language models (LLMs) , as they offer more flexibility than realistic benchmarks in scaling both input length and dataset size.However, existing synthetic tasks typically target narrow skill sets such as retrieving information from massive input, limiting their ability to comprehensively assess model capabilities.Furthermore, existing benchmarks often pair each task with a different input context, creating confounding factors that prevent fair crosstask comparison.To address these limitations, we introduce SYNC, a new evaluation suite of synthetic tasks spanning domains including graph understanding and translation.Each domain includes three tasks designed to test a wide range of capabilities-from retrieval, to multi-hop tracking, and to global context understanding that that requires chain-of-thought (CoT) reasoning.Crucially, all tasks share the same context, enabling controlled comparisons of model performance.We evaluate 14 LLMs on SYNC and observe substantial performance drops on more challenging tasks, underscoring the benchmark's difficulty.Additional experiments highlight the necessity of CoT reasoning and demonstrate that SYNC poses a robust challenge for future models. Shuyang Cao, Kaijian Zou, Lu Wang 0008 |
EMNLP | 1 |
| 2024 | AWESOME: GPU Memory-constrained Long Document Summarization using Memory Mechanism and Global Salient ContentabstractLong document summarization systems are critical for domains with lengthy and jargonladen text, yet they present significant challenges to researchers and developers with limited computing resources.Existing solutions mainly focus on efficient attentions or divideand-conquer strategies.The former reduces theoretical time complexity, but is still memoryheavy.The latter methods sacrifice global context, leading to uninformative and incoherent summaries.This work aims to leverage the memory-efficient nature of divide-and-conquer methods while preserving global context.Concretely, our framework AWESOME uses two novel mechanisms: (1) External memory mechanisms track previously encoded document segments and their corresponding summaries, to enhance global document understanding and summary coherence.(2) Global salient content is further identified beforehand to augment each document segment to support its summarization.Extensive experiments on diverse genres of text, including government reports, meeting transcripts, screenplays, scientific papers, and novels, show that AWESOME produces summaries with improved informativeness, faithfulness, and coherence than competitive baselines on longer documents, while having a smaller GPU memory footprint.Encoder Shuyang Cao, Lu Wang 0008 |
NAACL-HLT | 1 |
| 2023 | BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness MetricsabstractLiang Ma, Shuyang Cao, Robert L Logan IV, Di Lu, Shihao Ran, Ke Zhang, Joel Tetreault, Alejandro Jaimes. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shuyang Cao, Robert L. Logan IV, Di Lu 0003, Shihao Ran, Ke Zhang 0013, Joel R. Tetreault, Alejandro Jaimes |
ACL (1) | 2 |
| 2022 | HIBRIDS: Attention with Hierarchical Biases for Structure-aware Long Document SummarizationabstractDocument structure is critical for efficient information consumption.However, it is challenging to encode it efficiently into the modern Transformer architecture.In this work, we present HIBRIDS, which injects Hierarchical Biases foR Incorporating Document Structure into the calculation of attention scores.We further present a new task, hierarchical questionsummary generation, for summarizing salient content in the source document into a hierarchy of questions and summaries, where each follow-up question inquires about the content of its parent question-summary pair.We also annotate a new dataset with 6, 153 questionsummary hierarchies labeled on long government reports.Experiment results show that our model produces better question-summary hierarchies than comparisons on both hierarchy quality and content coverage, a finding also echoed by human judges.Additionally, our model improves the generation of longform summaries from lengthy government reports and Wikipedia articles, as measured by ROUGE scores. Shuyang Cao, Lu Wang 0008 |
ACL (1) | 1 |
| 2021 | Controllable Open-ended Question Generation with A New Question Type OntologyabstractShuyang Cao, Lu Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Shuyang Cao, Lu Wang 0008 |
ACL/IJCNLP (1) | 1 |
| 2021 | CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive SummarizationabstractWe study generating abstractive summaries that are faithful and factually consistent with the given articles.A novel contrastive learning formulation is presented, which leverages both reference summaries, as positive training data, and automatically generated erroneous summaries, as negative training data, to train summarization systems that are better at distinguishing between them.We further design four types of strategies for creating negative samples, to resemble errors made commonly by two state-of-the-art models, BART and PEGASUS, found in our new human annotations of summary errors.Experiments on XSum and CNN/Daily Mail show that our contrastive learning framework is robust across datasets and models.It consistently produces more factual summaries than strong comparisons with post error correction, entailmentbased reranking, and unlikelihood training, according to QA-based factuality evaluation.Human judges echo the observation and find that our model summaries correct more errors.REFERENCE: A "rare" short-eared owl found emaciated in Flintshire is now recuperating well, the RSPCA have said.SWAPENT: Flintshire → Bettisfield ⇒ A "rare" short-eared owl found emaciated in Bettisfield is now recuperating well, the RSPCA have said.MASKENT: A "rare" short-eared owl found emaciated in [MASK] is now recuperating well, the RSPCA have said.⇒ A "rare" short-eared owl found emaciated in a field in South Yorkshire is now recuperating well, the RSPCA have said.MASKREL: A "rare" short-eared owl found [MASK] in [MASK] is now recuperating well, the RSPCA have said.⇒ A "rare" short-eared owl found dead in London is now recuperating well, the RSPCA have said.REGENENT: A "rare" short-eared owl found emaciated in ⇒ A "rare" short-eared owl found emaciated in Nottinghamshire is now at a wildlife centre to recover.REGENREL: A "rare" short-eared owl found ⇒ A "rare" short-eared owl found in the grounds of a former coal mine is being cared for by the RSPCA in Somerset.SYSLOWCON: An injured golden owl found in a former coal mine in Lancashire is being cared for by the RSPCA. Shuyang Cao, Lu Wang 0008 |
EMNLP (1) | 1 |
| 2021 | Attention Head Masking for Inference Time Content Selection in Abstractive SummarizationabstractHow can we effectively inform content selection in Transformer-based abstractive summarization models?In this work, we present a simple-yet-effective attention head masking technique, which is applied on encoderdecoder attentions to pinpoint salient content at inference time.Using attention head masking, we are able to reveal the relation between encoder-decoder attentions and content selection behaviors of summarization models.We then demonstrate its effectiveness on three document summarization datasets based on both in-domain and cross-domain settings.Importantly, our models outperform prior state-ofthe-art models on CNN/Daily Mail and New York Times datasets.Moreover, our inferencetime masking technique is also data-efficient, requiring less than 20% of the training samples to outperform BART fine-tuned on the full CNN/DailyMail dataset. Shuyang Cao, Lu Wang 0008 |
NAACL-HLT | 1 |
| 2021 | Inference Time Style Control for SummarizationabstractHow to generate summaries of different styles without requiring corpora in the target styles, or training separate models?We present two novel methods that can be deployed during summary decoding on any pre-trained Transformer-based summarization model.(1) Decoder state adjustment instantly modifies decoder final states with externally trained style scorers, to iteratively refine the output against a target style.(2) Word unit prediction constrains the word usage to impose strong lexical control during generation.In experiments of summarizing with simplicity control, automatic evaluation and human judges both find our models producing outputs in simpler languages while still informative.We also generate news headlines with various ideological leanings, which can be distinguished by humans with a reasonable probability. Shuyang Cao, Lu Wang 0008 |
NAACL-HLT | 1 |
| 2021 | Efficient Attentions for Long Document SummarizationabstractLuyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji, Lu Wang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Luyang Huang, Shuyang Cao, Nikolaus Nova Parulian, Heng Ji 0001, Lu Wang 0008 |
NAACL-HLT | 2 |