Shuyang Cao

dblp:227/2764 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 7 first-author · 9 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 79% Question answering and dialogue systems · 19% Deep learning architectures and training · 3%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
question generation
1.122022
HIBRIDS: Attention with Hierarchical Biases for Structure-aware Long Document Summarization · ACL (1) 2022
Controllable Open-ended Question Generation with A New Question Type Ontology · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context understanding
0.912025
SYNC: A Synthetic Long-Context Understanding Benchmark for Controlled Comparisons of Model Capabilities · EMNLP 2025
Natural language and speech › Language models and text generation › evaluation of language models
faithfulness evaluation
0.712023
BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness Metrics · ACL (1) 2023
Natural language and speech › Language models and text generation › large language model evaluation
meta-evaluation
0.712023
BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness Metrics · ACL (1) 2023
Natural language and speech › Language models and text generation
text generation evaluation
0.712023
BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness Metrics · ACL (1) 2023
Natural language and speech › Language models and text generation › text summarization
document summarization
0.612022
HIBRIDS: Attention with Hierarchical Biases for Structure-aware Long Document Summarization · ACL (1) 2022
Natural language and speech › Language models and text generation › text summarization
long document summarization
0.612022
HIBRIDS: Attention with Hierarchical Biases for Structure-aware Long Document Summarization · ACL (1) 2022
Natural language and speech › Language models and text generation › text summarization
abstractive summarization
0.512021
CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive Summarization · EMNLP (1) 2021
Natural language and speech › Language models and text generation › text summarization › abstractive summarization
faithful summarization
0.512021
CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive Summarization · EMNLP (1) 2021
Machine learning › Deep learning architectures and training
attention mechanism
0.212022
HIBRIDS: Attention with Hierarchical Biases for Structure-aware Long Document Summarization · ACL (1) 2022
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability
factuality
0.112021
CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive Summarization · EMNLP (1) 2021

Methods — techniques the papers use, named apart from their topics

synthetic task generation · 0.9minimal pair construction · 0.7hierarchical attention bias injection · 0.6negative sampling · 0.5controllable generation · 0.5contrastive learning · 0.5
YearPublicationVenuePosition
2025 SYNC: A Synthetic Long-Context Understanding Benchmark for Controlled Comparisons of Model Capabilities
abstract
Recently, researchers have turned to synthetic tasks for evaluating long-context capabilities of large language models (LLMs) , as they offer more flexibility than realistic benchmarks in scaling both input length and dataset size.However, existing synthetic tasks typically target narrow skill sets such as retrieving information from massive input, limiting their ability to comprehensively assess model capabilities.Furthermore, existing benchmarks often pair each task with a different input context, creating confounding factors that prevent fair crosstask comparison.To address these limitations, we introduce SYNC, a new evaluation suite of synthetic tasks spanning domains including graph understanding and translation.Each domain includes three tasks designed to test a wide range of capabilities-from retrieval, to multi-hop tracking, and to global context understanding that that requires chain-of-thought (CoT) reasoning.Crucially, all tasks share the same context, enabling controlled comparisons of model performance.We evaluate 14 LLMs on SYNC and observe substantial performance drops on more challenging tasks, underscoring the benchmark's difficulty.Additional experiments highlight the necessity of CoT reasoning and demonstrate that SYNC poses a robust challenge for future models.
Shuyang Cao, Kaijian Zou, Lu Wang 0008
EMNLP1
2024 AWESOME: GPU Memory-constrained Long Document Summarization using Memory Mechanism and Global Salient Content
abstract
Long document summarization systems are critical for domains with lengthy and jargonladen text, yet they present significant challenges to researchers and developers with limited computing resources.Existing solutions mainly focus on efficient attentions or divideand-conquer strategies.The former reduces theoretical time complexity, but is still memoryheavy.The latter methods sacrifice global context, leading to uninformative and incoherent summaries.This work aims to leverage the memory-efficient nature of divide-and-conquer methods while preserving global context.Concretely, our framework AWESOME uses two novel mechanisms: (1) External memory mechanisms track previously encoded document segments and their corresponding summaries, to enhance global document understanding and summary coherence.(2) Global salient content is further identified beforehand to augment each document segment to support its summarization.Extensive experiments on diverse genres of text, including government reports, meeting transcripts, screenplays, scientific papers, and novels, show that AWESOME produces summaries with improved informativeness, faithfulness, and coherence than competitive baselines on longer documents, while having a smaller GPU memory footprint.Encoder
Shuyang Cao, Lu Wang 0008
NAACL-HLT1
2023 BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness Metrics
abstract
Liang Ma, Shuyang Cao, Robert L Logan IV, Di Lu, Shihao Ran, Ke Zhang, Joel Tetreault, Alejandro Jaimes. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Shuyang Cao, Robert L. Logan IV, Di Lu 0003, Shihao Ran, Ke Zhang 0013, Joel R. Tetreault, Alejandro Jaimes
ACL (1)2
2022 HIBRIDS: Attention with Hierarchical Biases for Structure-aware Long Document Summarization
abstract
Document structure is critical for efficient information consumption.However, it is challenging to encode it efficiently into the modern Transformer architecture.In this work, we present HIBRIDS, which injects Hierarchical Biases foR Incorporating Document Structure into the calculation of attention scores.We further present a new task, hierarchical questionsummary generation, for summarizing salient content in the source document into a hierarchy of questions and summaries, where each follow-up question inquires about the content of its parent question-summary pair.We also annotate a new dataset with 6, 153 questionsummary hierarchies labeled on long government reports.Experiment results show that our model produces better question-summary hierarchies than comparisons on both hierarchy quality and content coverage, a finding also echoed by human judges.Additionally, our model improves the generation of longform summaries from lengthy government reports and Wikipedia articles, as measured by ROUGE scores.
Shuyang Cao, Lu Wang 0008
ACL (1)1
2021 Controllable Open-ended Question Generation with A New Question Type Ontology
abstract
Shuyang Cao, Lu Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Shuyang Cao, Lu Wang 0008
ACL/IJCNLP (1)1
2021 CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive Summarization
abstract
We study generating abstractive summaries that are faithful and factually consistent with the given articles.A novel contrastive learning formulation is presented, which leverages both reference summaries, as positive training data, and automatically generated erroneous summaries, as negative training data, to train summarization systems that are better at distinguishing between them.We further design four types of strategies for creating negative samples, to resemble errors made commonly by two state-of-the-art models, BART and PEGASUS, found in our new human annotations of summary errors.Experiments on XSum and CNN/Daily Mail show that our contrastive learning framework is robust across datasets and models.It consistently produces more factual summaries than strong comparisons with post error correction, entailmentbased reranking, and unlikelihood training, according to QA-based factuality evaluation.Human judges echo the observation and find that our model summaries correct more errors.REFERENCE: A "rare" short-eared owl found emaciated in Flintshire is now recuperating well, the RSPCA have said.SWAPENT: Flintshire → Bettisfield ⇒ A "rare" short-eared owl found emaciated in Bettisfield is now recuperating well, the RSPCA have said.MASKENT: A "rare" short-eared owl found emaciated in [MASK] is now recuperating well, the RSPCA have said.⇒ A "rare" short-eared owl found emaciated in a field in South Yorkshire is now recuperating well, the RSPCA have said.MASKREL: A "rare" short-eared owl found [MASK] in [MASK] is now recuperating well, the RSPCA have said.⇒ A "rare" short-eared owl found dead in London is now recuperating well, the RSPCA have said.REGENENT: A "rare" short-eared owl found emaciated in ⇒ A "rare" short-eared owl found emaciated in Nottinghamshire is now at a wildlife centre to recover.REGENREL: A "rare" short-eared owl found ⇒ A "rare" short-eared owl found in the grounds of a former coal mine is being cared for by the RSPCA in Somerset.SYSLOWCON: An injured golden owl found in a former coal mine in Lancashire is being cared for by the RSPCA.
Shuyang Cao, Lu Wang 0008
EMNLP (1)1
2021 Attention Head Masking for Inference Time Content Selection in Abstractive Summarization
abstract
How can we effectively inform content selection in Transformer-based abstractive summarization models?In this work, we present a simple-yet-effective attention head masking technique, which is applied on encoderdecoder attentions to pinpoint salient content at inference time.Using attention head masking, we are able to reveal the relation between encoder-decoder attentions and content selection behaviors of summarization models.We then demonstrate its effectiveness on three document summarization datasets based on both in-domain and cross-domain settings.Importantly, our models outperform prior state-ofthe-art models on CNN/Daily Mail and New York Times datasets.Moreover, our inferencetime masking technique is also data-efficient, requiring less than 20% of the training samples to outperform BART fine-tuned on the full CNN/DailyMail dataset.
Shuyang Cao, Lu Wang 0008
NAACL-HLT1
2021 Inference Time Style Control for Summarization
abstract
How to generate summaries of different styles without requiring corpora in the target styles, or training separate models?We present two novel methods that can be deployed during summary decoding on any pre-trained Transformer-based summarization model.(1) Decoder state adjustment instantly modifies decoder final states with externally trained style scorers, to iteratively refine the output against a target style.(2) Word unit prediction constrains the word usage to impose strong lexical control during generation.In experiments of summarizing with simplicity control, automatic evaluation and human judges both find our models producing outputs in simpler languages while still informative.We also generate news headlines with various ideological leanings, which can be distinguished by humans with a reasonable probability.
Shuyang Cao, Lu Wang 0008
NAACL-HLT1
2021 Efficient Attentions for Long Document Summarization
abstract
Luyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji, Lu Wang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Luyang Huang, Shuyang Cao, Nikolaus Nova Parulian, Heng Ji 0001, Lu Wang 0008
NAACL-HLT2