Chuyan Zhou

dblp:437/7443 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 1 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › linguistic generalization
syntactic generalization
1.012026
GiLT: Augmenting Transformer Language Models with Dependency Graphs · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

dependency graph infusion · 1.0attention weight modulation · 1.0
YearPublicationVenuePosition
2026 GiLT: Augmenting Transformer Language Models with Dependency Graphs
abstract
Augmenting Transformers with linguistic structures effectively enhances the syntactic generalization performance of language models.Previous work in this direction focuses on syntactic tree structures of languages, in particular constituency tree structures.We propose Graph-Infused Layers Transformer Language Model (GiLT) which leverages dependency graphs for augmenting Transformer language models.Unlike most previous work, GiLT does not insert extra structural tokens in language modeling; instead, it injects structural information into language modeling by modulating attention weights in the Transformer with features extracted from the dependency graph that is incrementally constructed along with token prediction.In our experiments, GiLT with semantic dependency graphs achieves better syntactic generalization while maintaining competitive perplexity in comparison with Transformer language model baselines.In addition, GiLT can be finetuned from a pretrained language model to achieve improved downstream task performance.Our code is released at https://github.com/cookie-pie-oops/GiLT-LM.
Yida Zhao, Chuyan Zhou, Kewei Tu
ACL (1)3