VLDB 2026 Research / reviewers in the wild / expert
Kebin Fang
dblp:317/1382
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Language models and text generation · 77% Deep learning architectures and training · 23% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
token masking |
0.7 | 1 | 2023 | TLM: Token-Level Masking for Transformers · EMNLP 2023 |
Machine learning › Deep learning architectures and training › regularization
attention regularization |
0.2 | 1 | 2023 | TLM: Token-Level Masking for Transformers · EMNLP 2023 |
Methods — techniques the papers use, named apart from their topics
drophead · 0.7attention dropout · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | TLM: Token-Level Masking for TransformersabstractStructured dropout approaches, such as attention dropout and DropHead, have been investigated to regularize the multi-head attention mechanism in Transformers.In this paper, we propose a new regularization scheme based on token-level rather than structure-level to reduce overfitting.Specifically, we devise a novel Token-Level Masking (TLM) training strategy for Transformers to regularize the connections of self-attention, which consists of two masking techniques that are effective and easy to implement.The underlying idea is to manipulate the connections between tokens in the multi-head attention via masking, where the networks are forced to exploit partial neighbors' information to produce a meaningful representation.The generality and effectiveness of TLM are thoroughly evaluated via extensive experiments on 4 diversified NLP tasks across 18 datasets, including natural language understanding benchmark GLUE, ChineseGLUE, Chinese Grammatical Error Correction, and data-to-text generation.The results indicate that TLM can consistently outperform attention dropout and DropHead, e.g., it increases by 0.5 points relative to DropHead with BERT-large on GLUE.Moreover, TLM can establish a new record on the data-to-text benchmark Rotowire (18.93 BLEU).Our code will be publicly available at https://github.com/Young1993/tlm. Yangjun Wu, Kebin Fang, Dongxiang Zhang, Hao Zhang 0029, Gang Chen 0001 |
EMNLP | 2 |