Sihan Song

dblp:352/3018 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Information extraction and text analysis · 91% Vision and language · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › named entity recognition
low-resource named entity recognition
0.812024
RoPDA: Robust Prompt-Based Data Augmentation for Low-Resource Named Entity Recognition · AAAI 2024
Natural language and speech › Information extraction and text analysis
named entity recognition
0.812024
RoPDA: Robust Prompt-Based Data Augmentation for Low-Resource Named Entity Recognition · AAAI 2024
Natural language and speech › Information extraction and text analysis
prompt-based data augmentation
0.812024
RoPDA: Robust Prompt-Based Data Augmentation for Low-Resource Named Entity Recognition · AAAI 2024
Computer vision › Vision and language › vision-language model
prompt learning
0.212024
RoPDA: Robust Prompt-Based Data Augmentation for Low-Resource Named Entity Recognition · AAAI 2024

Methods — techniques the papers use, named apart from their topics

self-consistency filtering · 0.8mixup · 0.8continuous prompt tuning · 0.8
YearPublicationVenuePosition
2024 RoPDA: Robust Prompt-Based Data Augmentation for Low-Resource Named Entity Recognition
abstract
Data augmentation has been widely used in low-resource NER tasks to tackle the problem of data sparsity. However, previous data augmentation methods have the disadvantages of disrupted syntactic structures, token-label mismatch, and requirement for external knowledge or manual effort. To address these issues, we propose Robust Prompt-based Data Augmentation (RoPDA) for low-resource NER. Based on pre-trained language models (PLMs) with continuous prompt, RoPDA performs entity augmentation and context augmentation through five fundamental augmentation operations to generate label-flipping and label-preserving examples. To optimize the utilization of the augmented samples, we present two techniques: self-consistency filtering and mixup. The former effectively eliminates low-quality samples with a bidirectional mask, while the latter prevents performance degradation arising from the direct utilization of labelflipping samples. Extensive experiments on three popular benchmarks from different domains demonstrate that RoPDA significantly improves upon strong baselines, and also outperforms state-of-the-art semi-supervised learning methods when unlabeled data is included.
Sihan Song, Furao Shen, Jian Zhao 0013
AAAI1