Tatiana Shamardina

dblp:321/9874 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2022
0009-0008-9033-6646ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Information extraction and text analysis · 87% Knowledge representation and reasoning · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › error detection
grammatical error detection
0.612022
RuCoLA: Russian Corpus of Linguistic Acceptability · EMNLP 2022
Natural language and speech › Information extraction and text analysis
linguistic acceptability
0.612022
RuCoLA: Russian Corpus of Linguistic Acceptability · EMNLP 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge management › knowledge evaluation
linguistic knowledge evaluation
0.212022
RuCoLA: Russian Corpus of Linguistic Acceptability · EMNLP 2022

Methods — techniques the papers use, named apart from their topics

generative model · 0.6fine-tuning · 0.6
YearPublicationVenuePosition
2022 RuCoLA: Russian Corpus of Linguistic Acceptability
abstract
Linguistic acceptability (LA) attracts the attention of the research community due to its many uses, such as testing the grammatical knowledge of language models and filtering implausible texts with acceptability classifiers.However, the application scope of LA in languages other than English is limited due to the lack of high-quality resources.To this end, we introduce the Russian Corpus of Linguistic Acceptability (RuCoLA), built from the ground up under the well-established binary LA approach.RuCoLA consists of 9.8k indomain sentences from linguistic publications and 3.6k out-of-domain sentences produced by generative models.The out-of-domain set is created to facilitate the practical use of acceptability for improving language generation.Our paper describes the data collection protocol and presents a fine-grained analysis of acceptability classification experiments with a range of baseline approaches.In particular, we demonstrate that the most widely used language models still fall behind humans by a large margin, especially when detecting morphological and semantic errors.We release RuCoLA, the code of experiments, and a public leaderboard 1 to assess the linguistic competence of language models for Russian.
Vladislav Mikhailov, Tatiana Shamardina, Max Ryabinin, Alena Pestova, Ivan Smurov, Ekaterina Artemova
EMNLP2