Ekaterina Chernyak

dblp:147/6699 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 50% Information retrieval · 50%
Artificial intelligence
1 paper
Knowledge representation and reasoning · 100%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › predictive modeling › classification
multi-label classification
0.212015
An Approach to the Problem of Annotation of Research Publications · WSDM 2015
Information retrieval › evaluation › effectiveness metrics
relevance measure
0.212015
An Approach to the Problem of Annotation of Research Publications · WSDM 2015

Methods — techniques the papers use, named apart from their topics

suffix tree · 0.4TF-IDF · 0.4BM25 · 0.4
YearPublicationVenuePosition
2018 Relation Extraction Datasets in the Digital Humanities Domain and Their Evaluation with Word Embeddings
Gerhard Wohlgenannt, Ekaterina Chernyak, Dmitry I. Ilvovsky, Ariadna Barinova, Dmitry Mouromtsev
CICLing (1)2
2015 An Approach to the Problem of Annotation of Research Publications
abstract
An approach to multiple labelling research papers is explored. We develop techniques for annotating/labeling research papers in informatics and computer sciences with key phrases taken from the ACM Computing Classification System. The techniques utilize a phrase-to-text relevance measure so that only those phrases that are most relevant go to the annotation. Three phrase-to-text relevance measures are experimentally compared in this setting. The measures are: (a) cosine relevance score between conventional vector space representations of the texts coded with tf-idf weighting; (b) popular characteristic of probability of term generation BM25; and (c) an in-house characteristic of conditional probability of symbols averaged over matching fragments in suffix trees representing texts and phrases, CPAMF. In an experiment conducted over a set of texts published in journals of the ACM and manually annotated by their authors, CPAMF outperforms both the cosine measure and BM25 by a wide margin.
Ekaterina Chernyak
WSDM1