Miriam Connor

dblp:146/3999 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 56% Information retrieval · 44%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › text mining › text classification
web content classification
0.212015
Going In-Depth: Finding Longform on the Web · KDD 2015
Information retrieval
web search
0.212015
Going In-Depth: Finding Longform on the Web · KDD 2015
Data mining › text mining
text classification
0.112015
Going In-Depth: Finding Longform on the Web · KDD 2015

Methods — techniques the papers use, named apart from their topics

language and parse structure features · 0.2
YearPublicationVenuePosition
2015 Going In-Depth: Finding Longform on the Web
abstract
tl;dr: Longform articles are extended, in-depth pieces that often serve as feature stories in newspapers and magazines. In this work, we develop a system to automatically identify longform content across the web. Our novel classifier is highly accurate despite huge variation within longform in terms of topic, voice, and editorial taste. It is also scalable and interpretable, requiring a surprisingly small set of features based only on language and parse structures, length, and document interest. We implement our system at scale and use it to identify a corpus of several million longform documents. Using this corpus, we provide the first web-scale study with quantifiable and measurable information on longform, giving new insight into questions posed by the media on the past and current state of this famed literary medium.
Virginia Smith, Miriam Connor, Isabelle Stanton
KDD2
2014 A Gold Standard Dependency Corpus for English
Natalia Silveira, Timothy Dozat, Marie-Catherine de Marneffe, Samuel R. Bowman, Miriam Connor, John Bauer, Christopher D. Manning
LREC5