Sterling Stuart Stein

dblp:59/2695 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
0since 2021 · last 2006
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 1 first-authorArtificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data mining · 51% Information retrieval · 49%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › text mining
text classification
0.112006
The effect of OCR errors on stylistic text classification · SIGIR 2006
Information retrieval › text analysis › stylometry
authorship attribution
0.012003
Style mining of electronic messages for multiple authorship discrimination: first results · KDD 2003
Information retrieval › text analysis
stylometry
0.012003
Style mining of electronic messages for multiple authorship discrimination: first results · KDD 2003
Data mining
text mining
0.012003
Style mining of electronic messages for multiple authorship discrimination: first results · KDD 2003
Information retrieval
document processing
0.012006
The effect of OCR errors on stylistic text classification · SIGIR 2006

Methods — techniques the papers use, named apart from their topics

stylistic classification · 0.1winnow algorithm · 0.0multiclass classification · 0.0
YearPublicationVenuePosition
2006 The effect of OCR errors on stylistic text classification
abstract
Recently, interest is growing in non-topical text classification tasks such as genre classification, sentiment analysis, and authorship profiling. We study to what extent OCR errors affect stylistic text classification from scanned documents. We find that even a relatively high level of errors in the OCRed documents does not substantially affect stylistic classification accuracy.
Sterling Stuart Stein, Shlomo Argamon, Ophir Frieder
SIGIR1
2003 Style mining of electronic messages for multiple authorship discrimination: first results
abstract
This paper considers the use of computational stylistics for performing authorship attribution of electronic messages, addressing categorization problems with as many as 20 different classes (authors). Effective stylistic characterization of text is potentially useful for a variety of tasks, as language style contains cues regarding the authorship, purpose, and mood of the text, all of which would be useful adjuncts to information retrieval or knowledge-management tasks. We focus here on the problem of determining the author of an anonymous message, based only on the message text. Several multiclass variants of the Winnow algorithm were applied to a vector representation of the message texts to learn models for discriminating different authors. We present results comparing the classification accuracy of the different approaches. The results show that stylistic models can be accurately learned to determine an author's identity.
Shlomo Argamon, Marin Saric, Sterling Stuart Stein
KDD3