Andy Schlaikjer

dblp:151/3253 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › text mining
topic modeling
0.212014
Large-scale high-precision topic modeling on twitter · KDD 2014
Data mining › text mining › text classification
tweet classification
0.212014
Large-scale high-precision topic modeling on twitter · KDD 2014

Methods — techniques the papers use, named apart from their topics

two-stage training · 0.2human computation · 0.2
YearPublicationVenuePosition
2014 Large-scale high-precision topic modeling on twitter
abstract
We are interested in organizing a continuous stream of sparse and noisy texts, known as "tweets", in real time into an ontology of hundreds of topics with measurable and stringently high precision. This inference is performed over a full-scale stream of Twitter data, whose statistical distribution evolves rapidly over time. The implementation in an industrial setting with the potential of affecting and being visible to real users made it necessary to overcome a host of practical challenges. We present a spectrum of topic modeling techniques that contribute to a deployed system. These include non-topical tweet detection, automatic labeled data acquisition, evaluation with human computation, diagnostic and corrective learning and, most importantly, high-precision topic inference. The latter represents a novel two-stage training algorithm for tweet text classification and a close-loop inference mechanism for combining texts with additional sources of information. The resulting system achieves 93% precision at substantial overall coverage.
Shuang-Hong Yang, Alek Kolcz, Andy Schlaikjer, Pankaj Gupta 0002
KDD3