EDBT 2026 Demo / reviewers in the wild / expert
Andy Schlaikjer
dblp:151/3253
· DBLP profile ↗
1ranked-venue papers
0as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › text mining
topic modeling |
0.2 | 1 | 2014 | Large-scale high-precision topic modeling on twitter · KDD 2014 |
Data mining › text mining › text classification
tweet classification |
0.2 | 1 | 2014 | Large-scale high-precision topic modeling on twitter · KDD 2014 |
Methods — techniques the papers use, named apart from their topics
two-stage training · 0.2human computation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Large-scale high-precision topic modeling on twitterabstractWe are interested in organizing a continuous stream of sparse and noisy texts, known as "tweets", in real time into an ontology of hundreds of topics with measurable and stringently high precision. This inference is performed over a full-scale stream of Twitter data, whose statistical distribution evolves rapidly over time. The implementation in an industrial setting with the potential of affecting and being visible to real users made it necessary to overcome a host of practical challenges. We present a spectrum of topic modeling techniques that contribute to a deployed system. These include non-topical tweet detection, automatic labeled data acquisition, evaluation with human computation, diagnostic and corrective learning and, most importantly, high-precision topic inference. The latter represents a novel two-stage training algorithm for tweet text classification and a close-loop inference mechanism for combining texts with additional sources of information. The resulting system achieves 93% precision at substantial overall coverage. Shuang-Hong Yang, Alek Kolcz, Andy Schlaikjer, Pankaj Gupta 0002 |
KDD | 3 |