Zongru Wan

dblp:198/3340 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › evaluation
effectiveness metrics
0.622018
Further Insights on Drawing Sound Conclusions from Noisy Judgments · ACM Trans. Inf. Syst. 2018
Drawing Sound Conclusions from Noisy Judgments · WWW 2017
Information retrieval
evaluation
0.622018
Further Insights on Drawing Sound Conclusions from Noisy Judgments · ACM Trans. Inf. Syst. 2018
Drawing Sound Conclusions from Noisy Judgments · WWW 2017
Information retrieval › evaluation
statistical significance testing
0.312018
Further Insights on Drawing Sound Conclusions from Noisy Judgments · ACM Trans. Inf. Syst. 2018
Information retrieval › evaluation
relevance judgment
0.312017
Drawing Sound Conclusions from Noisy Judgments · WWW 2017
Information retrieval › evaluation › relevance judgment
crowdsourced relevance judgment
0.112018
Further Insights on Drawing Sound Conclusions from Noisy Judgments · ACM Trans. Inf. Syst. 2018

Methods — techniques the papers use, named apart from their topics

equations · 0.3algorithms · 0.3error correction equations · 0.3crowdsourcing · 0.3
YearPublicationVenuePosition
2018 Further Insights on Drawing Sound Conclusions from Noisy Judgments
abstract
The effectiveness of a search engine is typically evaluated using hand-labeled datasets, where the labels indicate the relevance of documents to queries. Often the number of labels needed is too large to be created by the best annotators, and so less expensive labels (e.g., from crowdsourcing) are used. This introduces errors in the labels, and thus errors in standard effectiveness metrics (such as P@k and DCG). These errors must be taken into consideration when using the metrics. Previous work has approached assessor error by taking aggregates over multiple inexpensive assessors. We take a different approach and introduce equations and algorithms that can adjust the metrics to the values they would have had if there were no annotation errors. This is especially important when two search engines are compared on their metrics. We give examples where one engine appeared to be statistically significantly better than the other, but the effect disappeared after the metrics were corrected for annotation error. In other words, the evidence supporting a statistical difference was illusory and caused by a failure to account for annotation error.
David Goldberg 0001, Andrew Trotman, Wei Min, Zongru Wan
ACM Trans. Inf. Syst.5
2017 Drawing Sound Conclusions from Noisy Judgments
abstract
The quality of a search engine is typically evaluated using hand-labeled data sets, where the labels indicate the relevance of documents to queries. Often the number of labels needed is too large to be created by the best annotators, and so less accurate labels (e.g. from crowdsourcing) must be used. This introduces errors in the labels, and thus errors in standard precision metrics (such as [email protected] and DCG); the lower the quality of the judge, the more errorful the labels, consequently the more inaccurate the metric. We introduce equations and algorithms that can adjust the metrics to the values they would have had if there were no annotation errors.
David Goldberg 0001, Andrew Trotman, Wei Min, Zongru Wan
WWW5