VLDB 2026 Research / reviewers in the wild / expert
Simon E. Overell
dblp:42/4629
· DBLP profile ↗
2ranked-venue papers
2as first author
0since 2021 · last 2009
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Web and social media mining · 56% Data mining · 44% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › predictive modeling › classification › multi-label classification
tag-based classification |
0.1 | 1 | 2009 | Classifying tags using open content resources · WSDM 2009 |
Web and social media mining › user-generated content
wikipedia |
0.1 | 1 | 2009 | Classifying tags using open content resources · WSDM 2009 |
Web and social media mining
social tagging |
0.0 | 1 | 2009 | Classifying tags using open content resources · WSDM 2009 |
Methods — techniques the papers use, named apart from their topics
wordnet-based classification · 0.1structural pattern extraction · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2009 | Classifying tags using open content resourcesabstractTagging has emerged as a popular means to annotate on-line objects such as bookmarks, photos and videos. Tags vary in semantic meaning and can describe different aspects of a media object. Tags describe the content of the media as well as locations, dates, people and other associated meta-data. Being able to automatically classify tags into semantic categories allows us to understand better the way users annotate media objects and to build tools for viewing and browsing the media objects. In this paper we present a generic method for classifying tags using third party open content resources, such as Wikipedia and the Open Directory. Our method uses structural patterns that can be extracted from resource meta-data. We describe the implementation of our method on Wikipedia using WordNet categories as our classification schema and ground truth. Two structural patterns found in Wikipedia are used for training and classification: categories and templates. We apply our system to classifying Flickr tags. Compared to a WordNet baseline our method increases the coverage of the Flickr vocabulary by 115%. We can classify many important entities that are not covered by WordNet, such as, London Eye, Big Island, Ronaldinho, geocaching and wii. Simon E. Overell, Börkur Sigurbjörnsson, Roelof van Zwol |
WSDM | 1 |
| 2008 | Using co-occurrence models for placename disambiguationabstractThis paper describes the generation of a model capturing information on how placenames co‐occur together. The advantages of the co‐occurrence model over traditional gazetteers are discussed and the problem of placename disambiguation is presented as a case study. We begin by outlining the problem of ambiguous placenames. We demonstrate how analysis of Wikipedia can be used in the generation of a co‐occurrence model. The accuracy of our model is compared to a handcrafted ground truth; then we evaluate alternative methods of applying this model to the disambiguation of placenames in free text (using the GeoCLEF evaluation forum). We conclude by showing how the inclusion of placenames in both the text and geographic parts of a query provides the maximum mean average precision and outline the benefits of a co‐occurrence model as a data source for the wider field of geographic information retrieval (GIR). Simon E. Overell, Stefan M. Rüger |
Int. J. Geogr. Inf. Sci. | 1 |