Börkur Sigurbjörnsson

dblp:32/564 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
0since 2021 · last 2010
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 13 · 3 first-authorArtificial intelligence and machine learning · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
8 papers
Information retrieval · 70% Web and social media mining · 12% Recommender systems · 6%

Topics — the 22 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › search interfaces
faceted search
0.222010
Faceted exploration of image search results · WWW 2010
Machine learned ranking of entity facets · SIGIR 2010
Information retrieval › document retrieval › structured document retrieval
XML retrieval
0.242006
Articulating information needs in XML query languages · ACM Trans. Inf. Syst. 2006
Multiple sources of evidence for XML retrieval · SIGIR 2004
Length normalization in XML retrieval · SIGIR 2004
Information retrieval
retrieval models
0.232010
Faceted exploration of image search results · WWW 2010
Articulating information needs in XML query languages · ACM Trans. Inf. Syst. 2006
Length normalization in XML retrieval · SIGIR 2004
Information retrieval › document retrieval
structured document retrieval
0.232006
Articulating information needs in XML query languages · ACM Trans. Inf. Syst. 2006
Length normalization in XML retrieval · SIGIR 2004
XML retrieval: what to retrieve? · SIGIR 2003
Information retrieval › ranking › search ranking
facet ranking
0.112010
Faceted exploration of image search results · WWW 2010
Information retrieval
image retrieval
0.112010
Faceted exploration of image search results · WWW 2010
Information retrieval › ranking
learning to rank
0.112010
Machine learned ranking of entity facets · SIGIR 2010
Information retrieval
search engines
0.112010
Faceted exploration of image search results · WWW 2010
Data mining › predictive modeling › classification › multi-label classification
tag-based classification
0.112009
Classifying tags using open content resources · WSDM 2009
Web and social media mining › user-generated content
wikipedia
0.112009
Classifying tags using open content resources · WSDM 2009
Recommender systems
tag recommendation
0.112008
Flickr tag recommendation based on collective knowledge · WWW 2008
Database theory
expressive power
0.112006
Articulating information needs in XML query languages · ACM Trans. Inf. Syst. 2006
Web and social media mining
social tagging
0.122010
Faceted exploration of image search results · WWW 2010
Classifying tags using open content resources · WSDM 2009
Information retrieval › retrieval models › term weighting
document length normalization
0.012004
Length normalization in XML retrieval · SIGIR 2004
Query processing and optimization › complex query processing
hybrid queries
0.012004
Multiple sources of evidence for XML retrieval · SIGIR 2004
Information retrieval › relevance feedback
click feedback
0.012010
Machine learned ranking of entity facets · SIGIR 2010
Recommender systems
click-through rate prediction
0.012010
Machine learned ranking of entity facets · SIGIR 2010
Web and social media mining
user-generated content
0.012010
Faceted exploration of image search results · WWW 2010
Web and social media mining
social media analysis
0.012008
Flickr tag recommendation based on collective knowledge · WWW 2008
Data models and query languages
XML query languages
0.012006
Articulating information needs in XML query languages · ACM Trans. Inf. Syst. 2006
Information retrieval › retrieval models
language model
0.012004
Length normalization in XML retrieval · SIGIR 2004
Information retrieval
document retrieval
0.012003
XML retrieval: what to retrieve? · SIGIR 2003

Methods — techniques the papers use, named apart from their topics

statistical analysis · 0.1query log analysis · 0.1gradient boosted decision trees · 0.1click-over-expected-clicks · 0.1wordnet-based classification · 0.1structural pattern extraction · 0.1tag co-occurrence analysis · 0.1collaborative filtering · 0.1user study · 0.1mathematical modeling · 0.1
YearPublicationVenuePosition
2010 Machine learned ranking of entity facets
abstract
The research described in this paper forms the backbone of a service that enables the faceted search experience of the Yahoo! search engine. We introduce an approach for a machine learned ranking of entity facets based on user click feedback and features extracted from three different ranking sources. The objective of the learned model is to predict the click-through rate on an entity facet. In an empirical evaluation we compare the performance of gradient boosted decision trees (GBDT) against a linear combination of features on two different click feedback models using the raw click-through rate (CTR), and click over expected clicks (COEC). The results show a significant improvement in retrieval performance, in terms of discounted cumulated gain, when ranking entity facets with GBDT trained on the COEC model. Most notably this is true when evaluated against the CTR test set.
Roelof van Zwol, Lluís Garcia Pueyo, Mridul Muralidharan, Börkur Sigurbjörnsson
SIGIR4
2010 Faceted exploration of image search results
abstract
This paper describes MediaFaces, a system that enables faceted exploration of media collections. The system processes semi-structured information sources to extract objects and facets, e.g. the relationships between two objects. Next, we rank the facets based on a statistical analysis of image search query logs, and the tagging behaviour of users annotating photos in Flickr. For a given object of interest, we can then retrieve the top-k most relevant facets and present them to the user. The system is currently deployed in production by Yahoo!'s image search engine1. We present the system architecture, its main components, and the application of the system as part of the image search experience.
Roelof van Zwol, Börkur Sigurbjörnsson, Ramu Adapala, Lluís Garcia Pueyo, Abhinav Katiyar, Kaushal Kurapati, Mridul Muralidharan, Sudar Muthu, Vanessa Murdock 0001, Polly Ng, Anand Ramani, Anuj Sahai, Sriram Thiru Sathish, Hari Vasudev, Upendra Vuyyuru
WWW2
2009 Compressing tags to find interesting media groups
abstract
On photo sharing websites like Flickr and Zooomr, users are offered the possibility to assign tags to their uploaded pictures. Using these tags to find interesting groups of semantically related pictures in the result set of a given query is a problem with obvious applications. We analyse this problem from a Minimum Description Length (MDL) perspective and develop an algorithm that finds the most interesting groups. The method is based on Krimp, which finds small sets of patterns that characterise the data using compression. These patterns are sets of tags, often assigned together to photos. The better a database compresses, the more structure it contains and thus the more homogeneous it is. Following this observation we devise a compression-based measure. Our experiments on Flickr data show that the most interesting and homogeneous groups are found. We show extensive examples and compare to clusterings on the Flickr website.
Matthijs van Leeuwen, Francesco Bonchi, Börkur Sigurbjörnsson, Arno Siebes
CIKM3
2009 Classifying tags using open content resources
abstract
Tagging has emerged as a popular means to annotate on-line objects such as bookmarks, photos and videos. Tags vary in semantic meaning and can describe different aspects of a media object. Tags describe the content of the media as well as locations, dates, people and other associated meta-data. Being able to automatically classify tags into semantic categories allows us to understand better the way users annotate media objects and to build tools for viewing and browsing the media objects. In this paper we present a generic method for classifying tags using third party open content resources, such as Wikipedia and the Open Directory. Our method uses structural patterns that can be extracted from resource meta-data. We describe the implementation of our method on Wikipedia using WordNet categories as our classification schema and ground truth. Two structural patterns found in Wikipedia are used for training and classification: categories and templates. We apply our system to classifying Flickr tags. Compared to a WordNet baseline our method increases the coverage of the Flickr vocabulary by 115%. We can classify many important entities that are not covered by WordNet, such as, London Eye, Big Island, Ronaldinho, geocaching and wii.
Simon E. Overell, Börkur Sigurbjörnsson, Roelof van Zwol
WSDM2
2008 Flickr tag recommendation based on collective knowledge
abstract
Online photo services such as Flickr and Zooomr allow users to share their photos with family, friends, and the online community at large. An important facet of these services is that users manually annotate their photos using so called tags, which describe the contents of the photo or provide additional contextual and semantical information. In this paper we investigate how we can assist users in the tagging phase. The contribution of our research is twofold. We analyse a representative snapshot of Flickr and present the results by means of a tag characterisation focussing on how users tags photos and what information is contained in the tagging. Based on this analysis, we present and evaluate tag recommendation strategies to support the user in the photo annotation task by recommending a set of tags that can be added to the photo. The results of the empirical evaluation show that we can effectively recommend relevant tags for a variety of photos with different levels of exhaustiveness of original tagging.
Börkur Sigurbjörnsson, Roelof van Zwol
WWW1
2006 Articulating information needs in XML query languages
abstract
Document-centric XML is a mixture of text and structure. With the increased availability of document-centric XML documents comes a need for query facilities in which both structural constraints and constraints on the content of the documents can be expressed. How does the expressiveness of languages for querying XML documents help users to express their information needs? We address this question from both an experimental and a theoretical point of view. Our experimental analysis compares a structure-ignorant with a structure-aware retrieval approach using the test suite of the INEX XML Retrieval Evaluation Initiative. Theoretically, we create two mathematical models of users' knowledge of a set of documents and define query languages which exactly fit these models. One of these languages corresponds to an XML version of fielded search, the other to the INEX query language.Our main experimental findings are: First, while structure is used in varying degrees of complexity, two-thirds of the queries can be expressed in a fielded-search-like format which does not use the hierarchical structure of the documents. Second, three-quarters of the queries use constraints on the context of the elements to be returned; these contextual constraints cannot be captured by ordinary keyword queries. Third, structure is used as a search hint, and not as a strict requirement, when judged against the underlying information need. Fourth, the use of structure in queries functions as a precision enhancing device.
Jaap Kamps, Maarten Marx, Maarten de Rijke, Börkur Sigurbjörnsson
ACM Trans. Inf. Syst.4
2005 Structured queries in XML retrieval
abstract
Document-centric XML is a mixture of text and structure. With the increased availability of document-centric XML content comes a need for query facilities in which both structural constraints and constraints on the content of the documents can be expressed. How does the expressiveness of languages for querying XML documents help users to express their information needs? We address this question from both an experimental and a theoretical point of view. Our experimental analysis compares a structure-ignorant with a structure-aware retrieval approach using the test-suite of the 2004 edition of the INEX XML retrieval evaluation initiative. Theoretically, we create mathematical models of users' knowledge of a set of documents and define query languages which exactly fit these models. One of these languages corresponds to an XML version of fielded search, the other to the INEX query language. Our main findings are: First, while structure is used in varying degrees of complexity, over half of the queries can be expressed in a fielded-search like format which does not use the hierarchical structure of the documents. Second, structure is used as a search hint, and not a strict requirement, when judged against the underlying information need. Third, the use of structure in queries functions as a precision enhancing device.
Jaap Kamps, Maarten Marx, Maarten de Rijke, Börkur Sigurbjörnsson
CIKM4
2005 The Importance of Length Normalization for XML Retrieval
Jaap Kamps, Maarten de Rijke, Börkur Sigurbjörnsson
Inf. Retr.3
2004 Processing content-oriented XPath queries
abstract
Document-centric XML collections contain text-rich documents, marked up with XML tags that add lightweight semantics to the text. Querying such collections calls for a hybrid query language: the text-rich nature of the documents suggests a content-oriented (IR) approach, while the mark-up allows users to add structural constraints to their IR queries. Hybrid queries tend to be more expressive, which should lead---in principle---to better retrieval performance. In practice, the processing of these hybrid queries within an IR systems turns out to be far from trivial, because a delicate balance between structural and content information needs to be sought. We propose an approach to processing such hybrid content-and-structure queries that decomposes a query into multiple content-only queries whose results are then combined in ways determined by the structural constraints of the original query. We evaluate our methods using the INEX 2003 test-suite, and show (1) that effective ways of processing of content-oriented XPath queries are non-trivial, (2) that there are differences in the effectiveness for different topics types, but (3) that with appropriate processing methods retrieval effectiveness can improve.
Börkur Sigurbjörnsson, Jaap Kamps, Maarten de Rijke
CIKM1
2004 Length normalization in XML retrieval
abstract
Abstract. XML retrieval is a departure from standard document retrieval in which each individual XML element, ranging from italicized words or phrases to full blown articles, is a retrievable unit. The distribution of XML element lengths is unlike what we usually observe in standard document collections, prompting us to revisit the issue of document length normalization. We perform a comparative analysis of arbitrary elements versus relevant elements, and show the importance of element length as a parameter for XML retrieval. Within the language modeling framework, we investigate a range of techniques that deal with length either directly or indirectly. We observe a length-bias introduced by the amount of smoothing, and show the importance of extreme length bias for XML retrieval. We also show that simply removing shorter elements from the index (by introducing a cut-off value) does not create an appropriate element length normalization. Even after restricting the minimal size of XML elements occurring in the index, the importance of an extreme explicit length bias remains. Keywords: XML retrieval, language models, length normalization, smoothing
Jaap Kamps, Maarten de Rijke, Börkur Sigurbjörnsson
SIGIR3
2004 Multiple sources of evidence for XML retrieval
abstract
Document-centric XML collections contain text-rich documents, marked up with XML tags. The tags add lightweight semantics to the text. Querying such collections calls for a hybrid query language: the text-rich nature of the documents suggest a content-oriented (IR) approach, while the mark-up allows users to add structural constraints to their IR queries. We will show how evidence for relevancy from different sources helps to answer such hybrid queries. We evaluate our methods using the INEX 2003 test set, and show that structural hints in hybrid queries help to improve retrieval effectiveness.
Börkur Sigurbjörnsson, Jaap Kamps, Maarten de Rijke
SIGIR1
2004 Best-Match Querying from Document-Centric XML
abstract
On the Web, there is a pervasive use of XML to give lightweight semantics to textual collections. Such document-centric XML collections require a query language that can gracefully handle structural constraints as well as constraints on the free text of the documents. Our main contributions are three-fold. First, we outline two fragments of XPath tailored to users that have varying degrees of understanding of the XML structure used, and give both syntactic and semantic characterizations of these fragments. Second, we extend XPath with an about function having a best-match semantics based on the relevance of the document component for the expressed information need. Third, we evaluate the resulting query language using the INEX 2003 test suite, and show that best-match approaches outperform exact-match approaches for evaluating content-and-structure queries.
Jaap Kamps, Maarten Marx, Maarten de Rijke, Börkur Sigurbjörnsson
WebDB4
2003 XML retrieval: what to retrieve?
abstract
The fundamental difference between standard information retrieval and XML retrieval is the unit of retrieval. In traditional IR, the unit of retrieval is fixed: it is the complete document. In XML retrieval, every XML element in a document is a retrievable unit. This makes XML retrieval more difficult: besides being relevant, a retrieved unit should be neither too large nor too small. The research presented here, a comparative analysis of two approaches to XML retrieval, aims to shed light on which XML elements should be retrieved. The experimental evaluation uses data from the Initiative for the Evaluation of XML retrieval (INEX 2002).
Jaap Kamps, Maarten Marx, Maarten de Rijke, Börkur Sigurbjörnsson
SIGIR4