Simon Jonassen

dblp:75/8908 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 67% Information retrieval · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
query result caching
0.112012
Prefetching query results and its impact on search engines · SIGIR 2012
Query processing and optimization › runtime optimization › prefetching
query result prefetching
0.112012
Prefetching query results and its impact on search engines · SIGIR 2012
Information retrieval
web search
0.112012
Prefetching query results and its impact on search engines · SIGIR 2012

Methods — techniques the papers use, named apart from their topics

query log analysis · 0.1machine learning · 0.1
YearPublicationVenuePosition
2014 Improving the performance of pipelined query processing with skipping - and its comparison to document-wise partitioning
Simon Jonassen, Svein Erik Bratsberg
World Wide Web1
2013 A term-based inverted index partitioning model for efficient distributed query processing
abstract
In a shared-nothing, distributed text retrieval system, queries are processed over an inverted index that is partitioned among a number of index servers. In practice, the index is either document-based or term-based partitioned. This choice is made depending on the properties of the underlying hardware infrastructure, query traffic distribution, and some performance and availability constraints. In query processing on retrieval systems that adopt a term-based index partitioning strategy, the high communication overhead due to the transfer of large amounts of data from the index servers forms a major performance bottleneck, deteriorating the scalability of the entire distributed retrieval system. In this work, to alleviate this problem, we propose a novel inverted index partitioning model that relies on hypergraph partitioning. In the proposed model, concurrently accessed index entries are assigned to the same index servers, based on the inverted index access patterns extracted from the past query logs. The model aims to minimize the communication overhead that will be incurred by future queries while maintaining the computational load balance among the index servers. We evaluate the performance of the proposed model through extensive experiments using a real-life text collection and a search query sample. Our results show that considerable performance gains can be achieved relative to the term-based index partitioning strategies previously proposed in literature. In most cases, however, the performance remains inferior to that attained by document-based partitioning.
Berkant Barla Cambazoglu, Enver Kayaaslan, Simon Jonassen, Cevdet Aykanat
ACM Trans. Web3
2012 Modeling Static Caching in Web Search Engines
Ricardo Baeza-Yates, Simon Jonassen
ECIR2
2012 Intra-query Concurrent Pipelined Processing for Distributed Full-Text Retrieval
Simon Jonassen, Svein Erik Bratsberg
ECIR1
2012 Prefetching query results and its impact on search engines
abstract
We investigate the impact of query result prefetching on the efficiency and effectiveness of web search engines. We propose offline and online strategies for selecting and ordering queries whose results are to be prefetched. The offline strategies rely on query log analysis and the queries are selected from the queries issued on the previous day. The online strategies select the queries from the result cache, relying on a machine learning model that estimates the arrival times of queries. We carefully evaluate the proposed prefetching techniques via simulation on a query log obtained from Yahoo! web search. We demonstrate that our strategies are able to improve various performance metrics, including the hit rate, query response time, result freshness, and query degradation rate, relative to a state-of-the-art baseline.
Simon Jonassen, Berkant Barla Cambazoglu, Fabrizio Silvestri
SIGIR1
2012 Improving the Performance of Pipelined Query Processing with Skipping
Simon Jonassen, Svein Erik Bratsberg
WISE1
2011 Efficient Compressed Inverted Index Skipping for Disjunctive Text-Queries
Simon Jonassen, Svein Erik Bratsberg
ECIR1
2011 Efficient Processing of Top-k Spatial Keyword Queries
João B. Rocha-Junior, Orestis Gkorgkas, Simon Jonassen, Kjetil Nørvåg
SSTD3
2010 A Combined Semi-pipelined Query Processing Architecture for Distributed Full-Text Retrieval
Simon Jonassen, Svein Erik Bratsberg
WISE1