Andrew Kane

dblp:73/10407 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
1since 2021 · last 2025
0009-0003-3102-4267ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 3 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Information retrieval · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
query processing
0.522018
Split-Lists and Initial Thresholds for WAND-based Search · SIGIR 2018
Skewed partial bitvectors for list intersection · SIGIR 2014
Information retrieval › document retrieval › domain-specific retrieval › mathematical information retrieval › math-aware search
mathematical formula retrieval
0.212016
Multi-Stage Math Formula Search: Using Appearance-Based Similarity Metrics at Scale · SIGIR 2016
Information retrieval
multi-stage retrieval
0.212016
Multi-Stage Math Formula Search: Using Appearance-Based Similarity Metrics at Scale · SIGIR 2016
Information retrieval
reranking
0.212016
Multi-Stage Math Formula Search: Using Appearance-Based Similarity Metrics at Scale · SIGIR 2016
Information retrieval › query processing
list intersection
0.212014
Skewed partial bitvectors for list intersection · SIGIR 2014
Distributed systems › peer-to-peer systems
distributed search
0.112018
Split-Lists and Initial Thresholds for WAND-based Search · SIGIR 2018
Information retrieval
ranking
0.112014
Skewed partial bitvectors for list intersection · SIGIR 2014
Information retrieval › ranking
ranking-based retrieval
0.112014
Skewed partial bitvectors for list intersection · SIGIR 2014

Methods — techniques the papers use, named apart from their topics

split-list WAND · 0.7initial thresholding · 0.7inverted index · 0.2dice coefficient · 0.2delta compression · 0.2URL ordering · 0.2
YearPublicationVenuePosition
2025 Exploiting Query Reformulation and Reciprocal Rank Fusion in Math-Aware Search Engines
abstract
Mathematical formulas introduce complications to the standard approaches used in information retrieval. By studying how traditional (sparse) search systems perform in matching queries to documents, we hope to gain insights into which features in the formulas and in the accompanying natural language text signal likely relevance.
Besat Kassaie, Andrew Kane, Frank Wm. Tompa
DocEng2
2018 Choosing Math Features for BM25 Ranking with Tangent-L
abstract
Combining text and mathematics when searching in a corpus with extensive mathematical notation remains an open problem. Recent results for Tangent-3 on the math and text retrieval task at NTCIR-12, for example, have room for improvement, even though formula retrieval appeared to be fairly successful.
Dallas J. Fraser, Andrew Kane, Frank Wm. Tompa
DocEng2
2018 Split-Lists and Initial Thresholds for WAND-based Search
abstract
We examine search engine performance for rank-safe query execution using the WAND and state-of-the-art BMW algorithms. Supported by extensive experiments, we suggest two approaches to improve query performance: initial list thresholds should be used when k values are large, and our split-list WAND approach should be used instead of the normal WAND or BMW approaches. We also recommend that reranking-based distributed systems use smaller k values when selecting the results to return from each partition.
Andrew Kane, Frank Wm. Tompa
SIGIR1
2017 Small-Term Distribution for Disk-Based Search
Andrew Kane, Frank Wm. Tompa
DocEng1
2016 Multi-Stage Math Formula Search: Using Appearance-Based Similarity Metrics at Scale
abstract
When using a mathematical formula for search (query-by-expression), the suitability of retrieved formulae often depends more upon symbol identities and layout than deep mathematical semantics. Using a Symbol Layout Tree representation for formula appearance, we propose the Maximum Subtree Similarity (MSS) for ranking formulae based upon the subexpression whose symbols and layout best match a query formula. Because MSS is too expensive to apply against a complete collection, the Tangent-3 system first retrieves expressions using an inverted index over symbol pair relationships, ranking hits using the Dice coefficient; the top-k formulae are then re-ranked by MSS. Tangent-3 obtains state-of-the-art performance on the NTCIR-11 Wikipedia formula retrieval benchmark, and is efficient in terms of both space and time. Retrieval systems for other graphical forms, including chemical diagrams, flowcharts, figures, and tables, may benefit from adopting this approach.
Richard Zanibbi, Kenny Davila, Andrew Kane, Frank Wm. Tompa
SIGIR3
2014 Skewed partial bitvectors for list intersection
abstract
This paper examines the space-time performance of in-memory conjunctive list intersection algorithms, as used in search engines, where integers represent document identifiers. We demonstrate that the combination of bitvectors, large skips, delta compressed lists and URL ordering produces superior results to using skips or bitvectors alone. We define semi-bitvectors, a new partial bitvector data structure that stores the front of the list using a bitvector and the remainder using skips and delta compression. To make it particularly effective, we propose that documents be ordered so as to skew the postings lists to have dense regions at the front. This can be accomplished by grouping documents by their size in a descending manner and then reordering within each group using URL ordering. In each list, the division point between bitvector and delta compression can occur at any group boundary. We explore the performance of semi-bitvectors using the GOV2 dataset for various numbers of groups, resulting in significant space-time improvements over existing approaches. Semi-bitvectors do not directly support ranking. Indeed, bitvectors are not believed to be useful for ranking based search systems, because frequencies and offsets cannot be included in their structure. To refute this belief, we propose several approaches to improve the performance of ranking-based search systems using bitvectors, and leave their verification for future work. These proposals suggest that bitvectors, and more particularly semi-bitvectors, warrant closer examination by the research community.
Andrew Kane, Frank Wm. Tompa
SIGIR1