Karl Gyllstrom

dblp:13/1086 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
1since 2021 · last 2023
0009-0003-6594-3552ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 6 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-authorArtificial intelligence and machine learning · 4 · 3 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
6 papers
Information retrieval · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
query suggestion
0.222010
Effects of popularity and quality on the usage of query suggestions during information search · CHI 2010
A comparison of query and term suggestion features for interactive searching · SIGIR 2009
Information retrieval
interactive information retrieval
0.222011
An examination of two delivery modes for interactive search system experiments: remote and laboratory · CHI 2011
A comparison of query and term suggestion features for interactive searching · SIGIR 2009
Information retrieval › evaluation
evaluation methodology
0.112011
An examination of two delivery modes for interactive search system experiments: remote and laboratory · CHI 2011
Information retrieval › search engines
search result augmentation
0.112010
A picture is worth a thousand search results: finding child-oriented multimedia results with collAge · SIGIR 2010
Computing education › computer science curriculum
information retrieval education
0.112009
Undergraduates' evaluations of assigned search topics · SIGIR 2009
Information retrieval
evaluation
0.112009
Undergraduates' evaluations of assigned search topics · SIGIR 2009
Information retrieval › query reformulation
query expansion
0.112009
A comparison of query and term suggestion features for interactive searching · SIGIR 2009
Information retrieval
query processing
0.112009
A comparison of query and term suggestion features for interactive searching · SIGIR 2009
Information retrieval › evaluation
test collection
0.112009
Undergraduates' evaluations of assigned search topics · SIGIR 2009
Information retrieval › evaluation › test collection
topic selection
0.112009
Undergraduates' evaluations of assigned search topics · SIGIR 2009
Information retrieval › search engines
desktop search
0.112007
Confluence: enhancing contextual desktop search · SIGIR 2007
Information retrieval › search engines
file search
0.112007
Confluence: enhancing contextual desktop search · SIGIR 2007
Usable security
user behavior
0.012010
Effects of popularity and quality on the usage of query suggestions during information search · CHI 2010
Information retrieval
query formulation
0.012009
A comparison of query and term suggestion features for interactive searching · SIGIR 2009

Methods — techniques the papers use, named apart from their topics

user study · 0.5comparative study · 0.2online user evaluation · 0.1user-generated suggestions · 0.1interactive information retrieval study · 0.1window focus event analysis · 0.1
YearPublicationVenuePosition
2023 Automatic and Precise Data Validation for Machine Learning
abstract
Machine learning (ML) models in production pipelines are frequently retrained on the latest partitions of large, continually- growing datasets. Due to engineering bugs, partitions in such datasets almost always have some corrupted features; thus, it's critical to find data issues and block retraining before downstream ML accuracy decreases. However, current ML data validation methods are difficult to operationalize: they yield too many false positive alerts, require manual tuning, or are infeasible at scale. In this pa- per, we present an automatic, precise, and scalable data validation system for ML pipelines, employing a simple idea that we call a Partition Summarization (PS) approach to data validation: each timestamp-based partition of data is summarized with data quality metrics, and summaries are compared to detect corrupted partitions. We demonstrate how to adapt PS for any data validation method in a robust manner and evaluate several adaptations-which by themselves provide limited precision. Finally, we present gate, our data validation method that leverages these adaptations, giving a 2.1× average improvement in precision over the baseline from prior work on a case study within our large tech company.
Shreya Shankar, Labib Fawaz, Karl Gyllstrom, Aditya G. Parameswaran
CIKM3
2012 The downside of markup: examining the harmful effects of CSS and javascript on indexing today's web
abstract
The continued development and maturation of advanced HTML features such as Cascading style sheets (CSS), Javascript, and AJAX, as well as their widespread adoption by browsers, has enabled web pages to flourish with sophistication and interactivity. Unfortunately, this presents challenges to the web search community, as a web page's representation in the browser (i.e., what users see) can diverge dramatically from its raw HTML content (i.e., what search engines index and retrieve). For example, interactive pages may contain content in regions that are not visible before a user action, such as focusing a tab, but which are nonetheless still contained within the raw HTML. We study this divergence by comparing raw HTML to its fully rendered form across a number of metrics spanning presentation, geometry, and content, using a large, representative sample of popular web pages. We find that a large divergence currently exists, and we show via a historical analysis that this divergence has grown more pronounced over the last decade. The general finding of our study is that continuing to index the web via simple HTML parsing will diminish the effectiveness of retrieval on the modern web, and that the IR community should work toward more sophisticated web page processing in indexing technology.
Karl Gyllstrom, Carsten Eickhoff, Arjen P. de Vries, Marie-Francine Moens
CIKM1
2012 EmSe: Supporting Children's Information Needs within a Hospital Environment
Leif Azzopardi, Douglas Dowie, Sergio Duarte Torres, Carsten Eickhoff, Richard Glassey, Karl Gyllstrom, Djoerd Hiemstra, Franciska de Jong, Frea Kruisinga, Kelly Ann Marshall, Marie-Francine Moens, Tamara Polajnar, Frans van der Sluis, Arjen P. de Vries
ECIR6
2011 An examination of two delivery modes for interactive search system experiments: remote and laboratory
abstract
We compare two delivery modes for interactive search system (ISS) experiments: remote and laboratory. Our study was completed by two groups of subjects from the same population. The first group completed the study remotely and the second group completed the study in the laboratory. We compare differences in participants, participation behaviors, search behaviors and evaluation behaviors. Overall, for most measures no significant differences were found, but there were some notable differences. Greater variance was observed in time taken and number of documents opened and saved by remote subjects. Lab subjects provided more favorable responses to exit questionnaire items and reported significantly higher satisfaction. Lab subjects also provided significantly longer responses to open questions, while remote subjects provided more null responses. These results suggest that many behaviors do not change significantly according to study mode and that results from remote ISS experiments are similar to those from laboratory experiments.
Diane Kelly 0001, Karl Gyllstrom
CHI2
2011 Examining the "leftness" property of Wikipedia categories
abstract
Wikipedia's rich category structure has helped make it one of the largest semantic taxonomies in existence, a property that has been central to much recent research. However, Wikipedia's category representation is simplistic: an article contains a single list of categories, with no data about their relative importance. We investigate the ordering of category lists to determine how a category's position in the list correlates with its relevance to the article and overall significance. We identify a number of interesting connections between a category's position and its persistence within the article, age, popularity, size, and descriptiveness.
Karl Gyllstrom, Marie-Francine Moens
CIKM1
2011 Web Search Query Assistance Functionality for Young Audiences
Carsten Eickhoff, Tamara Polajnar, Karl Gyllstrom, Sergio Duarte Torres, Richard Glassey
ECIR3
2011 Clash of the Typings - Finding Controversies and Children's Topics Within Queries
Karl Gyllstrom, Marie-Francine Moens
ECIR1
2010 Effects of popularity and quality on the usage of query suggestions during information search
abstract
Many search systems provide users with recommended queries during online information seeking. Although usage statistics are often used to recommend queries, this information is usually not displayed to the user. In this study, we investigate how the presentation of this information impacts use of query suggestions. Twenty-three subjects used an experimental search system to find documents about four topics. Eight query suggestions were provided for each topic: four were high quality queries and four were low quality queries. Fake usage information indicating how many other people used the queries was also provided. For half the queries this information was high and for the other half this information was low. Results showed that subjects could distinguish between high and low quality queries and were not influenced by the usage information. Qualitative data revealed that subjects felt favorable about the suggestions, but the usage information was less important for the search task used in this study.
Diane Kelly 0001, Amber L. Cushing, Maureen Dostert, Xi Niu, Karl Gyllstrom
CHI5
2010 Wisdom of the ages: toward delivering the children's web with the link-based agerank algorithm
abstract
Though children frequently use web search engines to learn, interact, and be entertained, modern web search engines are poorly suited to children's needs, requiring relatively complex querying and filtering of results in order to find pages oriented to young audiences. To address this limitation, we designed AgeRank, a link-based algorithm that ranks web pages according their appropriateness for young audiences. We show its effectiveness through a multipart evaluation that demonstrates AgeRank to be accurate in page-labeling, widely-spanning in page coverage, and with high potential to improve children's search. As a fast, scalable, and effective algorithm, AgeRank can be adopted by search engines seeking to more effectively address the needs of young users, or easily fitted to complementary machine-learning based classification approaches.
Karl Gyllstrom, Marie-Francine Moens
CIKM1
2010 Automatic generation of research trails in web history
abstract
We propose the concept of research trails to help web users create and reestablish context across fragmented research processes without requiring them to explicitly structure and organize the material. A research trail is an ordered sequence of web pages that were accessed as part of a larger investigation; they are automatically constructed by filtering and organizing users' activity history, using a combination of semantic and activity based criteria for grouping similar visited web pages. The design was informed by an ethnographic study of ordinary people doing research on the web, emphasizing a need to support research processes that are fragmented and where the research question is still in formation. This paper motivates and describes our algorithms for generating research trails.
Elin Rønby Pedersen, Karl Gyllstrom, Shengyin Gu, Peter Jin Hong
IUI2
2010 A picture is worth a thousand search results: finding child-oriented multimedia results with collAge
abstract
We present a simple and effective approach to complement search results for children’s web queries with child-oriented multimedia results, such as coloring pages and music sheets. Our approach determines appropriate media types for a query by searching Google’s database of frequent queries for co-occurrences of a query’s terms (e.g., “dinosaurs”) with preselected multimedia terms (e.g., “coloring pages”). We show the effectiveness of this approach through an online user evaluation.
Karl Gyllstrom, Marie-Francine Moens
SIGIR1
2009 Passages through time: chronicling users' information interaction history by recording when and what they read
abstract
The Passages system enhances information management by maintaining a detailed chronicle of all the text the user ever reads or edits, and making this chronicle available for rich temporal queries about the user's information workspace. Passages enables queries like, "which papers and web pages did I read when writing the 'related work' section of this paper?", and, "which of the emails in this folder have I skimmed, but not yet read in detail?" As time and interaction history are important attributes in users' recall of their personal information, effectively supporting them creates useful possibilities for information retrieval. We present methods to collect and make sense of the large volume of text with which the user interacts. We show through user evaluation the accuracy of Passages in building interaction history, and illustrate its capacity to both improve existing retrieval systems and enable novel ways to characterize document activity across time.
Karl Gyllstrom
IUI1
2009 Undergraduates' evaluations of assigned search topics
abstract
This paper evaluates undergraduate students' knowledge, interests and experiences with 20 topics from the TREC Robust Track collection. The goal is to characterize these topics along several dimensions to help researchers make more informed decisions about which topics are most appropriate to use in experimental IIR evaluations with undergraduate student subjects.
Earl W. Bailey, Diane Kelly 0001, Karl Gyllstrom
SIGIR3
2009 A comparison of query and term suggestion features for interactive searching
abstract
Query formulation is one of the most difficult and important aspects of information seeking and retrieval. Two techniques, term relevance feedback and query suggestion, provide methods to help users formulate queries, but each is limited in different ways. In this research we combine these two techniques by automatically creating query suggestions using term relevance feedback techniques. To evaluate our approach, we conducted an interactive information retrieval study with 55 subjects and 20 topics. Each subject completed four topics, half with a term suggestion system and half with a query suggestion system. We also investigated the source of the suggestions: approximately half of all subjects were provided with system-generated suggestions, while half were provided with user-generated suggestions. Results show that subjects used more query suggestions than term suggestions and saved more documents with these suggestions, even though there were no significant differences in performance. Subjects preferred the query suggestion system and rated it higher along a number of dimensions including its ability to help them think of new approaches to searching. Qualitative data provided insight into subjects' usage and ratings, and indicated that subjects often used the suggestions even when they did not click on them.
Diane Kelly 0001, Karl Gyllstrom, Earl W. Bailey
SIGIR2
2008 Seeing is retrieving: building information context from what the user sees
abstract
As the user's document and application workspace grows more diverse, supporting personal information management becomes increasingly important. This trend toward diversity renders it difficult to implement systems which are tailored to specific applications, file types, or other information sources.
Karl Gyllstrom, Craig A. N. Soules
IUI1
2007 Confluence: enhancing contextual desktop search
abstract
We present Confluence, an enhancement to a desktop file search tool called Confluence which extracts conceptual relationships between files by their temporal access patterns in the file system. A limitation of a purely file-based approach is that as file operations are increasingly abstracted by applications, their correlation to a user's activity weakens and thereby reduces the applicability of their temporal patterns. To deal with this problem, we augment the file event stream with a stream of window focus events from the UI layer. We present 3 algorithms that analyze this new stream, extracting the user's task information which informs the existing Confluence algorithms. We present results and conclusions from a preliminary user study on Confluence.
Karl Gyllstrom, Craig A. N. Soules, Alistair C. Veitch
SIGIR1