Craig Willis

dblp:58/10873 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0002-6148-7196ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-authorSoftware engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
evaluation
0.212016
What Makes a Query Temporally Sensitive? · SIGIR 2016
Information retrieval › document retrieval
temporal information retrieval
0.212016
What Makes a Query Temporally Sensitive? · SIGIR 2016
Information retrieval › evaluation
test collection
0.212016
What Makes a Query Temporally Sensitive? · SIGIR 2016
Information retrieval › retrieval models
boolean retrieval
0.212014
Learning sufficient queries for entity filtering · SIGIR 2014
Information retrieval › information filtering
document filtering
0.212014
Learning sufficient queries for entity filtering · SIGIR 2014
Information retrieval › information filtering
entity filtering
0.212014
Learning sufficient queries for entity filtering · SIGIR 2014
Information retrieval
retrieval models
0.112016
What Makes a Query Temporally Sensitive? · SIGIR 2016
Information retrieval › retrieval models › ad-hoc retrieval
time-aware retrieval
0.112016
What Makes a Query Temporally Sensitive? · SIGIR 2016
Information retrieval › machine learning for information retrieval
query learning
0.112014
Learning sufficient queries for entity filtering · SIGIR 2014

Methods — techniques the papers use, named apart from their topics

regression analysis · 0.2quantitative analysis · 0.2qualitative analysis · 0.2deterministic filtering · 0.2boolean query learning · 0.2
YearPublicationVenuePosition
2020 CHEESE: Cyber Human Ecosystem of Engaged Security Education
abstract
This Innovative Practice Full Paper presents CHEESE, a platform for cybersecurity education that complements formal classroom instruction with hands-on experience. With the ubiquitous use of computing devices and applications today, the protection of personal and privileged information is a persistent challenge. Modern software applications are typically complex pieces of code that borrow from various preexisting software libraries. Consequently, a flaw in one piece of software can have far-reaching and often unintended security implications that malicious actors can exploit. Thus, cybersecurity education needs to be transformed from a purely academic enterprise for cybersecurity researchers into a necessary skill that is imparted to the current and future IT workforce at large. CHEESE aims to impart such skills. CHEESE is composed of CHEESEHub, a public web-platform hosting demonstrations of cybersecurity concepts, a set of lessons complementing the demonstrations, and a community-driven approach to the contribution of new demonstrations and lessons. CHEESE is intended to supplement and enhance traditional cybersecurity education with hands-on training that has been shown to improve concept retention and understanding. Instructors can incorporate CHEESE into their teaching in several ways: by utilizing one or more of the demonstrations hosted on the publicly-accessible CHEESEHub in conjunction with the web-accessible lessons; by deploying their own version of CHEESEHub with a custom set of demonstrations and lessons; or by developing their own lesson plan which borrows from and combines one or more demonstrations on CHEESEHub. The use of CHEESEHub only requires a web-browser and can hence be employed in a wide variety of educational and training settings from K-12 schools through university.
Rajesh Kalyanam, Baijian Yang 0001, Craig Willis, Mike Lambert, Christine R. Kirkpatrick
FIE3
2019 Application of BagIt-Serialized Research Object Bundles for Packaging and Re-Execution of Computational Analyses
abstract
In this paper we describe our experience adopting the Research Object Bundle (RO-Bundle) format with BagIt serialization (BagIt-RO) for the design and implementation of "tales" in the Whole Tale platform. A tale is an executable research object intended for the dissemination of computational scientific findings that captures information needed to facilitate understanding, transparency, and re-execution for review and computational reproducibility at the time of publication. We describe the Whole Tale platform and requirements that led to our adoption of BagIt-RO, specifics of our implementation, and discuss migrating to the emerging Research Object Crate (RO-Crate) standard.
Kyle Chard, Thomas Thelen, Matthew J. Turk, Craig Willis, Niall Gaffney, Matthew B. Jones, Kacper Kowalik, Bertram Ludäscher, Timothy M. McPhillips, Jarek Nabrzyski, Victoria Stodden, Ian J. Taylor
eScience4
2019 Reproducibility by Other Means: Transparent Research Objects
abstract
Research Objects have the potential to significantly enhance the reproducibility of scientific research. One important way Research Objects can do this is by encapsulating the means for re-executing the computational components of studies, thus supporting the new form of reproducibility enabled by digital computing-exact repeatability. However, Research Objects also can make scientific research more reproducible by supporting transparency, a component of reproducibility orthogonal to re-executability. We describe here our vision for making Research Objects more transparent by providing means for disambiguating claims about reproducibility generally, and computational repeatability specifically. We show how support for science-oriented queries can enable researchers to assess the reproducibility of Research Objects and the individual methods and results they encapsulate.
Timothy M. McPhillips, Craig Willis, Michael R. Gryk, Santiago Núñez-Corrales, Bertram Ludäscher
eScience2
2018 Preserving Reproducibility: Provenance and Executable Containers in DataONE Data Packages
abstract
Many data packaging standards are available to researchers and data repository operators and the choice to use an existing standard or create a new one is challenging. We introduce the DataONE Data Package standard which is based on the existing OAI-ORE Resource Map standard. We describe the functionality Data Package provides, implementation considerations, compare it to existing standards, and discuss future extensions to the standard including the ability to describe execution environments via WholeTale "Tales"" and alternate serialization formats.
Bryce D. Mecum, Matthew B. Jones, David Vieglais, Craig Willis
eScience4
2016 What Makes a Query Temporally Sensitive?
abstract
This work takes an in-depth look at the factors that affect manual classifications of 'temporally sensitive' information needs. We use qualitative and quantitative techniques to analyze 660 topics from the Text Retrieval Conference (TREC) previously used in the experimental evaluation of temporal retrieval models. Regression analysis is used to identify factors in previous manual classifications. We explore potential problems with the previous classifications, considering principles and guidelines for future work on temporal retrieval models.
Craig Willis, Garrick Sherman, Miles Efron
SIGIR1
2014 Learning sufficient queries for entity filtering
abstract
Entity-centric document filtering is the task of analyzing a time-ordered stream of documents and emitting those that are relevant to a specified set of entities (e.g., people, places, organizations). This task is exemplified by the TREC Knowledge Base Acceleration (KBA) track and has broad applicability in other modern IR settings. In this paper, we present a simple yet effective approach based on learning high-quality Boolean queries that can be applied deterministically during filtering. We call these Boolean statements sufficient queries. We argue that using deterministic queries for entity-centric filtering can reduce confounding factors seen in more familiar "score-then-threshold" filtering methods. Experiments on two standard datasets show significant improvements over state-of-the-art baseline models.
Miles Efron, Craig Willis, Garrick Sherman
SIGIR2
2013 A random walk on an ontology: Using thesaurus structure for automatic subject indexing
abstract
Relationships between terms and features are an essential component of thesauri, ontologies, and a range of controlled vocabularies. In this article, we describe ways to identify important concepts in documents using the relationships in a thesaurus or other vocabulary structures. We introduce a methodology for the analysis and modeling of the indexing process based on a weighted random walk algorithm. The primary goal of this research is the analysis of the contribution of thesaurus structure to the indexing process. The resulting models are evaluated in the context of automatic subject indexing using four collections of documents pre‐indexed with 4 different thesauri (AGROVOC [UN Food and Agriculture Organization], high‐energy physics taxonomy [HEP], National Agricultural Library Thesaurus [NALT], and medical subject headings [MeSH]). We also introduce a thesaurus‐centric matching algorithm intended to improve the quality of candidate concepts. In all cases, the weighted random walk improves automatic indexing performance over matching alone with an increase in average precision (AP) of 9% for HEP, 11% for MeSH, 35% for NALT, and 37% for AGROVOC. The results of the analysis support our hypothesis that subject indexing is in part a browsing process, and that using the vocabulary and its structure in a thesaurus contributes to the indexing process. The amount that the vocabulary structure contributes was found to differ among the 4 thesauri, possibly due to the vocabulary used in the corresponding thesauri and the structural relationships between the terms. Each of the thesauri and the manual indexing associated with it is characterized using the methods developed here.
Craig Willis, Robert M. Losee
J. Assoc. Inf. Sci. Technol.1
2012 Analysis and synthesis of metadata goals for scientific data
abstract
The proliferation of discipline‐specific metadata schemes contributes to artificial barriers that can impede interdisciplinary and transdisciplinary research. The authors considered this problem by examining thedomains,objectives, andarchitecturesof nine metadata schemes used to document scientific data in the physical, life, and social sciences. They used a mixed‐methods content analysis andGreenberg's ( ) metadata objectives, principles, domains, and architectural layout (MODAL) framework, and derived 22 metadata‐related goals from textual content describing each metadata scheme. Relationships are identified between the domains (e.g., scientific discipline and type of data) and the categories of scheme objectives. For each strong correlation (>0.6), a Fisher's exact test for nonparametric data was used to determine significance (p < .05). Significant relationships were found between the domains and objectives of the schemes. Schemes describing observational data are more likely to have “scheme harmonization” (compatibility and interoperability with related schemes) as an objective; schemes with the objective “abstraction” (a conceptual model exists separate from the technical implementation) also have the objective “sufficiency” (the scheme defines a minimal amount of information to meet the needs of the community); and schemes with the objective “data publication” do not have the objective “element refinement.” The analysis indicates that many metadata‐driven goals expressed by communities are independent of scientific discipline or the type of data, although they are constrained by historical community practices and workflows as well as the technological environment at the time of scheme creation. The analysis reveals 11 fundamental metadata goals for metadata documenting scientific data in support of sharing research data across disciplines and domains. The authors report these results and highlight the need for more metadata‐related research, particularly in the context of recent funding agency policy changes.
Craig Willis, Jane Greenberg, Hollie White
J. Assoc. Inf. Sci. Technol.1