VLDB 2026 Research / reviewers in the wild / expert
Craig Willis
dblp:58/10873
· DBLP profile ↗
8ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0002-6148-7196ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 3 first-authorSoftware engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
evaluation |
0.2 | 1 | 2016 | What Makes a Query Temporally Sensitive? · SIGIR 2016 |
Information retrieval › document retrieval
temporal information retrieval |
0.2 | 1 | 2016 | What Makes a Query Temporally Sensitive? · SIGIR 2016 |
Information retrieval › evaluation
test collection |
0.2 | 1 | 2016 | What Makes a Query Temporally Sensitive? · SIGIR 2016 |
Information retrieval › retrieval models
boolean retrieval |
0.2 | 1 | 2014 | Learning sufficient queries for entity filtering · SIGIR 2014 |
Information retrieval › information filtering
document filtering |
0.2 | 1 | 2014 | Learning sufficient queries for entity filtering · SIGIR 2014 |
Information retrieval › information filtering
entity filtering |
0.2 | 1 | 2014 | Learning sufficient queries for entity filtering · SIGIR 2014 |
Information retrieval
retrieval models |
0.1 | 1 | 2016 | What Makes a Query Temporally Sensitive? · SIGIR 2016 |
Information retrieval › retrieval models › ad-hoc retrieval
time-aware retrieval |
0.1 | 1 | 2016 | What Makes a Query Temporally Sensitive? · SIGIR 2016 |
Information retrieval › machine learning for information retrieval
query learning |
0.1 | 1 | 2014 | Learning sufficient queries for entity filtering · SIGIR 2014 |
Methods — techniques the papers use, named apart from their topics
regression analysis · 0.2quantitative analysis · 0.2qualitative analysis · 0.2deterministic filtering · 0.2boolean query learning · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | CHEESE: Cyber Human Ecosystem of Engaged Security EducationabstractThis Innovative Practice Full Paper presents CHEESE, a platform for cybersecurity education that complements formal classroom instruction with hands-on experience. With the ubiquitous use of computing devices and applications today, the protection of personal and privileged information is a persistent challenge. Modern software applications are typically complex pieces of code that borrow from various preexisting software libraries. Consequently, a flaw in one piece of software can have far-reaching and often unintended security implications that malicious actors can exploit. Thus, cybersecurity education needs to be transformed from a purely academic enterprise for cybersecurity researchers into a necessary skill that is imparted to the current and future IT workforce at large. CHEESE aims to impart such skills. CHEESE is composed of CHEESEHub, a public web-platform hosting demonstrations of cybersecurity concepts, a set of lessons complementing the demonstrations, and a community-driven approach to the contribution of new demonstrations and lessons. CHEESE is intended to supplement and enhance traditional cybersecurity education with hands-on training that has been shown to improve concept retention and understanding. Instructors can incorporate CHEESE into their teaching in several ways: by utilizing one or more of the demonstrations hosted on the publicly-accessible CHEESEHub in conjunction with the web-accessible lessons; by deploying their own version of CHEESEHub with a custom set of demonstrations and lessons; or by developing their own lesson plan which borrows from and combines one or more demonstrations on CHEESEHub. The use of CHEESEHub only requires a web-browser and can hence be employed in a wide variety of educational and training settings from K-12 schools through university. Rajesh Kalyanam, Baijian Yang 0001, Craig Willis, Mike Lambert, Christine R. Kirkpatrick |
FIE | 3 |
| 2019 | Application of BagIt-Serialized Research Object Bundles for Packaging and Re-Execution of Computational AnalysesabstractIn this paper we describe our experience adopting the Research Object Bundle (RO-Bundle) format with BagIt serialization (BagIt-RO) for the design and implementation of "tales" in the Whole Tale platform. A tale is an executable research object intended for the dissemination of computational scientific findings that captures information needed to facilitate understanding, transparency, and re-execution for review and computational reproducibility at the time of publication. We describe the Whole Tale platform and requirements that led to our adoption of BagIt-RO, specifics of our implementation, and discuss migrating to the emerging Research Object Crate (RO-Crate) standard. Kyle Chard, Thomas Thelen, Matthew J. Turk, Craig Willis, Niall Gaffney, Matthew B. Jones, Kacper Kowalik, Bertram Ludäscher, Timothy M. McPhillips, Jarek Nabrzyski, Victoria Stodden, Ian J. Taylor |
eScience | 4 |
| 2019 | Reproducibility by Other Means: Transparent Research ObjectsabstractResearch Objects have the potential to significantly enhance the reproducibility of scientific research. One important way Research Objects can do this is by encapsulating the means for re-executing the computational components of studies, thus supporting the new form of reproducibility enabled by digital computing-exact repeatability. However, Research Objects also can make scientific research more reproducible by supporting transparency, a component of reproducibility orthogonal to re-executability. We describe here our vision for making Research Objects more transparent by providing means for disambiguating claims about reproducibility generally, and computational repeatability specifically. We show how support for science-oriented queries can enable researchers to assess the reproducibility of Research Objects and the individual methods and results they encapsulate. Timothy M. McPhillips, Craig Willis, Michael R. Gryk, Santiago Núñez-Corrales, Bertram Ludäscher |
eScience | 2 |
| 2018 | Preserving Reproducibility: Provenance and Executable Containers in DataONE Data PackagesabstractMany data packaging standards are available to researchers and data repository operators and the choice to use an existing standard or create a new one is challenging. We introduce the DataONE Data Package standard which is based on the existing OAI-ORE Resource Map standard. We describe the functionality Data Package provides, implementation considerations, compare it to existing standards, and discuss future extensions to the standard including the ability to describe execution environments via WholeTale "Tales"" and alternate serialization formats. Bryce D. Mecum, Matthew B. Jones, David Vieglais, Craig Willis |
eScience | 4 |
| 2016 | What Makes a Query Temporally Sensitive?abstractThis work takes an in-depth look at the factors that affect manual classifications of 'temporally sensitive' information needs. We use qualitative and quantitative techniques to analyze 660 topics from the Text Retrieval Conference (TREC) previously used in the experimental evaluation of temporal retrieval models. Regression analysis is used to identify factors in previous manual classifications. We explore potential problems with the previous classifications, considering principles and guidelines for future work on temporal retrieval models. Craig Willis, Garrick Sherman, Miles Efron |
SIGIR | 1 |
| 2014 | Learning sufficient queries for entity filteringabstractEntity-centric document filtering is the task of analyzing a time-ordered stream of documents and emitting those that are relevant to a specified set of entities (e.g., people, places, organizations). This task is exemplified by the TREC Knowledge Base Acceleration (KBA) track and has broad applicability in other modern IR settings. In this paper, we present a simple yet effective approach based on learning high-quality Boolean queries that can be applied deterministically during filtering. We call these Boolean statements sufficient queries. We argue that using deterministic queries for entity-centric filtering can reduce confounding factors seen in more familiar "score-then-threshold" filtering methods. Experiments on two standard datasets show significant improvements over state-of-the-art baseline models. Miles Efron, Craig Willis, Garrick Sherman |
SIGIR | 2 |
| 2013 | A random walk on an ontology: Using thesaurus structure for automatic subject indexingabstractRelationships between terms and features are an essential component of thesauri, ontologies, and a range of controlled vocabularies. In this article, we describe ways to identify important concepts in documents using the relationships in a thesaurus or other vocabulary structures. We introduce a methodology for the analysis and modeling of the indexing process based on a weighted random walk algorithm. The primary goal of this research is the analysis of the contribution of thesaurus structure to the indexing process. The resulting models are evaluated in the context of automatic subject indexing using four collections of documents pre‐indexed with 4 different thesauri (AGROVOC [UN Food and Agriculture Organization], high‐energy physics taxonomy [HEP], National Agricultural Library Thesaurus [NALT], and medical subject headings [MeSH]). We also introduce a thesaurus‐centric matching algorithm intended to improve the quality of candidate concepts. In all cases, the weighted random walk improves automatic indexing performance over matching alone with an increase in average precision (AP) of 9% for HEP, 11% for MeSH, 35% for NALT, and 37% for AGROVOC. The results of the analysis support our hypothesis that subject indexing is in part a browsing process, and that using the vocabulary and its structure in a thesaurus contributes to the indexing process. The amount that the vocabulary structure contributes was found to differ among the 4 thesauri, possibly due to the vocabulary used in the corresponding thesauri and the structural relationships between the terms. Each of the thesauri and the manual indexing associated with it is characterized using the methods developed here. Craig Willis, Robert M. Losee |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2012 | Analysis and synthesis of metadata goals for scientific dataabstractThe proliferation of discipline‐specific metadata schemes contributes to artificial barriers that can impede interdisciplinary and transdisciplinary research. The authors considered this problem by examining thedomains,objectives, andarchitecturesof nine metadata schemes used to document scientific data in the physical, life, and social sciences. They used a mixed‐methods content analysis andGreenberg's ( ) metadata objectives, principles, domains, and architectural layout (MODAL) framework, and derived 22 metadata‐related goals from textual content describing each metadata scheme. Relationships are identified between the domains (e.g., scientific discipline and type of data) and the categories of scheme objectives. For each strong correlation (>0.6), a Fisher's exact test for nonparametric data was used to determine significance (p < .05). Significant relationships were found between the domains and objectives of the schemes. Schemes describing observational data are more likely to have “scheme harmonization” (compatibility and interoperability with related schemes) as an objective; schemes with the objective “abstraction” (a conceptual model exists separate from the technical implementation) also have the objective “sufficiency” (the scheme defines a minimal amount of information to meet the needs of the community); and schemes with the objective “data publication” do not have the objective “element refinement.” The analysis indicates that many metadata‐driven goals expressed by communities are independent of scientific discipline or the type of data, although they are constrained by historical community practices and workflows as well as the technological environment at the time of scheme creation. The analysis reveals 11 fundamental metadata goals for metadata documenting scientific data in support of sharing research data across disciplines and domains. The authors report these results and highlight the need for more metadata‐related research, particularly in the context of recent funding agency policy changes. Craig Willis, Jane Greenberg, Hollie White |
J. Assoc. Inf. Sci. Technol. | 1 |