EDBT 2026 Demo / reviewers in the wild / expert
Gordon V. Cormack
dblp:c/GVCormack
· DBLP profile ↗
69ranked-venue papers
32as first author
3since 2021 · last 2026
0000-0002-5890-0293ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 50 · 26 first-author · 3 since 2021Software engineering, systems software and programming languages · 10 · 3 first-authorArtificial intelligence and machine learning · 6 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 2 first-authorTheory of computation · 4 · 1 first-authorSystems, architecture and hardware · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AI ethics education: A scoping review of pedagogy, curriculum, and assessmentabstractBackground Artificial intelligence (AI) is increasingly embedded in social and institutional decision-making, creating demand for ethically literate practitioners. Universities have responded by introducing AI ethics instruction, but the structure, content, pedagogy, and evaluation of these efforts remain unevenly documented. Objective To map and synthesize research on university level AI ethics education by characterizing course design, pedagogy, ethical themes, and assessment methods, and identifying evidence gaps that limit knowledge consolidation and instructional refinement. Methods We conducted a scoping review using Continuous Active Learning to screen 50,766 records up to 2024. 43 studies met inclusion after title, abstract, and full text review. We coded instructional design, curricular themes, pedagogical methods, and evaluation approaches using descriptive frequency counts and qualitative synthesis. Results Most included studies were conceptual or descriptive, with relatively few empirical evaluations. Instruction was concentrated in computing and engineering and primarily targeted undergraduate learners. Ethics content was more often embedded within technical courses than delivered as standalone offerings. Reported pedagogy relied heavily on lecture and case-based discussion, with fewer studies describing participatory formats such as simulations or role-play. Curricular emphasis clustered around bias/fairness and privacy, with comparatively less attention to governance, explainability, and trust. Evaluation most often relied on self-report and reflective methods, while validated instruments and performance-based assessments were less common, and behavioral or applied outcomes were rarely assessed. Conclusions The literature suggests a field oriented toward awareness-building more than measurable ethical competence. Clearer competency claims, stronger assessment transparency, and greater alignment between instructional design and evaluation would improve comparability across studies and support evidence-informed course development. Calvin Hillis, Maushumi Bhattacharjee, Batool AlMousawi, Riley Martens, Tarik Eltanahy, Sara Ono, Marcus Hui, Ba' Pham, Michelle Swab, Gordon V. Cormack, Maura R. Grossman, Ebrahim Bagheri, Zack Marshall |
Inf. Process. Manag. | 10 |
| 2024 | Unbiased Validation of Technology-Assisted Review for eDiscoveryabstractAlthough it is well established that recall estimates are valid only when based on independent relevance assessments, and useful only to compare the relative effectiveness of competing methods, these conditions are seldom met when validating eDiscovery efforts in litigation. We present two unbiased validation strategies that embed blind relevance assessments into a technology-assisted review (TAR) process, so as to compare its recall to that which would have been achieved by exhaustive manual review. We illustrate the use of these strategies within the context of TAR occasioned by litigation over accounting practices preceding the collapse of a major insurance company. Gordon V. Cormack, Maura R. Grossman, Andrew Harbison, Tom O'Halloran, Bronagh McManus |
SIGIR | 1 |
| 2023 | Technology-Assisted Review for Spreadsheets and Noisy TextabstractIn a large-scale eDiscovery effort, human assessors participated in a technology-assisted review ("TAR") process employing a modified version of Grossman and Cormack's Continuous Active Learning® ("CAL®") tool to review Excel spreadsheets and poor-quality OCR text (defined as 30-50% Markov error rate). In the legal industry, these documents are typically considered inappropriate for the application of TAR and, consequently, are usually the subject of exhaustive manual review. Our results assuage this concern by showing that a CAL TAR process, using feature engineering techniques adapted from spam filtering, can achieve satisfactory results on Excel spreadsheets and noisy OCR text. Our findings are cause for optimism in the legal industry--- adding these document classes to TAR datasets will make large reviews more manageable and less costly. Tom O'Halloran, Bronagh McManus, Andrew Harbison, Maura R. Grossman, Gordon V. Cormack |
DocEng | 5 |
| 2020 | Evaluating sentence-level relevance feedback for high-recall information retrieval
Haotian Zhang 0001, Gordon V. Cormack, Maura R. Grossman, Mark D. Smucker |
Inf. Retr. J. | 2 |
| 2019 | Unbiased Low-Variance Estimators for Precision and Related Information Retrieval Effectiveness MeasuresabstractThis work describes an estimator from which unbiased measurements of precision, rank-biased precision, and cumulative gain may be derived from a uniform or non-uniform sample of relevance assessments. Adversarial testing supports the theory that our estimator yields unbiased low-variance measurements from sparse samples, even when used to measure results that are qualitatively different from those returned by known information retrieval methods. Our results suggest that test collections using sampling to select documents for relevance assessment yield more accurate measurements than test collections using pooling, especially for the results of retrieval methods not contributing to the pool. Gordon V. Cormack, Maura R. Grossman |
SIGIR | 1 |
| 2019 | Quantifying Bias and Variance of System RankingsabstractWhen used to assess the accuracy of system rankings, Kendall's tau and other rank correlation measures conflate bias and variance as sources of error. We derive from tau a distance between rankings in Euclidean space, from which we can determine the magnitude of bias, variance, and error. Using bootstrap estimation, we show that shallow pooling has substantially higher bias and insubstantially lower variance than probability-proportional-to-size sampling, coupled with the recently released dynAP estimator. Gordon V. Cormack, Maura R. Grossman |
SIGIR | 1 |
| 2019 | Dynamic Sampling Meets PoolingabstractA team of six assessors used Dynamic Sampling (Cormack and Grossman 2018) and one hour of assessment effort per topic to form, without pooling, a test collection for the TREC 2018 Common Core Track. Later, official relevance assessments were rendered by NIST for documents selected by depth-10 pooling augmented by move-to-front (MTF) pooling (Cormack et al. 1998), as well as the documents selected by our Dynamic Sampling effort. MAP estimates rendered from dynamically sampled assessments using the xinfAP statistical evaluator are comparable to those rendered from the complete set of official assessments using the standard trec_eval tool. MAP estimates rendered using only documents selected by pooling, on the other hand, differ substantially. The results suggest that the use of Dynamic Sampling without pooling can, for an order of magnitude less assessment effort, yield information-retrieval effectiveness estimates that exhibit lower bias, lower error, and comparable ability to rank system effectiveness. Gordon V. Cormack, Haotian Zhang 0001, Nimesh Ghelani, Mustafa Abualsaud, Mark D. Smucker, Maura R. Grossman, Shahin Rahbariasl, Amira Ghenai |
SIGIR | 1 |
| 2018 | Effective User Interaction for High-Recall Retrieval: Less is MoreabstractHigh-recall retrieval --- finding all or nearly all relevant documents --- is critical to applications such as electronic discovery, systematic review, and the construction of test collections for information retrieval tasks. The effectiveness of current methods for high-recall information retrieval is limited by their reliance on human input, either to generate queries, or to assess the relevance of documents. Past research has shown that humans can assess the relevance of documents faster and with little loss in accuracy by judging shorter document surrogates, e.g.\ extractive summaries, in place of full documents. To test the hypothesis that short document surrogates can reduce assessment time and effort for high-recall retrieval, we conducted a 50-person, controlled, user study. We designed a high-recall retrieval system using continuous active learning (CAL) that could display either full documents or short document excerpts for relevance assessment. In addition, we tested the value of integrating a search engine with CAL. In the experiment, we asked participants to try to find as many relevant documents as possible within one hour. We observed that our study participants were able to find significantly more relevant documents when they used the system with document excerpts as opposed to full documents. We also found that allowing participants to compose and execute their own search queries did not improve their ability to find relevant documents and, by some measures, impaired performance. These results suggest that for high-recall systems to maximize performance, system designers should think carefully about the amount and nature of user interaction incorporated into the system. Haotian Zhang 0001, Mustafa Abualsaud, Nimesh Ghelani, Mark D. Smucker, Gordon V. Cormack, Maura R. Grossman |
CIKM | 5 |
| 2018 | The Quest for Total RecallabstractThe objective of high-recall information retrieval (HRIR) is to identify substantially all information relevant to an information need, where the consequences of missing or untimely results may have serious legal, policy, health, social, safety, defence, or financial implications. To find acceptance in practice, HRIR technologies must be more effective---and must be shown to be more effective---than current practice, according to the legal, statutory, regulatory, ethical, or professional standards governing the application domain. Such domains include, but are not limited to, electronic discovery in legal proceedings; distinguishing between public and non-public records in the curation of government archives; systematic review for meta-analysis in evidence-based medicine; separating irregularities and intentional misstatements from unintentional errors in accounting restatements; performing "due diligence" in connection with pending mergers, acquisitions, and financing transactions; and surveillance and compliance activities involving massive datasets. HRIR differs from ad hoc information retrieval where the objective is to identify the best, rather than all relevant information, and from classification or categorization where the objective is to separate relevant from non-relevant information based on previously labeled training examples. HRIR is further differentiated from established information retrieval applications by the need to quantify "substantially all relevant information"; an objective for which existing evaluation strategies and measures, such as precision and recall, are not particularly well suited. Gordon V. Cormack, Maura R. Grossman |
DocEng | 1 |
| 2018 | A System for Efficient High-Recall RetrievalabstractThe goal of high-recall information retrieval (HRIR) is to find all or nearly all relevant documents for a search topic. In this paper, we present the design of our system that affords efficient high-recall retrieval. HRIR systems commonly rely on iterative relevance feedback. Our system uses a state-of-the-art implementation of continuous active learning (CAL), and is designed to allow other feedback systems to be attached with little work. Our system allows users to judge documents as fast as possible with no perceptible interface lag. We also support the integration of a search engine for users who would like to interactively search and judge documents. In addition to detailing the design of our system, we report on user feedback collected as part of a 50 participants user study. While we have found that users find the most relevant documents when we restrict user interaction, a majority of participants prefer having flexibility in user interaction. Our work has implications on how to build effective assessment systems and what features of the system are believed to be useful by users. Mustafa Abualsaud, Nimesh Ghelani, Haotian Zhang 0001, Mark D. Smucker, Gordon V. Cormack, Maura R. Grossman |
SIGIR | 5 |
| 2018 | Beyond PoolingabstractDynamic Sampling is a novel, non-uniform, statistical sampling strategy in which documents are selected for relevance assessment based on the results of prior assessments. Unlike static and dynamic pooling methods that are commonly used to compile relevance assessments for the creation of information retrieval test collections, Dynamic Sampling yields a statistical sample from which substantially unbiased estimates of effectiveness measures may be derived. In contrast to static sampling strategies, which make no use of relevance assessments, Dynamic Sampling is able to select documents from a much larger universe, yielding superior test collections for a given budget of relevance assessments. These assertions are supported by simulation studies using secondary data from the TREC 2017 Common Core Track. Gordon V. Cormack, Maura R. Grossman |
SIGIR | 1 |
| 2017 | Navigating Imprecision in Relevance Assessments on the Road to Total Recall: Roger and MeabstractTechnology-assisted review ("TAR") systems seek to achieve "total recall"; that is, to approach, as nearly as possible, the ideal of 100% recall and 100% precision, while minimizing human review effort. The literature reports that TAR methods using relevance feedback can achieve considerably greater than the 65% recall and 65% precision reported by Voorhees as the "practical upper bound on retrieval performance... since that is the level at which humans agree with one another" (Variations in Relevance Judgments and the Measurement of Retrieval Effectiveness, 2000). This work argues that in order to build - as well as to, evaluate - TAR systems that approach 100% recall and 100% precision, it is necessary to model human assessment, not as absolute ground truth, but as an indirect indicator of the amorphous property known as "relevance." The choice of model impacts both the evaluation of system effectiveness, as well as the simulation of relevance feedback. Models are presented that better fit available data than the infallible ground-truth model. These models suggest ways to improve TAR-system effectiveness so that hybrid human-computer systems can improve on both the accuracy and efficiency of human review alone. This hypothesis is tested by simulating TAR using two datasets: the TREC 4 AdHoc collection, and a dataset consisting of 401,960 email messages that were manually reviewed and classified by a single individual, Roger, in his official capacity as Senior State Records Archivist. The results using the TREC 4 data show that TAR achieves higher recall and higher precision than the assessments by either of two independent NIST assessors, and blind adjudication of the email dataset, conducted by Roger, more than two years after his original review, shows that he could have achieved the same recall and better precision, while reviewing substantially fewer than 401,960 emails, had he employed TAR in place of exhaustive manual review. Gordon V. Cormack, Maura R. Grossman |
SIGIR | 1 |
| 2017 | Automatic and Semi-Automatic Document Selection for Technology-Assisted ReviewabstractAbstract In the TREC Total Recall Track (2015-2016), participating teams could employ either fully automatic or human-assisted ("semi-automatic") methods to select documents for relevance assessment by a simulated human reviewer. According to the TREC 2016 evaluation, the fully automatic baseline method achieved a recall-precision breakeven ("R-precision") score of 0.71, while the two semi-automatic efforts achieved scores of 0.67 and 0.51. In this work, we investigate the extent to which the observed effectiveness of the different methods may be confounded by chance, by inconsistent adherence to the Track guidelines, by selection bias in the evaluation method, or by discordant relevance assessments. We find no evidence that any of these factors could yield relative effectiveness scores inconsistent with the official TREC 2016 ranking. Maura R. Grossman, Gordon V. Cormack, Adam Roegiest |
SIGIR | 2 |
| 2017 | Ten Blue Links on MarsabstractThis paper explores a simple question: How would we provide a high-quality search experience on Mars, where the fundamental physical limit is speed-of-light propagation delays on the order of tens of minutes? On Earth, users are accustomed to nearly instantaneous responses from web services. Is it possible to overcome orders-of-magnitude longer latency to provide a tolerable user experience on Mars? In this paper, we formulate the searching from Mars problem as a tradeoff between "effort" (waiting for responses from Earth) and "data transfer" (pre-fetching or caching data on Mars). The contribution of our work is articulating this design space and presenting two case studies that explore the effectiveness of baseline techniques, using publicly available data from the TREC Total Recall and Sessions Tracks. We intend for this research problem to be aspirational as well as inspirational---even if one is not convinced by the premise of Mars colonization, there are Earth-based scenarios such as searching from rural villages in India that share similar constraints, thus making the problem worthy of exploration and attention from researchers. Charles L. A. Clarke, Gordon V. Cormack, Jimmy Lin, Adam Roegiest |
WWW | 2 |
| 2016 | Scalability of Continuous Active Learning for Reliable High-Recall Text ClassificationabstractFor finite document collections, continuous active learning ('CAL') has been observed to achieve high recall with high probability, at a labeling cost asymptotically proportional to the number of relevant documents. As the size of the collection increases, the number of relevant documents typically increases as well, thereby limiting the applicability of CAL to low-prevalence high-stakes classes, such as evidence in legal proceedings, or security threats, where human effort proportional to the number of relevant documents is justified. We present a scalable version of CAL ('S-CAL') that requires O(log N) labeling effort and O(N log N) computational effort---where N is the number of unlabeled training examples---to construct a classifier whose effectiveness for a given labeling cost compares favorably with previously reported methods. At the same time, S-CAL offers calibrated estimates of class prevalence, recall, and precision, facilitating both threshold setting and determination of the adequacy of the classifier. Gordon V. Cormack, Maura R. Grossman |
CIKM | 1 |
| 2016 | Engineering Quality and Reliability in Technology-Assisted ReviewabstractThe objective of technology-assisted review ("TAR") is to find as much relevant information as possible with reasonable effort. Quality is a measure of the extent to which a TAR method achieves this objective, while reliability is a measure of how consistently it achieves an acceptable result. We are concerned with how to define, measure, and achieve high quality and high reliability in TAR. When quality is defined using the traditional goal-post method of specifying a minimum acceptable recall threshold, the quality and reliability of a TAR method are both, by definition, equal to the probability of achieving the threshold. Assuming this definition of quality and reliability, we show how to augment any TAR method to achieve guaranteed reliability, for a quantifiable level of additional review effort. We demonstrate this result by augmenting the TAR method supplied as the baseline model implementation for the TREC 2015 Total Recall Track, measuring reliability and effort for 555 topics from eight test collections. While our empirical results corroborate our claim of guaranteed reliability, we observe that the augmentation strategy may entail disproportionate effort, especially when the number of relevant documents is low. To address this limitation, we propose stopping criteria for the model implementation that may be applied with no additional review effort, while achieving empirical reliability that compares favorably to the provably reliable method. We further argue that optimizing reliability according to the traditional goal-post method is inconsistent with certain subjective aspects of quality, and that optimizing a Taguchi quality loss function may be more apt. Gordon V. Cormack, Maura R. Grossman |
SIGIR | 1 |
| 2016 | Impact of Review-Set Selection on Human Assessment for Text ClassificationabstractIn a laboratory study, human assessors were significantly more likely to judge the same documents as relevant when they were presented for assessment within the context of documents selected using random or uncertainty sampling, as compared to relevance sampling. The effect is substantial and significant [0.54 vs. 0.42, p<0.0002] across a population of documents including both relevant and non-relevant documents, for several definitions of ground truth. This result is in accord with Smucker and Jethani's SIGIR 2010 finding that documents were more likely to be judged relevant when assessed within low-precision versus high-precision ranked lists. Our study supports the notion that relevance is malleable, and that one should take care in assuming any labeling to be ground truth, whether for training, tuning, or evaluating text classifiers. Adam Roegiest, Gordon V. Cormack |
SIGIR | 2 |
| 2016 | An Architecture for Privacy-Preserving and Replicable High-Recall Retrieval ExperimentsabstractWe demonstrate the infrastructure used in the TREC 2015 Total Recall track to facilitate controlled simulation of "assessor in the loop" high-recall retrieval experimentation. The implementation and corresponding design decisions are presented for this platform. This includes the necessary considerations to ensure that experiments are privacy-preserving when using test collections that cannot be distributed. Furthermore, we describe the use of virtual machines as a means of system submission in order to to promote replicable experiments while also ensuring the security of system developers and data providers. Adam Roegiest, Gordon V. Cormack |
SIGIR | 2 |
| 2016 | Sampling Strategies and Active Learning for Volume EstimationabstractThis paper tackles the challenge of accurately and efficiently estimating the number of relevant documents in a collection for a particular topic. One real-world application is estimating the volume of social media posts (e.g., tweets) pertaining to a topic, which is fundamental to tracking the popularity of politicians and brands, the potential sales of a product, etc. Our insight is to leverage active learning techniques to find all the "easy" documents, and then to use sampling techniques to infer the number of relevant documents in the residual collection. We propose a simple yet effective technique for determining this "switchover" point, which intuitively can be understood as the "knee" in an effort vs. recall gain curve, as well as alternative sampling strategies beyond the knee. We show on several TREC datasets and a collection of tweets that our best technique yields more accurate estimates (with the same effort) than several alternatives. Haotian Zhang 0001, Jimmy Lin, Gordon V. Cormack, Mark D. Smucker |
SIGIR | 3 |
| 2015 | Multi-Faceted Recall of Continuous Active Learning for Technology-Assisted ReviewabstractContinuous active learning achieves high recall for technology-assisted review, not only for an overall information need, but also for various facets of that information need, whether explicit or implicit. Through simulations using Cormack and Grossman's TAR Evaluation Toolkit (SIGIR 2014), we show that continuous active learning, applied to a multi-faceted topic, efficiently achieves high recall for each facet of the topic. Our results assuage the concern that continuous active learning may achieve high overall recall at the expense of excluding identifiable categories of relevant information. Gordon V. Cormack, Maura R. Grossman |
SIGIR | 1 |
| 2015 | Impact of Surrogate Assessments on High-Recall RetrievalabstractWe are concerned with the effect of using a surrogate assessor to train a passive (i.e., batch) supervised-learning method to rank documents for subsequent review, where the effectiveness of the ranking will be evaluated using a different assessor deemed to be authoritative. Previous studies suggest that surrogate assessments may be a reasonable proxy for authoritative assessments for this task. Nonetheless, concern persists in some application domains---such as electronic discovery---that errors in surrogate training assessments will be amplified by the learning method, materially degrading performance. We demonstrate, through a re-analysis of data used in previous studies, that, with passive supervised-learning methods, using surrogate assessments for training can substantially impair classifier performance, relative to using the same deemed-authoritative assessor for both training and assessment. In particular, using a single surrogate to replace the authoritative assessor for training often yields a ranking that must be traversed much lower to achieve the same level of recall as the ranking that would have resulted had the authoritative assessor been used for training. We also show that steps can be taken to mitigate, and sometimes overcome, the impact of surrogate assessments for training: relevance assessments may be diversified through the use of multiple surrogates; and, a more liberal view of relevance can be adopted by having the surrogate label borderline documents as relevant. By taking these steps, rankings derived from surrogate assessments can match, and sometimes exceed, the performance of the ranking that would have been achieved, had the authority been used for training. Finally, we show that our results still hold when the role of surrogate and authority are interchanged, indicating that the results may simply reflect differing conceptions of relevance between surrogate and authority, as opposed to the authority having special skill or knowledge lacked by the surrogate. Adam Roegiest, Gordon V. Cormack, Charles L. A. Clarke, Maura R. Grossman |
SIGIR | 2 |
| 2014 | Evaluation of machine-learning protocols for technology-assisted review in electronic discoveryabstractAbstract Using a novel evaluation toolkit that simulates a human reviewer in the loop, we compare the effectiveness of three machine-learning protocols for technology-assisted review as used in document review for discovery in legal proceedings. Our comparison addresses a central question in the deployment of technology-assisted review: Should training documents be selected at random, or should they be selected using one or more non-random methods, such as keyword search or active learning? On eight review tasks -- four derived from the TREC 2009 Legal Track and four derived from actual legal matters -- recall was measured as a function of human review effort. The results show that entirely non-random training methods, in which the initial training documents are selected using a simple keyword search, and subsequent training documents are selected by active learning, require substantially and significantly less human review effort (P<0.01) to achieve any given level of recall, than passive learning, in which the machine-learning algorithm plays no role in the selection of training documents. Among passive-learning methods, significantly less human review effort (P<0.01) is required when keywords are used instead of random sampling to select the initial training documents. Among active-learning methods, continuous active learning with relevance feedback yields generally superior results to simple active learning with uncertainty sampling, while avoiding the vexing issue of "stabilization" -- determining when training is adequate, and therefore may stop. Gordon V. Cormack, Maura R. Grossman |
SIGIR | 1 |
| 2011 | Efficient and effective spam filtering and re-ranking for large web datasets
Gordon V. Cormack, Mark D. Smucker, Charles L. A. Clarke |
Inf. Retr. | 1 |
| 2010 | Semi-supervised spam filtering using aggressive consistency learningabstractA graph based semi-supervised method for email spam filtering, based on the local and global consistency method, yields low error rates with very few labeled examples. The motivating application of this method is spam filters with access to very few labeled message. For example, during the initial deployment of a spam filter, only a handful of labeled examples are available but unlabeled examples are plentiful. We demonstrate the performance of our approach on TREC 2007 and CEAS 2008 email corpora. Our results compare favorably with the best-known methods, using as few as just two labeled examples: one spam and one non-spam. Mona Mojdeh, Gordon V. Cormack |
SIGIR | 2 |
| 2009 | Genre-based decomposition of email class noiseabstractCorruption of data by class-label noise is an important practical concern impacting many classification problems. Studies of data cleaning techniques often assume a uniform label noise model, however, which is seldom realized in practice. Relatively little is understood, as to how the natural label noise distribution can be measured or simulated. Using email spam-filtering data, we demonstrate that class noise can have substantial content specific bias. We also demonstrate that noise detection techniques based on classifier confidence tend to identify instances that human assessors are likely to label in error. We show that genre modeling can be very informative in identifying potential areas of mislabeling. Moreover, we are able to show that genre decomposition can also be used to substantially improve spam filtering accuracy, with our results outperforming the best published figures for the trec05-p1 and ceas-2008 benchmark collections. Alek Kolcz, Gordon V. Cormack |
KDD | 2 |
| 2009 | On the relative age of spam and ham training samples for email filteringabstractEmail spam filters are commonly trained on a sample of spam and ham (non-spam) messages. We investigate the effect on filter performance of using samples of spam and ham messages sent months before those to be filtered. Our results show that filter performance deteriorates with the overall age of spam and ham samples, but at different rates. Spam and ham samples of different ages may be mixed to advantage, provided temporal cues are elided Gordon V. Cormack, José-Marcio Martins da Cruz |
SIGIR | 1 |
| 2009 | Reciprocal rank fusion outperforms condorcet and individual rank learning methodsabstractReciprocal Rank Fusion (RRF), a simple method for combining the document rankings from multiple IR systems, consistently yields better results than any individual system, and better results than the standard method Condorcet Fuse. This result is demonstrated by using RRF to combine the results of several TREC experiments, and to build a meta-learner that ranks the LETOR 3 dataset better than any previously reported method Gordon V. Cormack, Charles L. A. Clarke, Stefan Büttcher |
SIGIR | 1 |
| 2009 | Spam filter evaluation with imprecise ground truthabstractWhen trained and evaluated on accurately labeled datasets, online email spam filters are remarkably effective, achieving error rates an order of magnitude better than classifiers in similar applications. But labels acquired from user feedback or third-party adjudication exhibit higher error rates than the best filters -- even filters trained using the same source of labels. It is appropriate to use naturally occuring labels -- including errors -- as training data in evaluating spam filters. Erroneous labels are problematic, however, when used as ground truth to measure filter effectiveness. Any measurement of the filter's error rate will be augmented and perhaps masked by the label error rate. Using two natural sources of labels, we demonstrate automatic and semi-automatic methods that reduce the influence of labeling errors on evaluation, yielding substantially more precise measurements of true filter error rates. Gordon V. Cormack, Alek Kolcz |
SIGIR | 1 |
| 2009 | Swapping documents and terms
Charles L. A. Clarke, Gordon V. Cormack, Thomas R. Lynam, Chris Buckley, Donna K. Harman |
Inf. Retr. | 2 |
| 2008 | Novelty and diversity in information retrieval evaluationabstractEvaluation measures act as objective functions to be optimized by information retrieval systems. Such objective functions must accurately reflect user requirements, particularly when tuning IR systems and learning ranking functions. Ambiguity in queries and redundancy in retrieved documents are poorly reflected by current evaluation measures. In this paper, we present a framework for evaluation that systematically rewards novelty and diversity. We develop this framework into a specific evaluation measure, based on cumulative gain. We demonstrate the feasibility of our approach using a test collection based on the TREC question answering track. Charles L. A. Clarke, Maheedhar Kolla, Gordon V. Cormack, Olga Vechtomova, Azin Ashkan, Stefan Büttcher, Ian MacKinnon |
SIGIR | 3 |
| 2008 | Semi-supervised spam filtering: does it work?abstractThe results of the 2006 ECML/PKDD Discovery Challenge suggest that semi-supervised learning methods work well for spam filtering when the source of available labeled examples differs from those to be classified. We have attempted to reproduce these results using data from the 2005 and 2007 TREC Spam Track, and have found the opposite effect: methods like self-training and transductive support vector machines yield inferior classifiers to those constructed using supervised learning on the labeled data alone. We investigate differences between the ECML/PKDD and TREC data sets and methodologies that may account for the opposite results. Mona Mojdeh, Gordon V. Cormack |
SIGIR | 2 |
| 2007 | Spam filtering for short messagesabstractWe consider the problem of content-based spam filtering for short text messages that arise in three contexts: mobile (SMS) communication, blog comments, and email summary information such as might be displayed by a low-bandwidth client. Short messages often consist of only a few words, and therefore present a challenge to traditional bag-of-words based spam filters. Using three corpora of short messages and message fields derived from real SMS, blog, and spam messages, we evaluate feature-based and compression-model-based spam filters. We observe that bag-of-words filters can be improved substantially using different features, while compression-model filters perform quite well as-is. We conclude that content filtering for short messages is surprisingly effective. Gordon V. Cormack, José María Gómez Hidalgo, Enrique Puertas Sanz |
CIKM | 1 |
| 2007 | A larger decidable semiunification problemabstractWe present a graph-theoretic framework in which to study instances of the semiunification problem (SUP), which is known to be undecidable, but has several known and important decidable subsets. One such subset, the acyclic semiunification problem (ASUP), has proved useful in the study of polymorphic type inference. We present graph-theoretic criteria in our framework that exactly characterize the ASUP acyclicity constraint. We then use our framework to find a decidable subset of SUP (which we call R-ASUP), which has a more natural description than ASUP, and strictly contains it. Brad Lushman, Gordon V. Cormack |
PPDP | 2 |
| 2007 | Feature engineering for mobile (SMS) spam filteringabstractMobile spam in an increasing threat that may be addressed using filtering systems like those employed against email spam. We believe that email filtering techniques require some adaptation to reach good levels of performance on SMS spam, especially regarding message representation. In order to test this assumption, we have performed experiments on SMS filtering using top performing email spam filters on mobile spam messages using a suitable feature representation, with results supporting our hypothesis. Gordon V. Cormack, José María Gómez Hidalgo, Enrique Puertas Sanz |
SIGIR | 1 |
| 2007 | Validity and power of t-test for comparing MAP and GMAPabstractWe examine the validity and power of the t-test, Wilcoxon test, and sign test in determining whether or not the difference in performance between two IR systems is significant. Empirical tests conducted on subsets of the TREC2004 Robust Retrieval collection indicate that the p-values computed by these tests for the difference in mean average precision (MAP) between two systems are very accurate fora wide range of sample sizes and significance estimates. Similarly, these tests have good power, with the t-test proving superior overall. The t-test is also valid for comparing geometric mean average precision (GMAP), exhibiting slightly superior accuracy and slightly inferior power than for MAPcomparison. Gordon V. Cormack, Thomas R. Lynam |
SIGIR | 1 |
| 2007 | Power and bias of subset pooling strategiesabstractWe define a method to estimate the random and systematic errors resulting from incomplete relevance assessments.Mean Average Precision (MAP) computed over a large number of topics with a shallow assessment pool substantially outperforms -- for the same adjudication effort MAP computed over fewer topics with deeper pools, and [email protected] computed with pools of the same depth. Move-to-front pooling,previously reported to yield substantially better rank correlation, yields similar power, and lower bias, compared tofixed-depth pooling. Gordon V. Cormack, Thomas R. Lynam |
SIGIR | 1 |
| 2007 | Online supervised spam filter evaluationabstractEleven variants of six widely used open-source spam filters are tested on a chronological sequence of 49086 e-mail messages received by an individual from August 2003 through March 2004. Our approach differs from those previously reported in that the test set is large, comprises uncensored raw messages, and is presented to each filter sequentially with incremental feedback. Misclassification rates and Receiver Operating Characteristic Curve measurements are reported, with statistical confidence intervals. Quantitative results indicate that content-based filters can eliminate 98% of spam while incurring 0.1% legitimate email loss. Qualitative results indicate that the risk of loss depends on the nature of the message, and that messages likely to be lost may be those that are less critical. More generally, our methodology has been encapsulated in a free software toolkit, which may used to conduct similar experiments. Gordon V. Cormack, Thomas R. Lynam |
ACM Trans. Inf. Syst. | 1 |
| 2006 | Statistical precision of information retrieval evaluationabstractWe introduce and validate bootstrap techniques to compute confidence intervals that quantify the effect of test-collection variability on average precision (AP) and mean average precision (MAP) IR effectiveness measures. We consider the test collection in IR evaluation to be a representative of a population of materially similar collections, whose documents are drawn from an infinite pool with similar characteristics. Our model accurately predicts the degree of concordance between system results on randomly selected halves of the TREC-6 ad hoc corpus. We advance a framework for statistical evaluation that uses the same general framework to model other sources of chance variation as a source of input for meta-analysis techniques. Gordon V. Cormack, Thomas R. Lynam |
SIGIR | 1 |
| 2006 | On-line spam filter fusionabstractWe show that a set of independently developed spam filters may be combined in simple ways to provide substantially better filtering than any of the individual filters. The results of fifty-three spam filters evaluated at the TREC 2005 Spam Track were combined post-hoc so as to simulate the parallel on-line operation of the filters. The combined results were evaluated using the TREC methodology, yielding more than a factor of two improvement over the best filter. The simplest method -- averaging the binary classifications returned by the individual filters -- yields a remarkably good result. A new method -- averaging log-odds estimates based on the scores returned by the individual filters -- yields a somewhat better result, and provides input to SVM- and logistic-regression-based stacking methods. The stacking methods appear to provide further improvement, but only for very large corpora. Of the stacking methods, logistic regression yields the better result. Finally, we show that it is possible to select a priori small subsets of the filters that, when combined, still outperform the best individual filter by a substantial margin. Thomas R. Lynam, Gordon V. Cormack, David R. Cheriton |
SIGIR | 2 |
| 2006 | Spam Filtering Using Statistical Data Compression ModelsabstractSpam filtering poses a special problem in text categorization, of which the defining characteristic is that filters face an active adversary, which constantly attempts to evade filtering. Since spam evolves continuously and most practical applications are based on online user feedback, the task calls for fast, incremental and robust learning algorithms. In this paper, we investigate a novel approach to spam filtering based on adaptive statistical data compression models. The nature of these models allows them to be employed as probabilistic text classifiers based on character-level or binary sequences. By modeling messages as sequences, tokenization and other error-prone preprocessing steps are omitted altogether, resulting in a method that is very robust. The models are also fast to construct and incrementally updateable. We evaluate the filtering performance of two different compression algorithms; dynamic Markov compression and prediction by partial matching. The results of our empirical evaluation indicate that compression models outperform currently established spam filters, as well as a number of methods proposed in previous studies. Andrej Bratko, Gordon V. Cormack, Bogdan Filipic, Thomas R. Lynam, Blaz Zupan |
J. Mach. Learn. Res. | 2 |
| 2004 | A multi-system analysis of document and term selection for blind feedbackabstractExperiments were conducted to explore the impact of combining various components of eight leading information retrieval systems. Each system demonstrated improved effectiveness with the use of blind feedback, in which the results of a preliminary retrieval step were used to augment the efficacy of a secondary retrieval step. The hybrid combination of primary and secondary retrieval steps from different systems in a number of cases yielded better effectiveness than either of the constituent systems alone. This positive combining effect was observed when entire documents were passed between the two retrieval steps, but not when only the expansion terms were passed. Several combinations of primary and secondary retrieval steps were fused using the CombMNZ algorithm; all yielded significant effectiveness improvement over the individual systems, with the best yielding a an improvement of 13% (p = 10-6) over the best individual system and an improvement of 4% (p = 10-5) over a simple fusion of the eight systems. Thomas R. Lynam, Chris Buckley, Charles L. A. Clarke, Gordon V. Cormack |
CIKM | 4 |
| 2003 | Proof of correctness of Ressel's adOPTed algorithm
Brad Lushman, Gordon V. Cormack |
Inf. Process. Lett. | 2 |
| 2002 | The impact of corpus size on question answering performanceabstractUsing our question answering system, questions from the TREC 2001 evaluation were executed over a series of Web data collections, with the sizes of the collections increasing from 25 gigabytes up to nearly a terabyte. Charles L. A. Clarke, Gordon V. Cormack, M. Laszlo, Thomas R. Lynam, Egidio L. Terra |
SIGIR | 2 |
| 2001 | Exploiting Redundancy in Question AnsweringabstractOur goal is to automatically answer brief factual questions of the form ``When was the Battle of Hastings?'' or ``Who wrote The Wind in the Willows?''. Since the answer to nearly any such question can now be found somewhere on the Web, the problem reduces to finding potential answers in large volumes of data and validating their accuracy. We apply a method for arbitrary passage retrieval to the first half of the problem and demonstrate that answer redundancy can be used to address the second half. The success of our approach depends on the idea that the volume of available Web data is large enough to supply the answer to most factual questions multiple times and in multiple contexts. A query is generated from a question and this query is used to select short passages that may contain the answer from a large collection of Web data. These passages are analyzed to identify candidate answers. The frequency of these candidates within the passages is used to ``vote'' for the most likely answer. The approach is experimentally tested on questions taken from the TREC-9 question-answering test collection. As an additional demonstration, the approach is extended to answer multiple choice trivia questions of the form typically asked in trivia quizzes and television game shows. Charles L. A. Clarke, Gordon V. Cormack, Thomas R. Lynam |
SIGIR | 2 |
| 2000 | Relevance ranking for one to three term queries
Charles L. A. Clarke, Gordon V. Cormack, Elizabeth A. Tudhope |
Inf. Process. Manag. | 2 |
| 2000 | Passage-based query refinement (MultiText experiments for TREC-6)
Gordon V. Cormack, Charles L. A. Clarke, Christopher R. Palmer, Samuel S. L. To |
Inf. Process. Manag. | 1 |
| 2000 | Shortest-substring retrieval and rankingabstractWe present a model for arbitrary passage retrieval using Boolean queries. The model is applied to the task of ranking documents, or other structural elements, in the order of their expected relevance. Features such as phrase matching, truncation, and stemming integrate naturally into the model. Properties of Boolean algebra are obeyed, and the exact-match semantics of Boolean retrieval are preserved. Simple inverted-list file structures provide an efficient implementation. Retrieval effectiveness is comparable to that of standard ranking techniques. Since global statistics are not used, the method is of particular value in distributed environments. Since ranking is based on arbitrary passages, the structural elements to be ranked may be specified at query time and do not need to be restricted to predefined elements. Charles L. A. Clarke, Gordon V. Cormack |
ACM Trans. Inf. Syst. | 2 |
| 1999 | The MultiText Retrieval System (demonstration abstract)abstractNo abstract available. Gordon V. Cormack, Charles L. A. Clarke, Christopher R. Palmer, Robert C. Good |
SIGIR | 1 |
| 1999 | Estimating Precision by Random Sampling (poster abstract)abstractNo abstract available. Gordon V. Cormack, Ondrej Lhoták, Christopher R. Palmer |
SIGIR | 1 |
| 1998 | Operation Transforms for a Distributed Shared SpreadsheetabstractThe Distributed Operation Dansform (dOPT), pr~ posed by Efi and Gibbs, is used to dfie concurrently updatable shared objects.Eh and Gibbs give the OP eration tr~orrus that d&e a simple shared text editor supporting single chmacter insertions and deletions on a hear btim.We report here on the construction of operation transforms for a more sophisticated group nwe appficatiou a shared spreadsheet.We identify a set of abstract operations that &aracterize the operations on a spreadsheet.Using Cormack's Cdtius for Concurrent Update, which extends and corrects dOPT, m'egive the tr~forms on these operations necessary to d~e a shared spreadsheet.~re use the transforms to btid a shared version of SC, the Unix spreadsheet due to Goshg. Christopher R. Palmer, Gordon V. Cormack |
CSCW | 2 |
| 1998 | Efficient Construction of Large Test CollectionsabstractTest collections with a million or more documents are needed for the evaluation of modern information retrieval systems.Yet their construction requires a great deal of effort.Judgements must be rendered as to whether or not documents are relevant to each of a set of queries.Exhaustive judging, in which every document is examined and a judgement rendered, is infeasible for collections of this size.Current practice is represented by the "pooling method", as used in the TREC conference series, in which only the first k documents from each of a number of sources are judged.We propose two methods, Intemctive Searching and Judging and Moveto-front Pooling, that yield effective test collections while requiring many fewer judgements.Interactive Searching and Judging selects documents to be judged using an interactive search system, and may be used by a small research team to develop an effective test collection using minimal resources.Move-to-Front Pooling directly improves on the standard pooling method by using a variable number of documents from each source depending on its retrieval performance.Move-to-Front Pooling would be an appropriate replacement for the standard pooling method in future collection development efforts involving many independent groups. Gordon V. Cormack, Christopher R. Palmer, Charles L. A. Clarke |
SIGIR | 1 |
| 1997 | On the Use of Regular Expressions for Searching TextabstractThe use of regular expressions for text search is widely known and well understood. It is then surprising that the standard techniques and tools prove to be of limited use for searching structured text formatted with SGML or similar markup languages. Our experience with structured text search has caused us to reexamine the current practice. The generally accepted rule of “leftmost longest match” is an unfortunate choice and is at the root of the difficulties. We instead propose a rule which is semantically cleaner. This rule is generally applicable to a variety of text search applications, including source code analysis, and has interesting properties in its own right. We have written a publicly available search tool implementing the theory in the article, which has proved valuable in a variety of circumstances. Charles L. A. Clarke, Gordon V. Cormack |
ACM Trans. Program. Lang. Syst. | 2 |
| 1996 | Kinded Type Inference for Parametric Overloading
Dominic Duggan, Gordon V. Cormack, John Ophel |
Acta Informatica | 2 |
| 1995 | A Calculus for Concurrent Update (Abstract)abstractNo abstract available. Gordon V. Cormack |
PODC | 1 |
| 1995 | An Algebra for Structured Text Search and a Framework for its ImplementationabstractA query algebra is presented that expresses searches on structured text. In addition to traditional full-text boolean queries that search a pre-defined collection of documents, the algebra permits queries that harness document structure. The algebra manipulates arbitrary intervals of text, which are recognized in the text from implicit or explicit markup. The algebra has seven operators, which combined intervals to yield new ones: containing, not containing, contained in, not contained in, one of, both of, followed by. The ultimate result of a query is the set of intervals that satisfy it. An implementation framework is given based on four primitive access functions. Each access function finds the solution to a query nearest to a given position in the database. Recursive definitions for the seven operators are given in terms of these access functions. Search time is at worst proportional to the time required to evaluate the access functions for occurrences of the elementary terms in a query. Charles L. A. Clarke, Gordon V. Cormack, Forbes J. Burkowski |
Comput. J. | 2 |
| 1994 | Access Control for Private Declarations in Ada
Gordon V. Cormack |
Comput. Lang. | 2 |
| 1992 | Constructing Word-Based Text Compression AlgorithmsabstractText compression algorithms are normally defined in terms of a source alphabet Sigma of 8-bit ASCII codes. The authors consider choosing Sigma to be an alphabet whose symbols are the words of English or, in general, alternate maximal strings of alphanumeric characters and nonalphanumeric characters. The compression algorithm would be able to take advantage of longer-range correlations between words and thus achieve better compression. The large size of Sigma leads to some implementation problems, but these are overcome to construct word-based LZW, word-based adaptive Huffman, and word-based context modelling compression algorithms.> R. Nigel Horspool, Gordon V. Cormack |
Data Compression Conference | 2 |
| 1990 | Use of Perfect Hashing in a Paged Memory Management Unit
Forbes J. Burkowski, Gordon V. Cormack |
ICPP (1) | 2 |
| 1990 | Type-Dependent Parameter InferenceabstractAn algorithm is presented to infer the type and operation parameters of polymorphic functions. Operation parameters are named and typed at the function definition, but are selected from the set of overloaded definitions available wherever the function is used. These parameters are always implicit, implying that the complexity of using a function does not increase with the generality of its type. Gordon V. Cormack, Andrew K. Wright |
PLDI | 1 |
| 1990 | Modular Attribute GrammarsabstractAttribute grammars provide a formal declarative notation for describing the semantics and translation of programming languages. Describing any real programming language is a significant software engineering challenge. From a software engineering viewpoint, current notations for attribute grammars have two flaws: tedious repetition of essentially the same attribute computations is inevitable, and the various components of the description cannot be decomposed into modules - they must be merged (and hence closely coupled) with the syntax specification. This paper describes a tool that generates attribute grammars from pattern-oriented specifications. These specifications can be grouped according to the separation of concerns arising from individual aspects of the compilation process. Implementation and use of the attribute grammar generate MAGGIE is described. G. D. P. Dueck, Gordon V. Cormack |
Comput. J. | 2 |
| 1989 | Architectural Support for Synchronous Task CommunicationabstractThis paper describes the motivation for a set of intertask communication primitives, the hardware support of these primitives, the architecture used in the Sylvan project which studies these issues, and the experience gained from various experiments conducted in this area. We start by describing how these facilities have been implemented in a multiprocessor configuration that utilizes a shared backplane. This configuration represents a single node in the system. The latter part of the paper discusses a distributed multiple node system and the extension of the primitives that are used in this expanded environment. Forbes J. Burkowski, Gordon V. Cormack, G. D. P. Dueck |
ASPLOS | 2 |
| 1989 | An LR Substring Parser for Noncorrecting Syntax Error RecoveryabstractFor a context-free grammar G, a construction is given to produce an LR parser that recognizes any substring of the language generated by G. The construction yields a conflict-free (deterministic) parser for the bounded context class of grammars (Floyd, 1964). The same construction yields either a left-to-right or right-to-left substring parser, as required to implement Non-correcting Syntax Error Recovery as proposed by Richter (1985). Experience in constructing a substring parser for Pascal is described. Gordon V. Cormack |
PLDI | 1 |
| 1989 | Scannerless NSLR(1) Parsing of Programming LanguagesabstractThe disadvantages of traditional two-phase parsing (a scanner phase preprocessing input for a parser phase) are discussed. We present metalanguage enhancements for context-free grammars that allow the syntax of programming languages to be completely described in a single grammar. The enhancements consist of two new grammar rules, the exclusion rule, and the adjacency-restriction rule. We also present parser construction techniques for building parsers from these enhanced grammars, that eliminate the need for a scanner phase. Daniel J. Salomon, Gordon V. Cormack |
PLDI | 2 |
| 1988 | A Micro-Kernel for Concurrency in CabstractAbstract A micro‐kernel that supports concurrent execution of C procedures within a single user process is described. A micro‐kernel provides only four primitives, which have been used to build a number of higher‐level abstractions, including support for distributed processing. The micro‐kernel differs from other efforts in that it is small and efficient, it is written entirely as a non‐privileged user program, and it provides fine‐grained unpredictable interleaving of execution. Gordon V. Cormack |
Softw. Pract. Exp. | 1 |
| 1987 | Data Compression Using Dynamic Markov ModellingabstractA method of dynamically constructing Markov chain models that describe the characteristics of binary messages is developed. Such models can be used to predict future message characters and can therefore be used as a basis for data compression. To this end, the Markov modelling technique is combined with Guazzo's arithmetic coding scheme to produce a powerful method of data compression. The method has the advantage of being adaptive: messages may be encoded or decoded with just a single pass through the data. Experimental results reported here indicate that the Markov modelling approach generally achieves much better data compression than that observed with competing methods on typical computer data. Gordon V. Cormack, R. Nigel Horspool |
Comput. J. | 1 |
| 1987 | Structured Program Lookahead
Thomas Strothotte, Gordon V. Cormack |
Comput. Lang. | 2 |
| 1987 | Hashing as a Compaction Technique for LR Parser TablesabstractAbstract Authors of papers on LR parser table compaction and authors of books on compiler construction appear to have either overlooked or discounted the possibility of using hashing. In fact, hashing is easy to implement as a compaction technique and gives reasonable performance. It produces tables that are as compact as some of the other techniques reported in the literature while permitting efficient table lookups. R. Nigel Horspool, Gordon V. Cormack |
Softw. Pract. Exp. | 2 |
| 1985 | Practical Perfect HashingabstractA practical method is presented that permits retrieval from a table in constant time. The method is suitable for large tables and consumes, in practice, O(n) space for n table elements. In addition, the table and the hashing function can be constructed in O(n) expected time. Variations of the method that offer different compromises between storage usage and update time are presented. Gordon V. Cormack, R. Nigel Horspool, Matthias Kaiserswerth |
Comput. J. | 1 |
| 1984 | Algorithms for Adaptive Huffman Codes
Gordon V. Cormack, R. Nigel Horspool |
Inf. Process. Lett. | 1 |