EDBT 2026 Demo / reviewers in the wild / expert
Maura R. Grossman
dblp:122/5875
· DBLP profile ↗
22ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0003-2279-4262ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 18 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AI ethics education: A scoping review of pedagogy, curriculum, and assessmentabstractBackground Artificial intelligence (AI) is increasingly embedded in social and institutional decision-making, creating demand for ethically literate practitioners. Universities have responded by introducing AI ethics instruction, but the structure, content, pedagogy, and evaluation of these efforts remain unevenly documented. Objective To map and synthesize research on university level AI ethics education by characterizing course design, pedagogy, ethical themes, and assessment methods, and identifying evidence gaps that limit knowledge consolidation and instructional refinement. Methods We conducted a scoping review using Continuous Active Learning to screen 50,766 records up to 2024. 43 studies met inclusion after title, abstract, and full text review. We coded instructional design, curricular themes, pedagogical methods, and evaluation approaches using descriptive frequency counts and qualitative synthesis. Results Most included studies were conceptual or descriptive, with relatively few empirical evaluations. Instruction was concentrated in computing and engineering and primarily targeted undergraduate learners. Ethics content was more often embedded within technical courses than delivered as standalone offerings. Reported pedagogy relied heavily on lecture and case-based discussion, with fewer studies describing participatory formats such as simulations or role-play. Curricular emphasis clustered around bias/fairness and privacy, with comparatively less attention to governance, explainability, and trust. Evaluation most often relied on self-report and reflective methods, while validated instruments and performance-based assessments were less common, and behavioral or applied outcomes were rarely assessed. Conclusions The literature suggests a field oriented toward awareness-building more than measurable ethical competence. Clearer competency claims, stronger assessment transparency, and greater alignment between instructional design and evaluation would improve comparability across studies and support evidence-informed course development. Calvin Hillis, Maushumi Bhattacharjee, Batool AlMousawi, Riley Martens, Tarik Eltanahy, Sara Ono, Marcus Hui, Ba' Pham, Michelle Swab, Gordon V. Cormack, Maura R. Grossman, Ebrahim Bagheri, Zack Marshall |
Inf. Process. Manag. | 11 |
| 2025 | Precarity and Solidarity: Preliminary Results on a Study of Queer and Disabled Fiction Writers' Experiences with Generative AIabstractWe present a mixed-methods study of professional fiction writers' experiences with generative AI (genAI), primarily focused on queer and disabled writers. Queer and disabled writers are markedly more pessimistic than others about the impact of genAI on their industry, although pessimism is the majority attitude for all groups. We explore how genAI exacerbates existing causes of precarity for writers, reasons why writers are opposed to its use, and strategies used by marginalized fiction writers to safeguard their industry. Carolyn Elizabeth Lamb, Dan Brown 0001, Maura R. Grossman |
IJCAI | 3 |
| 2025 | Self-Disclosure and Beyond: Takeaways from an Online and In-Person Computing Ethics CourseabstractWe evaluate the amount and nature of self-disclosure in two versions of a 400-level computing ethics course focusing on discrimination and surveillance. The study involved 30 participants enrolled in two identical course offerings, taught by the same pair of instructors, but delivered in different formats: online versus in-person. Our analysis concentrated on the extent and contents of self-disclosure by both students and instructors. By using both quantitative and qualitative methods, we observed a higher prevalence of self-disclosure by both students and instructors in the online section. Notably, an analysis of demographic data revealed that minority group members were particularly active in self-disclosure in both formats. Overall, our findings suggest that an online setting may be more effective for delivering computing ethics courses where a primary goal is increasing open discussion and self-disclosure among participants. Helen Weixu Chen, Maura R. Grossman, Dan Brown 0001 |
SIGCSE (2) | 2 |
| 2024 | Unbiased Validation of Technology-Assisted Review for eDiscoveryabstractAlthough it is well established that recall estimates are valid only when based on independent relevance assessments, and useful only to compare the relative effectiveness of competing methods, these conditions are seldom met when validating eDiscovery efforts in litigation. We present two unbiased validation strategies that embed blind relevance assessments into a technology-assisted review (TAR) process, so as to compare its recall to that which would have been achieved by exhaustive manual review. We illustrate the use of these strategies within the context of TAR occasioned by litigation over accounting practices preceding the collapse of a major insurance company. Gordon V. Cormack, Maura R. Grossman, Andrew Harbison, Tom O'Halloran, Bronagh McManus |
SIGIR | 2 |
| 2024 | Mining User Study Data to Judge the Merit of a Model for Supporting User-Specific Explanations of AI SystemsabstractABSTRACT In this paper, we present a model for supporting user‐specific explanations of AI systems. We then discuss a user study that was conducted to gauge whether the decisions for adjusting output to users with certain characteristics was confirmed to be of value to participants. We focus on the merit of having explanations attuned to particular psychological profiles of users, and the value of having different options for the level of explanation that is offered (including allowing for no explanation, as one possibility). Following the description of the study, we present an approach for mining data from user participant responses in order to determine whether the model that was developed for varying the output to users was well‐founded. While our results in this respect are preliminary, we explain how using varied machine learning methods is of value as a concrete step toward validation of specific approaches for AI explanation. We conclude with a discussion of related work and some ideas for new directions with the research, in the future. Owen Chambers, Robin Cohen, Maura R. Grossman, Liam Hebert, Elias Awad |
Comput. Intell. | 3 |
| 2023 | Technology-Assisted Review for Spreadsheets and Noisy TextabstractIn a large-scale eDiscovery effort, human assessors participated in a technology-assisted review ("TAR") process employing a modified version of Grossman and Cormack's Continuous Active Learning® ("CAL®") tool to review Excel spreadsheets and poor-quality OCR text (defined as 30-50% Markov error rate). In the legal industry, these documents are typically considered inappropriate for the application of TAR and, consequently, are usually the subject of exhaustive manual review. Our results assuage this concern by showing that a CAL TAR process, using feature engineering techniques adapted from spam filtering, can achieve satisfactory results on Excel spreadsheets and noisy OCR text. Our findings are cause for optimism in the legal industry--- adding these document classes to TAR datasets will make large reviews more manageable and less costly. Tom O'Halloran, Bronagh McManus, Andrew Harbison, Maura R. Grossman, Gordon V. Cormack |
DocEng | 4 |
| 2021 | Are Machine Learning Corpora "Fair Dealing" under Canadian Law?
Dan Brown 0001, Lauren Byl, Maura R. Grossman |
ICCC | 3 |
| 2020 | Evaluating sentence-level relevance feedback for high-recall information retrieval
Haotian Zhang 0001, Gordon V. Cormack, Maura R. Grossman, Mark D. Smucker |
Inf. Retr. J. | 3 |
| 2019 | Unbiased Low-Variance Estimators for Precision and Related Information Retrieval Effectiveness MeasuresabstractThis work describes an estimator from which unbiased measurements of precision, rank-biased precision, and cumulative gain may be derived from a uniform or non-uniform sample of relevance assessments. Adversarial testing supports the theory that our estimator yields unbiased low-variance measurements from sparse samples, even when used to measure results that are qualitatively different from those returned by known information retrieval methods. Our results suggest that test collections using sampling to select documents for relevance assessment yield more accurate measurements than test collections using pooling, especially for the results of retrieval methods not contributing to the pool. Gordon V. Cormack, Maura R. Grossman |
SIGIR | 2 |
| 2019 | Quantifying Bias and Variance of System RankingsabstractWhen used to assess the accuracy of system rankings, Kendall's tau and other rank correlation measures conflate bias and variance as sources of error. We derive from tau a distance between rankings in Euclidean space, from which we can determine the magnitude of bias, variance, and error. Using bootstrap estimation, we show that shallow pooling has substantially higher bias and insubstantially lower variance than probability-proportional-to-size sampling, coupled with the recently released dynAP estimator. Gordon V. Cormack, Maura R. Grossman |
SIGIR | 2 |
| 2019 | Dynamic Sampling Meets PoolingabstractA team of six assessors used Dynamic Sampling (Cormack and Grossman 2018) and one hour of assessment effort per topic to form, without pooling, a test collection for the TREC 2018 Common Core Track. Later, official relevance assessments were rendered by NIST for documents selected by depth-10 pooling augmented by move-to-front (MTF) pooling (Cormack et al. 1998), as well as the documents selected by our Dynamic Sampling effort. MAP estimates rendered from dynamically sampled assessments using the xinfAP statistical evaluator are comparable to those rendered from the complete set of official assessments using the standard trec_eval tool. MAP estimates rendered using only documents selected by pooling, on the other hand, differ substantially. The results suggest that the use of Dynamic Sampling without pooling can, for an order of magnitude less assessment effort, yield information-retrieval effectiveness estimates that exhibit lower bias, lower error, and comparable ability to rank system effectiveness. Gordon V. Cormack, Haotian Zhang 0001, Nimesh Ghelani, Mustafa Abualsaud, Mark D. Smucker, Maura R. Grossman, Shahin Rahbariasl, Amira Ghenai |
SIGIR | 6 |
| 2018 | Effective User Interaction for High-Recall Retrieval: Less is MoreabstractHigh-recall retrieval --- finding all or nearly all relevant documents --- is critical to applications such as electronic discovery, systematic review, and the construction of test collections for information retrieval tasks. The effectiveness of current methods for high-recall information retrieval is limited by their reliance on human input, either to generate queries, or to assess the relevance of documents. Past research has shown that humans can assess the relevance of documents faster and with little loss in accuracy by judging shorter document surrogates, e.g.\ extractive summaries, in place of full documents. To test the hypothesis that short document surrogates can reduce assessment time and effort for high-recall retrieval, we conducted a 50-person, controlled, user study. We designed a high-recall retrieval system using continuous active learning (CAL) that could display either full documents or short document excerpts for relevance assessment. In addition, we tested the value of integrating a search engine with CAL. In the experiment, we asked participants to try to find as many relevant documents as possible within one hour. We observed that our study participants were able to find significantly more relevant documents when they used the system with document excerpts as opposed to full documents. We also found that allowing participants to compose and execute their own search queries did not improve their ability to find relevant documents and, by some measures, impaired performance. These results suggest that for high-recall systems to maximize performance, system designers should think carefully about the amount and nature of user interaction incorporated into the system. Haotian Zhang 0001, Mustafa Abualsaud, Nimesh Ghelani, Mark D. Smucker, Gordon V. Cormack, Maura R. Grossman |
CIKM | 6 |
| 2018 | The Quest for Total RecallabstractThe objective of high-recall information retrieval (HRIR) is to identify substantially all information relevant to an information need, where the consequences of missing or untimely results may have serious legal, policy, health, social, safety, defence, or financial implications. To find acceptance in practice, HRIR technologies must be more effective---and must be shown to be more effective---than current practice, according to the legal, statutory, regulatory, ethical, or professional standards governing the application domain. Such domains include, but are not limited to, electronic discovery in legal proceedings; distinguishing between public and non-public records in the curation of government archives; systematic review for meta-analysis in evidence-based medicine; separating irregularities and intentional misstatements from unintentional errors in accounting restatements; performing "due diligence" in connection with pending mergers, acquisitions, and financing transactions; and surveillance and compliance activities involving massive datasets. HRIR differs from ad hoc information retrieval where the objective is to identify the best, rather than all relevant information, and from classification or categorization where the objective is to separate relevant from non-relevant information based on previously labeled training examples. HRIR is further differentiated from established information retrieval applications by the need to quantify "substantially all relevant information"; an objective for which existing evaluation strategies and measures, such as precision and recall, are not particularly well suited. Gordon V. Cormack, Maura R. Grossman |
DocEng | 2 |
| 2018 | A System for Efficient High-Recall RetrievalabstractThe goal of high-recall information retrieval (HRIR) is to find all or nearly all relevant documents for a search topic. In this paper, we present the design of our system that affords efficient high-recall retrieval. HRIR systems commonly rely on iterative relevance feedback. Our system uses a state-of-the-art implementation of continuous active learning (CAL), and is designed to allow other feedback systems to be attached with little work. Our system allows users to judge documents as fast as possible with no perceptible interface lag. We also support the integration of a search engine for users who would like to interactively search and judge documents. In addition to detailing the design of our system, we report on user feedback collected as part of a 50 participants user study. While we have found that users find the most relevant documents when we restrict user interaction, a majority of participants prefer having flexibility in user interaction. Our work has implications on how to build effective assessment systems and what features of the system are believed to be useful by users. Mustafa Abualsaud, Nimesh Ghelani, Haotian Zhang 0001, Mark D. Smucker, Gordon V. Cormack, Maura R. Grossman |
SIGIR | 6 |
| 2018 | Beyond PoolingabstractDynamic Sampling is a novel, non-uniform, statistical sampling strategy in which documents are selected for relevance assessment based on the results of prior assessments. Unlike static and dynamic pooling methods that are commonly used to compile relevance assessments for the creation of information retrieval test collections, Dynamic Sampling yields a statistical sample from which substantially unbiased estimates of effectiveness measures may be derived. In contrast to static sampling strategies, which make no use of relevance assessments, Dynamic Sampling is able to select documents from a much larger universe, yielding superior test collections for a given budget of relevance assessments. These assertions are supported by simulation studies using secondary data from the TREC 2017 Common Core Track. Gordon V. Cormack, Maura R. Grossman |
SIGIR | 2 |
| 2017 | Navigating Imprecision in Relevance Assessments on the Road to Total Recall: Roger and MeabstractTechnology-assisted review ("TAR") systems seek to achieve "total recall"; that is, to approach, as nearly as possible, the ideal of 100% recall and 100% precision, while minimizing human review effort. The literature reports that TAR methods using relevance feedback can achieve considerably greater than the 65% recall and 65% precision reported by Voorhees as the "practical upper bound on retrieval performance... since that is the level at which humans agree with one another" (Variations in Relevance Judgments and the Measurement of Retrieval Effectiveness, 2000). This work argues that in order to build - as well as to, evaluate - TAR systems that approach 100% recall and 100% precision, it is necessary to model human assessment, not as absolute ground truth, but as an indirect indicator of the amorphous property known as "relevance." The choice of model impacts both the evaluation of system effectiveness, as well as the simulation of relevance feedback. Models are presented that better fit available data than the infallible ground-truth model. These models suggest ways to improve TAR-system effectiveness so that hybrid human-computer systems can improve on both the accuracy and efficiency of human review alone. This hypothesis is tested by simulating TAR using two datasets: the TREC 4 AdHoc collection, and a dataset consisting of 401,960 email messages that were manually reviewed and classified by a single individual, Roger, in his official capacity as Senior State Records Archivist. The results using the TREC 4 data show that TAR achieves higher recall and higher precision than the assessments by either of two independent NIST assessors, and blind adjudication of the email dataset, conducted by Roger, more than two years after his original review, shows that he could have achieved the same recall and better precision, while reviewing substantially fewer than 401,960 emails, had he employed TAR in place of exhaustive manual review. Gordon V. Cormack, Maura R. Grossman |
SIGIR | 2 |
| 2017 | Automatic and Semi-Automatic Document Selection for Technology-Assisted ReviewabstractAbstract In the TREC Total Recall Track (2015-2016), participating teams could employ either fully automatic or human-assisted ("semi-automatic") methods to select documents for relevance assessment by a simulated human reviewer. According to the TREC 2016 evaluation, the fully automatic baseline method achieved a recall-precision breakeven ("R-precision") score of 0.71, while the two semi-automatic efforts achieved scores of 0.67 and 0.51. In this work, we investigate the extent to which the observed effectiveness of the different methods may be confounded by chance, by inconsistent adherence to the Track guidelines, by selection bias in the evaluation method, or by discordant relevance assessments. We find no evidence that any of these factors could yield relative effectiveness scores inconsistent with the official TREC 2016 ranking. Maura R. Grossman, Gordon V. Cormack, Adam Roegiest |
SIGIR | 1 |
| 2016 | Scalability of Continuous Active Learning for Reliable High-Recall Text ClassificationabstractFor finite document collections, continuous active learning ('CAL') has been observed to achieve high recall with high probability, at a labeling cost asymptotically proportional to the number of relevant documents. As the size of the collection increases, the number of relevant documents typically increases as well, thereby limiting the applicability of CAL to low-prevalence high-stakes classes, such as evidence in legal proceedings, or security threats, where human effort proportional to the number of relevant documents is justified. We present a scalable version of CAL ('S-CAL') that requires O(log N) labeling effort and O(N log N) computational effort---where N is the number of unlabeled training examples---to construct a classifier whose effectiveness for a given labeling cost compares favorably with previously reported methods. At the same time, S-CAL offers calibrated estimates of class prevalence, recall, and precision, facilitating both threshold setting and determination of the adequacy of the classifier. Gordon V. Cormack, Maura R. Grossman |
CIKM | 2 |
| 2016 | Engineering Quality and Reliability in Technology-Assisted ReviewabstractThe objective of technology-assisted review ("TAR") is to find as much relevant information as possible with reasonable effort. Quality is a measure of the extent to which a TAR method achieves this objective, while reliability is a measure of how consistently it achieves an acceptable result. We are concerned with how to define, measure, and achieve high quality and high reliability in TAR. When quality is defined using the traditional goal-post method of specifying a minimum acceptable recall threshold, the quality and reliability of a TAR method are both, by definition, equal to the probability of achieving the threshold. Assuming this definition of quality and reliability, we show how to augment any TAR method to achieve guaranteed reliability, for a quantifiable level of additional review effort. We demonstrate this result by augmenting the TAR method supplied as the baseline model implementation for the TREC 2015 Total Recall Track, measuring reliability and effort for 555 topics from eight test collections. While our empirical results corroborate our claim of guaranteed reliability, we observe that the augmentation strategy may entail disproportionate effort, especially when the number of relevant documents is low. To address this limitation, we propose stopping criteria for the model implementation that may be applied with no additional review effort, while achieving empirical reliability that compares favorably to the provably reliable method. We further argue that optimizing reliability according to the traditional goal-post method is inconsistent with certain subjective aspects of quality, and that optimizing a Taguchi quality loss function may be more apt. Gordon V. Cormack, Maura R. Grossman |
SIGIR | 2 |
| 2015 | Multi-Faceted Recall of Continuous Active Learning for Technology-Assisted ReviewabstractContinuous active learning achieves high recall for technology-assisted review, not only for an overall information need, but also for various facets of that information need, whether explicit or implicit. Through simulations using Cormack and Grossman's TAR Evaluation Toolkit (SIGIR 2014), we show that continuous active learning, applied to a multi-faceted topic, efficiently achieves high recall for each facet of the topic. Our results assuage the concern that continuous active learning may achieve high overall recall at the expense of excluding identifiable categories of relevant information. Gordon V. Cormack, Maura R. Grossman |
SIGIR | 2 |
| 2015 | Impact of Surrogate Assessments on High-Recall RetrievalabstractWe are concerned with the effect of using a surrogate assessor to train a passive (i.e., batch) supervised-learning method to rank documents for subsequent review, where the effectiveness of the ranking will be evaluated using a different assessor deemed to be authoritative. Previous studies suggest that surrogate assessments may be a reasonable proxy for authoritative assessments for this task. Nonetheless, concern persists in some application domains---such as electronic discovery---that errors in surrogate training assessments will be amplified by the learning method, materially degrading performance. We demonstrate, through a re-analysis of data used in previous studies, that, with passive supervised-learning methods, using surrogate assessments for training can substantially impair classifier performance, relative to using the same deemed-authoritative assessor for both training and assessment. In particular, using a single surrogate to replace the authoritative assessor for training often yields a ranking that must be traversed much lower to achieve the same level of recall as the ranking that would have resulted had the authoritative assessor been used for training. We also show that steps can be taken to mitigate, and sometimes overcome, the impact of surrogate assessments for training: relevance assessments may be diversified through the use of multiple surrogates; and, a more liberal view of relevance can be adopted by having the surrogate label borderline documents as relevant. By taking these steps, rankings derived from surrogate assessments can match, and sometimes exceed, the performance of the ranking that would have been achieved, had the authority been used for training. Finally, we show that our results still hold when the role of surrogate and authority are interchanged, indicating that the results may simply reflect differing conceptions of relevance between surrogate and authority, as opposed to the authority having special skill or knowledge lacked by the surrogate. Adam Roegiest, Gordon V. Cormack, Charles L. A. Clarke, Maura R. Grossman |
SIGIR | 4 |
| 2014 | Evaluation of machine-learning protocols for technology-assisted review in electronic discoveryabstractAbstract Using a novel evaluation toolkit that simulates a human reviewer in the loop, we compare the effectiveness of three machine-learning protocols for technology-assisted review as used in document review for discovery in legal proceedings. Our comparison addresses a central question in the deployment of technology-assisted review: Should training documents be selected at random, or should they be selected using one or more non-random methods, such as keyword search or active learning? On eight review tasks -- four derived from the TREC 2009 Legal Track and four derived from actual legal matters -- recall was measured as a function of human review effort. The results show that entirely non-random training methods, in which the initial training documents are selected using a simple keyword search, and subsequent training documents are selected by active learning, require substantially and significantly less human review effort (P<0.01) to achieve any given level of recall, than passive learning, in which the machine-learning algorithm plays no role in the selection of training documents. Among passive-learning methods, significantly less human review effort (P<0.01) is required when keywords are used instead of random sampling to select the initial training documents. Among active-learning methods, continuous active learning with relevance feedback yields generally superior results to simple active learning with uncertainty sampling, while avoiding the vexing issue of "stabilization" -- determining when training is adequate, and therefore may stop. Gordon V. Cormack, Maura R. Grossman |
SIGIR | 2 |