VLDB 2026 Research / reviewers in the wild / expert
John Lalor
dblp:159/0111 · also John P. Lalor, John Patrick Lalor
· DBLP profile ↗
15ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0003-0848-4786ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | No Simple Answer to Data Complexity: An Examination of Instance-Level Complexity Metrics for Classification TasksabstractRyan A. Cook, John P. Lalor, Ahmed Abbasi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ryan A. Cook, John Lalor, Ahmed Abbasi |
NAACL (Long Papers) | 2 |
| 2025 | Hierarchical Deep Document ModelabstractTopic modeling is a commonly used text analysis tool for discovering latent topics in a text corpus. However, while topics in a text corpus often exhibit a hierarchical structure (e.g., cellphone is a sub-topic of electronics), most topic modeling methods assume a flat topic structure that ignores the hierarchical dependency among topics, or utilize a predefined topic hierarchy. In this work, we present a novel Hierarchical Deep Document Model (HDDM) to learn topic hierarchies using a variational autoencoder framework. We propose a novel objective function, sum of log likelihood, instead of the widely used evidence lower bound, to facilitate the learning of hierarchical latent topic structure. The proposed objective function can directly model and optimize the hierarchical topic-word distributions at all topic levels. We conduct experiments on four real-world text datasets to evaluate the topic modeling capability of the proposed HDDM method compared to state-of-the-art hierarchical topic modeling benchmarks. Experimental results show that HDDM achieves considerable improvement over benchmarks and is capable of learning meaningful topics and topic hierarchies. To further demonstrate the practical utility of HDDM, we apply it to a real-world medical notes dataset for clinical prediction. Experimental results show that HDDM can better summarize topics in medical notes, resulting in more accurate clinical predictions. Yi Yang 0042, John Lalor, Ahmed Abbasi, Daniel Dajun Zeng |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Should Fairness be a Metric or a Model? A Model-based Framework for Assessing Bias in Machine Learning PipelinesabstractFairness measurement is crucial for assessing algorithmic bias in various types of machine learning (ML) models, including ones used for search relevance, recommendation, personalization, talent analytics, and natural language processing. However, the fairness measurement paradigm is currently dominated by fairness metrics that examine disparities in allocation and/or prediction error as univariate key performance indicators (KPIs) for a protected attribute or group. Although important and effective in assessing ML bias in certain contexts such as recidivism, existing metrics don’t work well in many real-world applications of ML characterized by imperfect models applied to an array of instances encompassing a multivariate mixture of protected attributes, that are part of a broader process pipeline. Consequently, the upstream representational harm quantified by existing metrics based on how the model represents protected groups doesn’t necessarily relate to allocational harm in the application of such models in downstream policy/decision contexts. We propose FAIR-Frame, a model-based framework for parsimoniously modeling fairness across multiple protected attributes in regard to the representational and allocational harm associated with the upstream design/development and downstream usage of ML models. We evaluate the efficacy of our proposed framework on two testbeds pertaining to text classification using pretrained language models. The upstream testbeds encompass over fifty thousand documents associated with twenty-eight thousand users, seven protected attributes and five different classification tasks. The downstream testbeds span three policy outcomes and over 5.41 million total observations. Results in comparison with several existing metrics show that the upstream representational harm measures produced by FAIR-Frame and other metrics are significantly different from one another, and that FAIR-Frame’s representational fairness measures have the highest percentage alignment and lowest error with allocational harm observed in downstream applications. Our findings have important implications for various ML contexts, including information retrieval, user modeling, digital platforms, and text classification, where responsible and trustworthy AI is becoming an imperative. John Lalor, Ahmed Abbasi, Kezia Oketch, Yi Yang 0042, Nicole Forsgren |
ACM Trans. Inf. Syst. | 1 |
| 2023 | <tt>py-irt</tt>: A Scalable Item Response Theory Library for Pythonabstractpy-irt is a Python library for fitting Bayesian item response theory (IRT) models. At present, there is no Python package for fitting large-scale IRT models. py-irt estimates latent traits of subjects and items, making it appropriate for use in IRT tasks as well as in ideal point models. py-irt is built on top of the Pyro and PyTorch frameworks and uses GPU-accelerated training to scale to large data sets. It is the first Python package for large-scale IRT model fitting. py-irt is easy to use for practitioners and also allows for researchers to build and fit custom IRT models. py-irt is available as open-source software and can be installed from GitHub or the Python Package Index. History: Accepted by Ted Ralphs, Area Editor for software tools. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplementary Information [ https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2022.1250 ] or is available from the IJOC GitHub software repository ( https://github.com/INFORMSJoC ) at [ http://dx.doi.org/10.5281/zenodo.6818509 ]. John Lalor, Pedro Rodríguez 0001 |
INFORMS J. Comput. | 1 |
| 2022 | Benchmarking Intersectional Biases in NLPabstractJohn Lalor, Yi Yang, Kendall Smith, Nicole Forsgren, Ahmed Abbasi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. John Lalor, Yi Yang 0042, Kendall Smith, Nicole Forsgren, Ahmed Abbasi |
NAACL-HLT | 1 |
| 2021 | Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?abstractPedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, Jordan Boyd-Graber. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Pedro Rodríguez 0001, Joe Barrow, Alexander Miserlis Hoyle, John Lalor, Robin Jia, Jordan L. Boyd-Graber |
ACL/IJCNLP (1) | 4 |
| 2021 | Constructing a Psychometric Testbed for Fair Natural Language ProcessingabstractPsychometric measures of ability, attitudes, perceptions, and beliefs are crucial for understanding user behavior in various contexts including health, security, e-commerce, and finance.Traditionally, psychometric dimensions have been measured and collected using survey-based methods.Inferring such constructs from user-generated text could allow timely, unobtrusive collection and analysis.In this work we construct a corpus for psychometric natural language processing (NLP) related to important dimensions such as trust, anxiety, numeracy, and literacy, in the health domain.We discuss our multi-step process to align user text with their survey-based response items and provide an overview of the resulting testbed, which encompasses surveybased psychometric measures and accompanying user-generated text from 8,502 respondents.Our testbed also encompasses selfreported demographic information, including race, sex, age, income, and education, allowing for measuring bias and benchmarking fairness of text classification methods.We report preliminary results on use of the text to predict/categorize users' survey response labels and on the fairness of these models.We also discuss the important implications of our work and resulting testbed for future NLP research on psychometrics and fairness. Ahmed Abbasi, David G. Dobolyi, John Lalor, Richard G. Netemeyer, Kendall Smith, Yi Yang 0042 |
EMNLP (1) | 3 |
| 2019 | Efficient Semi-Supervised Learning for Natural Language Understanding by Optimizing DiversityabstractExpanding new functionalities efficiently is an ongoing challenge for single-turn task-oriented dialogue systems. In this work, we explore functionality-specific semi-supervised learning via self-training. We consider methods that augment training data automatically from unlabeled data sets in a functionality-targeted manner. In addition, we examine multiple techniques for efficient selection of augmented utterances to reduce training time and increase diversity. First, we consider paraphrase detection methods that attempt to find utterance variants of labeled training data with good coverage. Second, we explore sub-modular optimization based on n-grams features for utterance selection. Experiments show that functionality-specific self-training is very effective for improving system performance. In addition, methods optimizing diversity can reduce training data in many cases to 50% with little impact on performance. Eunah Cho, He Xie, John Lalor, William M. Campbell |
ASRU | 3 |
| 2019 | Learning Latent Parameters without Human Response Patterns: Item Response Theory with Artificial CrowdsabstractIncorporating Item Response Theory (IRT) into NLP tasks can provide valuable information about model performance and behavior. Traditionally, IRT models are learned using human response pattern (RP) data, presenting a significant bottleneck for large data sets like those required for training deep neural networks (DNNs). In this work we propose learning IRT models using RPs generated from artificial crowds of DNN models. We demonstrate the effectiveness of learning IRT models using DNN-generated data through quantitative and qualitative analyses for two NLP tasks. Parameters learned from human and machine RPs for natural language inference and sentiment analysis exhibit medium to large positive correlations. We demonstrate a use-case for latent difficulty item parameters, namely training set filtering, and show that using difficulty to sample training data outperforms baseline methods. Finally, we highlight cases where human expectation about item difficulty does not match difficulty as estimated from the machine RPs. John Lalor, Hao Wu 0055, Hong Yu 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Detecting Hypoglycemia Incidents from Patients' Secure Messages
Jinying Chen, John Lalor, Hong Yu 0001 |
AMIA | 2 |
| 2018 | Understanding Deep Learning Performance through an Examination of Test Set Difficulty: A Psychometric Case StudyabstractInterpreting the performance of deep learning models beyond test set accuracy is challenging. Characteristics of individual data points are often not considered during evaluation, and each data point is treated equally. We examine the impact of a test set question's difficulty to determine if there is a relationship between difficulty and performance. We model difficulty using well-studied psychometric methods on human response patterns. Experiments on Natural Language Inference (NLI) and Sentiment Analysis (SA) show that the likelihood of answering a question correctly is impacted by the question's difficulty. As DNNs are trained with more data, easy examples are learned more quickly than hard examples. John Lalor, Hao Wu 0055, Tsendsuren Munkhdalai, Hong Yu 0001 |
EMNLP | 1 |
| 2017 | Generating a Test of Electronic Health Record Narrative Comprehension with Item Response Theory
John Lalor, Hao Wu 0055, Kathleen M. Mazor, Hong Yu 0001 |
AMIA | 1 |
| 2016 | Building an Evaluation Scale using Item Response TheoryabstractEvaluation of NLP methods requires testing against a previously vetted gold-standard test set and reporting standard metrics (accuracy/precision/recall/F1). The current assumption is that all items in a given test set are equal with regards to difficulty and discriminating power. We propose Item Response Theory (IRT) from psychometrics as an alternative means for gold-standard test-set generation and NLP system evaluation. IRT is able to describe characteristics of individual items - their difficulty and discriminating power - and can account for these characteristics in its estimation of human intelligence or ability for an NLP task. In this paper, we demonstrate IRT by generating a gold-standard test set for Recognizing Textual Entailment. By collecting a large number of human responses and fitting our IRT model, we show that our IRT model compares NLP systems with the performance in a human population and is able to provide more insight into system performance than standard evaluation metrics. We show that a high accuracy score does not always imply a high IRT score, which depends on the item characteristics and the response pattern. John Lalor, Hao Wu 0055, Hong Yu 0001 |
EMNLP | 1 |
| 2015 | A Computer Science Linked-courses Learning CommunityabstractPrevious work has shown that factors such as student engagement and involvement can impact progress for computer science majors. One promising approach for improving student engagement is learning communities, which have a long history in academia but are relatively uncommon in computing. In this article we describe a linked-courses learning community for women and men of color majoring in development-focused computing degrees. We provide logistical information about the first offering of the learning community and assess the effectiveness of the community via a student survey. Our results show that students in the learning community are more likely to report that they have support for success in computer science courses and that they are a part of a community of programmers. Amber Settle, John Lalor, Theresa A. Steinbach |
ITiCSE | 2 |
| 2015 | Reconsidering the Impact of CS1 on Novice AttitudesabstractStudent success in an introductory programing course is crucial, both because it influences retention and because student attitudes and habits in a first course can have a lasting impact on student success in computer science as a field. In this paper we present results about student attitudes and habits before and after a CS1 class. Statistically significant attitude differences were found in three areas: students were less likely to report they were good at programming, more likely to agree they are challenged by programming problems they can't understand immediately, and are less likely to report that computer science allows them to be creative. Statistically significant differences in female and first-quarter responses were also found. Amber Settle, John Lalor, Theresa A. Steinbach |
SIGCSE | 2 |