VLDB 2026 Research / reviewers in the wild / expert
Wesley Morris
dblp:340/6520 · also Wesley G. Morris
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-6316-6479ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Guidelines for Designing AI Technologies to Support Adult LearningabstractAI-powered educational technologies have demonstrated measurable benefits for learners, but their design and evaluation have largely centered on K-12 contexts. As a result, many AI-supported learning systems remain poorly aligned with the needs, constraints, and goals of adult learners. To better understand how AI systems function in adult education, this paper examines the deployment of several AI learning technologies developed within a multidisciplinary, national research institute in the United States focused on adult learning and online education. Drawing on longitudinal deployment data, we conducted a reflexive thematic analysis to identify recurring challenges and design considerations across systems. These insights were synthesized into a set of 19 design guidelines intended to inform future AI-supported adult learning technologies. We demonstrate the utility of these guidelines through a heuristic evaluation of the deployed systems. Lastly, we present a guideline exploration tool that aids in the ideation of technologies by connecting the guidelines to stakeholder statements surfaced in the analysis process. Jennifer M. Reddig, Glen R. Smith Jr., Sanaz Ahmadzadeh Siyahrood, Wesley Morris, Yoojin Bae, Kaitlyn Crutcher, John Kos, Rahul K. Dass, Momin Naushad Siddiqui, Daniel Weitekamp III, Ploy Thajchayapong, Sandeep Kakar, Alex Endert, Scott Crossley, Min Kyu Kim, Chris Dede, Ashok K. Goel 0001, Christopher J. MacLellan |
DIS | 4 |
| 2026 | The Effects of AI Feedback on College Students' Reading and Writing Performance in an Intelligent Text FrameworkabstractThis study investigates how AI feedback influences college students’ reading and writing performance within an intelligent text aligned with the Interactive Constructive Active and Passive (ICAP) framework. Using a within-subjects design, students completed two read-to-write tasks — constructed response items (CRIs) and summaries — under two conditions: (1) Strategic Thinking And Interactive Reading Support (STAIRS), which provided AI feedback to support interactive engagement, and (2) Random Reread, a control condition without AI feedback to support constructive engagement. STAIRS offered automated scoring of read-to-write tasks, targeted rereading, and a chatbot that engaged students in interactive dialogues to support revision. Results showed that STAIRS significantly improved students’ summary revisions and subsequent summary performance, suggesting that AI feedback functions not only as guidance for improving current summaries but also benefits future performance. However, students in the STAIRS condition did not show higher CRI or summative quiz scores. CRI scores may have been limited by a ceiling effect, while the lack of quiz benefits may reflect a tradeoff in the study design, where the classroom tasks did not fully match the system’s activities. These mixed outcomes highlight the need to further explore how AI feedback may support different aspects of reading with an intelligent text. Alyssa Friend Wise, Wesley Morris, Langdon Holmes, Scott R. Hinze, Scott A. Crossley |
LAK | 3 |
| 2025 | Uncovering Differential Sensitivity Toward Linguistic Features of Cohesion in Large Language Models
Wesley Morris, Langdon Holmes, Joon Suh Choi, Scott A. Crossley |
AIED (6) | 1 |
| 2025 | Exploratory Assessment of Learning in an Intelligent Text Framework: iTELL RCTabstractThis study explores users' learning gains and experiences from reading within three versions of the same text. The versions included 1) a traditional digital text, 2) a productive text that required participants to produce knowledge about what they read, but without any feedback on the knowledge produced, and 3) an interactive text that required participants to produce knowledge and provided AI feedback and interaction. Learning gains were assessed in a randomized control trial where crowd-sourced users were assigned to one of the three versions and were examined based on the quality of constructed responses and summaries provided as well as through differences from a pre-test and a post-test. User experiences were investigated using survey results. Results indicated that users were generally satisfied with interacting with all versions the text, except the summary portion of the interactive text. Conversely, results indicated that users of the interactive text consistently wrote better summaries than in the productive condition, and they revised summaries to a greater degree and to a greater success. Lastly, results showed that knowledge gains occurred in all reading conditions and that readers in the interactive condition who spent more time reading the text showed stronger test scores overall. Scott Crossley, Wesley Morris, Joon Suh Choi, Langdon Holmes |
L@S | 2 |
| 2024 | Plagiarism Detection Using Keystroke Logs
Scott A. Crossley, Joon Suh Choi, Langdon Holmes, Wesley Morris |
EDM | 5 |
| 2024 | iScore: Visual Analytics for Interpreting How Language Models Automatically Score SummariesabstractThe recent explosion in popularity of large language models (LLMs) has inspired learning engineers to incorporate them into adaptive educational tools that automatically score summary writing. Understanding and evaluating LLMs is vital before deploying them in critical learning environments, yet their unprecedented size and expanding number of parameters inhibits transparency and impedes trust when they underperform. Through a collaborative user-centered design process with several learning engineers building and deploying summary scoring LLMs, we characterized fundamental design challenges and goals around interpreting their models, including aggregating large text inputs, tracking score provenance, and scaling LLM interpretability methods. To address their concerns, we developed iScore, an interactive visual analytics tool for learning engineers to upload, score, and compare multiple summaries simultaneously. Tightly integrated views allow users to iteratively revise the language in summaries, track changes in the resulting LLM scores, and visualize model weights at multiple levels of abstraction. To validate our approach, we deployed iScore with three learning engineers over the course of a month. We present a case study where interacting with iScore led a learning engineer to improve their LLM’s score accuracy by three percentage points. Finally, we conducted qualitative interviews with the learning engineers that revealed how iScore enabled them to understand, evaluate, and build trust in their LLMs during deployment. Adam Coscia, Langdon Holmes, Wesley Morris, Joon Suh Choi, Scott A. Crossley, Alex Endert |
IUI | 3 |
| 2023 | Using Transformer Language Models to Validate Peer-Assigned Essay Scores in Massive Open Online Courses (MOOCs)abstractMassive Open Online Courses (MOOCs) such as those offered by Coursera are popular ways for adults to gain important skills, advance their careers, and pursue their interests. Within these courses, students are often required to compose, submit, and peer review written essays, providing a valuable pedagogical experience for the student and a wealth of natural language data for the educational researcher. However, the scores provided by peers do not always reflect the actual quality of the text, generating questions about the reliability and validity of the scores. This study evaluates methods to increase the reliability of MOOC peer-review ratings through a series of validation tests on peer-reviewed essays. Reliability of reviewers was based on correlations between text length and essay quality. Raters were pruned based on score variance and the lexical diversity observed in their comments to create sub-sets of raters. Each subset was then used as training data to finetune distilBERT large language models to automatically score essay quality as a measure of validation. The accuracy of each language model for each subset was evaluated. We find that training language models on data subsets produced by more reliable raters based on a combination of score variance and lexical diversity produce more accurate essay scoring models. The approach developed in this study should allow for enhanced reliability of peer-reviewed scoring in MOOCS affording greater credibility within the systems. Wesley Morris, Scott A. Crossley, Langdon Holmes, Anne Trumbore |
LAK | 1 |