VLDB 2026 Research / reviewers in the wild / expert
Renzhe Yu
dblp:221/4533
· DBLP profile ↗
22ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0002-2375-3537ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 21 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 7 since 2021Systems, architecture and hardware · 8 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 8 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing LLM-Based Data Annotation with Error DecompositionabstractLarge language models (LLMs) offer a scalable alternative to human coding for data annotation tasks, enabling the scale-up of research across data-intensive domains such as learning analytics. While LLMs are already achieving near-human accuracy on objective annotation tasks, their performance on subjective annotation tasks, such as those involving psychological constructs, is less consistent and more prone to errors. Standard evaluation practices typically collapse all annotation errors into a single alignment metric, but this simplified approach may obscure different kinds of errors that affect final analytical conclusions in different ways. Here, we propose a diagnostic evaluation paradigm that incorporates a human-in-the-loop step to separate task-inherent ambiguity from model-driven inaccuracies and assess annotation quality in terms of their potential downstream impacts. We refine this paradigm on ordinal annotation tasks, which are common in subjective annotation. The refined paradigm includes: (1) a diagnostic taxonomy that categorizes LLM annotation errors along two dimensions: source (model-specific vs. task-inherent) and type (boundary ambiguity vs. conceptual misidentification); (2) a lightweight human annotation test to estimate task-inherent ambiguity from LLM annotations; and (3) a computational method to decompose observed LLM annotation errors following our taxonomy. We validate this paradigm on four educational annotation tasks, demonstrating both its conceptual validity and practical utility. Theoretically, our work provides empirical evidence for why excessively high alignment is unrealistic in specific annotation tasks and why single alignment metrics inadequately reflect the quality of LLM annotations. In practice, our paradigm can be a low-cost diagnostic tool that assesses the suitability of a given task for LLM annotation and provides actionable insights for further technical optimization. Vedant Khatri, Yijun Dai, Xiner Liu, Siyan Li, Xuanming Zhang, Renzhe Yu |
LAK | 7 |
| 2026 | Calibrating Social Bias in LLM and Human Responses in Postsecondary Online Discussion ForumsabstractRecent research has increasingly documented that large language models (LLMs) can exhibit social bias in consequential social contexts such as education. Yet this literature often evaluates LLM bias in isolation without calibrating it against human bias, even though human bias has been extensively documented and constitutes one important source of LLM bias through pretraining on human-generated text. Such calibration is important for understanding when and how LLMs may reproduce human patterns of bias, with strong practical implications for deploying general-purpose LLMs and customizing models on task-specific datasets. The current study examines these issues in postsecondary education contexts, comparing racial and gender bias in human- and LLM-generated responses to postsecondary online discussion posts. Using a small pilot dataset sampled at a public four-year institution in the United States, we find that humans exhibit small and insignificant levels of social bias and that LLM bias does not differ significantly from this human baseline. These analyses motivate a larger-scale investigation that may inform debates about the appropriateness of AI in online learning environments and the risks of fine-tuning LLMs on local educational data. Daniel March, Chengyuan Yao, Renzhe Yu |
L@S | 4 |
| 2026 | Asynchronous Discussion Forums Show More Analytical but Less Diverse Student Engagement in the Age of AI
Yijun Dai, Chenxi Shi, Siyan Li, Chengyuan Yao, Renzhe Yu |
L@S | 6 |
| 2025 | Evaluating an AI Tutor for Bias Across Different Foundation Models
Aditya Vinodh, Emma Harvey, Husni Almoubayyed, Renzhe Yu, Christopher Brooks 0001, Allison Koenecke, René F. Kizilcec |
AIED (6) | 4 |
| 2025 | From Course to Skill: Evaluating Large Language Model Performance in Curricular Analytics
Xinjin Li, Yingqi Huan, Veronica Minaya, Renzhe Yu |
AIED (6) | 5 |
| 2025 | Understanding Predictive Models of Student Success with a Multiverse Analysis
Yunxuan Tang, Emma Harvey, Chengyuan Yao, Renzhe Yu, René F. Kizilcec, Christopher Brooks 0001 |
EDM | 4 |
| 2025 | Predicting Adolescent Suicidal Risk from Multi-task-based Speech: An Ensemble Learning Approach
Renzhe Yu, Yanshen Tan, Yiyi Li, Quan Qian |
INTERSPEECH | 2 |
| 2025 | Towards Fair and Privacy-Aware Transfer Learning for Educational Predictive Modeling: A Case Study on Retention Prediction in Community Colleges
Chengyuan Yao, Carmen Cortez, Renzhe Yu |
LAK | 3 |
| 2025 | Investigating Systematic Variation in Academic Procrastination Behavior by Course, Assignment, and Student CharacteristicsabstractProcrastination has been linked to lower academic performance and sociodemographic achievement gaps in a variety of educational contexts, posing challenges to student success and educational equity. While prior research acknowledges that learning environments play a crucial role in shaping student procrastination alongside personal traits, there is a lack of solid empirical evidence on the connection between specific variations in learning environments and academic procrastination. This study provides a large-scale evaluation of the relationship between course and assignment characteristics and student procrastination behavior using a sample of 33,514 students across 3,169 courses at a US university. Using fixed effects linear regression models, we find that students tend to procrastinate less in courses with larger enrollment, non-introductory content, and well-structured deadlines. Procrastination is also lower for assignments with spaced-out deadlines, weekend deadlines, and a quiz or discussion post format. However, these patterns do not apply equally across all student groups. Male, ethnic minority, and first-generation college students exhibit higher levels of procrastination than their peers, especially for courses and assignments with specific characteristics. We suggest two instructional design strategies to help manage procrastination across student populations: (1) allowing more time before the first assignment deadline, and (2) ensuring adequate spacing between deadlines. This study provides large-scale evidence of the complex relationship between learning environment design, student characteristics, and procrastination. Yan Tao, Nathan Maidi, Renzhe Yu, René F. Kizilcec |
L@S | 3 |
| 2024 | Temporal and Between-Group Variability in College Dropout PredictionabstractLarge-scale administrative data is a common input in early warning systems for college dropout in higher education. Still, the terminology and methodology vary significantly across existing studies, and the implications of different modeling decisions are not fully understood. This study provides a systematic evaluation of contributing factors and predictive performance of machine learning models over time and across different student groups. Drawing on twelve years of administrative data at a large public university in the US, we find that dropout prediction at the end of the second year has a 20% higher AUC than at the time of enrollment in a Random Forest model. Also, most predictive factors at the time of enrollment, including demographics and high school performance, are quickly superseded in predictive importance by college performance and in later stages by enrollment behavior. Regarding variability across student groups, college GPA has more predictive value for students from traditionally disadvantaged backgrounds than their peers. These results can help researchers and administrators understand the comparative value of different data sources when building early warning systems and optimizing decisions under specific policy goals. Dominik Glandorf, Hye Rin Lee, Gabe Avakian Orona, Marina Pumptow, Renzhe Yu, Christian Fischer 0007 |
LAK | 5 |
| 2024 | Contexts Matter but How? Course-Level Correlates of Performance and Fairness Shift in Predictive Model TransferabstractLearning analytics research has highlighted that contexts matter for predictive models, but little research has explicated how contexts matter for models’ utility. Such insights are critical for real-world applications where predictive models are frequently deployed across instructional and institutional contexts. Building upon administrative records and behavioral traces from 37,089 students across 1,493 courses, we provide a comprehensive evaluation of performance and fairness shifts of predictive models when transferred across different course contexts. We specifically quantify how differences in various contextual factors moderate model portability. Our findings indicate an average decline in model performance and inconsistent directions in fairness shifts, without a direct trade-off, when models are transferred across different courses within the same institution. Among the course-to-course contextual differences we examined, differences in admin features account for the largest portion of both performance and fairness loss. Differences in student composition can simultaneously amplify drops in performance and fairness while differences in learning design have a greater impact on performance degradation. Given these complexities, our results highlight the importance of considering multiple dimensions of course contexts and evaluating fairness shifts in addition to performance loss when conducting transfer learning of predictive models in education. Joseph Olson, Nicole Pochinki, Zhijian Zheng, Renzhe Yu |
LAK | 5 |
| 2024 | Technology-Based Instructional Strategies Show Promise in Improving Self-Regulated Learning Skills at Broad-Access Postsecondary InstitutionsabstractSelf-regulated learning (SRL) is critical for student success in online postsecondary education. Many technology-based interventions have been studied to improve SRL skills, but few were situated in broad-access institutions that disproportionately serve systemically marginalized student populations in STEM fields. This study presents preliminary findings from a rapid-cycle evaluation that tests two technology-supported instructional strategies (videos and prompts) designed to improve SRL in online learning. Using fine-grained clickstream data from 141 students across ten sections of five courses taught at a minority-serving community college, we generate measures of SRL behavior and correlate them with students' exposure to tested strategies. Our results indicate modestly positive relationships between both videos and prompts and SRL behavior. In addition, prompts are more strongly correlated with SRL behavior for first-generation and female students than for their peers. These initial findings reveal the promise and complexity of implementing effective and equitable technology-supported interventions to develop SRL skills and mindsets among diverse student populations in online STEM education. Renzhe Yu, Hui Yang 0025, Xiaoying Lin, Chengyuan Yao, Paul Burkander, Krystal Thomas, Jessica Mislevy |
L@S | 1 |
| 2023 | Semantic Topic Chains for Modeling Temporality of Themes in Online Student Discussion Forums
Harshita Chopra, Yiwen Lin, Mohammad Amin Samadi, Jacqueline G. Cavazos, Renzhe Yu, Spencer Jaquay, Nia Nixon |
EDM | 5 |
| 2022 | FATED 2022: Fairness, Accountability, and Transparency in Educational Data
Collin F. Lynch, Mirko Marras, Mykola Pechenizkiy, Anna N. Rafferty, Steven Ritter 0001, Vinitra Swamy, Renzhe Yu |
EDM | 7 |
| 2022 | Large-Scale Student Data Reveal Sociodemographic Gaps in Procrastination BehaviorabstractUniversity students have to manage complex and demanding schedules to keep up with coursework across multiple classes while navigating formative personal, cultural, and financial events. Procrastination, the act of deferring study effort until the task deadline, is therefore a prevalent phenomenon, but whether it is more common among historically disadvantaged students is unknown. If systematic differences in procrastination behavior exist across sociodemographic groups, they may also contribute to achievement gaps, considering that procrastination is largely negatively associated with academic performance in prior research. We therefore investigate these questions in the context of assignment submission using campus-wide learning management system (LMS) data from a large U.S. research university. We analyze 2,631,893 submission records by 25,659 students across 2,153 courses and propose a context-agnostic procrastination score for each student in each course based on their assignment submission times relative to classmates. Based on this procrastination score, we find significantly higher levels of procrastination behavior among males, racial minorities, and first-generation college students than their peers. However, these differences only explain performance gaps to a very limited extent and the negative association between procrastination behavior and performance remains relatively stable across student groups. This large-scale behavioral study advances the understanding of academic procrastination through an equity lens and informs the development of scalable interventions to mitigate the negative effects of procrastination. Sunil Sabnis, Renzhe Yu, René F. Kizilcec |
L@S | 2 |
| 2021 | Should College Dropout Prediction Models Include Protected Attributes?abstractEarly identification of college dropouts can provide tremendous value for improving student success and institutional effectiveness, and predictive analytics are increasingly used for this purpose. However, ethical concerns have emerged about whether including protected attributes in these prediction models discriminates against underrepresented student groups and exacerbates existing inequities. We examine this issue in the context of a large U.S. research university with both residential and fully online degree-seeking students. Based on comprehensive institutional records for the entire student population across multiple years (N = 93,457), we build machine learning models to predict student dropout after one academic year of study and compare the overall performance and fairness of model predictions with or without four protected attributes (gender, URM, first-generation student, and high financial need). We find that including protected attributes does not impact the overall prediction performance and it only marginally improves the algorithmic fairness of predictions. These findings suggest that including protected attributes is preferable. We offer guidance on how to evaluate the impact of including protected attributes in a local context, where institutional stakeholders seek to leverage predictive analytics to support student success. Renzhe Yu, René F. Kizilcec |
L@S | 1 |
| 2020 | LIWCs the Same, Not the Same: Gendered Linguistic Signals of Performance and Experience in Online STEM Courses
Yiwen Lin, Renzhe Yu, Nia Nixon |
AIED (1) | 2 |
| 2020 | Towards Accurate and Fair Prediction of College Success: Evaluating Different Sources of Student Data
Renzhe Yu, Qiujie Li, Christian Fischer 0007, Shayan Doroudi, Di Xu 0005 |
EDM | 1 |
| 2020 | Interpretable Models Do Not Compromise Accuracy or Fairness in Predicting College SuccessabstractThe presence of "big data" in higher education has led to the increasing popularity of predictive analytics for guiding various stakeholders on appropriate actions to support student success. In developing such applications, model selection is a central issue. As such, this study presents a comprehensive examination of five commonly used machine learning models in student success prediction. Using administrative and learning management system (LMS) data for nearly 2,000 college students at a public university, we employ the models to predict short-term and long-term academic success. Beyond the tradeoff between model interpretability and accuracy, we also focus on the fairness of these models with regard to different student populations. Our findings suggest that more interpretable models such as logistic regression do not necessarily compromise predictive accuracy. Also, they lead to no more, if not less, prediction bias against disadvantaged student groups than complicated models. Moreover, prediction biases against certain groups persist even in the fairest model. These results thus recommend using simpler algorithms in conjunction with human evaluation in instructional and institutional applications of student success prediction when valid student features are in place. Catherine Kung, Renzhe Yu |
L@S | 2 |
| 2019 | Utilizing Learning Analytics to Map Students' Self-Reported Study Strategies to Click Behaviors in STEM CoursesabstractInformed by cognitive theories of learning, this work examined how students' self-reported study patterns (spacing vs. cramming) corresponded to their engagement with the Learning Management System (LMS) across two years in a large biology course. We specifically focused on how students accessed non-mandatory resources (lecture videos, lecture slides) and considered whether this pattern differed by underrepresented minority (URM) status. Overall, students who self-reported utilizing spacing strategies throughout the course had higher grades than students who reported cramming throughout the course. When examining LMS engagement, only a small percentage of students accessed the lecture videos and lecture slides. Applying a negative binomial regression model to daily counts of click activities, we also found that students who utilized spacing strategies accessed LMS resources more often but not earlier before major deadlines. Moreover, this finding was not different for underrepresented students. Our results provide some initial evidence showing how spacing behaviors correspond to accessing learning resources. However, given the lack of general engagement with LMS resources, our results underscore the value of encouraging students to utilize these resources when studying course material. Fernando Rodriguez, Renzhe Yu, Mariela Janet Rivas, Mark Warschauer, Brian K. Sato |
LAK | 2 |
| 2018 | Understanding Student Procrastination via Mixture Models
Renzhe Yu, Fernando Rodriguez, Rachel B. Baker, Padhraic Smyth, Mark Warschauer |
EDM | 2 |
| 2018 | Representing and predicting student navigational pathways in online college coursesabstractRepresentation and prediction of student navigational pathways, typically based on neural network (NN) methods, have seen their potential of improving instruction and learning under insufficient human knowledge about learner behavior. However, they are prominently studied in MOOCs and less probed within more institutionalized higher education scenarios. This work extends such research to the context of online college courses. Comparing student navigational sequences through course pages to documents in natural language processing, we apply a skip-gram model to learn vector embedding of course pages, and visualize the learnt vectors to understand the extent to which students' learning pathways align with pre-designed course structure. We find that students who get different letter grades in the end exhibit different levels of adherence to designed sequence. Next, we fit the embedded sequences into a long short-term memory architecture and test its ability to predict next page that a student visits given her prior sequence. The highest accuracy reaches 50.8% and largely outperforms the frequency-based baseline of 41.3%. These results show that neural network methods have the potential to help instructors understand students' learning behaviors and facilitate automated instructional support. Renzhe Yu, Daokun Jiang, Mark Warschauer |
L@S | 1 |