VLDB 2026 Research / reviewers in the wild / expert
Qinyi Liu
dblp:81/10594
· DBLP profile ↗
8ranked-venue papers
6as first author
8since 2021 · last 2026
0009-0003-4973-0901ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 6 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 7 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Brief but Impactful: How Human Tutoring Interactions Shape Engagement in Online LearningabstractLearning analytics can guide human tutors to efficiently address motivational barriers to learning that AI systems struggle to support. Students become more engaged when they receive human attention. However, what occurs during short interventions, and when are they most effective? We align student–tutor dialogue transcripts with MATHia tutoring system log data to study brief human-tutor interactions on Zoom drawn from 2,075 hours of 191 middle school students’ classroom math practice. Mixed-effect models reveal that engagement, measured as successful solution steps per minute, is higher during a human-tutor visit and remains elevated afterward. Visit length exhibits diminishing returns: engagement rises during and shortly after visits, irrespective of visit length. Timing also matters: later visits yield larger immediate lifts than earlier ones, though an early visit remains important to counteract engagement decline. We create analytics that identify which tutor-student dialogues raise engagement the most. Qualitative analysis reveals that interactions with concrete, stepwise scaffolding with explicit work organization elevate engagement most strongly. We discuss implications for resource-constrained tutoring, prioritizing several brief, well-timed check-ins by a human tutor while ensuring at least one early contact. Our analytics can guide the prioritization of students for support and surface effective tutor moves in real-time. Conrad Borchers, Ashish Gurung, Qinyi Liu, Danielle R. Thomas, Mohammad Khalil, Kenneth R. Koedinger |
LAK | 3 |
| 2026 | Measuring the Impact of Student Gaming Behaviors on Learner ModelingabstractThe expansion of large-scale online education platforms has yielded vast amounts of student interaction data for knowledge tracing (KT). KT models estimate students’ concept mastery from interaction data, but the models’ performance is sensitive to input data quality. Gaming behaviors, such as excessive hint use, may misrepresent students’ knowledge and undermine model reliability. However, systematic investigations of how different types of gaming behaviors affect KT remain scarce, and existing studies rely on costly manual analysis that does not capture behavioral diversity. In this study, we conceptualize gaming behaviors as a form of data poisoning, defined as the deliberate submission of incorrect or misleading interaction data to corrupt a model’s learning process. We design Data Poisoning Attacks (DPA) to simulate diverse gaming patterns and systematically evaluate their impact on KT model performance. Moreover, drawing on advances in DPA detection, we explore unsupervised approaches to enhance the generalizability of gaming behavior detection. We find that KT models performance tend to decrease especially for random guess behaviors. Our findings provide insights into the vulnerabilities of KT models and highlight the potential of adversarial methods for improving the robustness of learning analytics systems. Qinyi Liu, Lin Li 0039, Valdemar Svábenský, Conrad Borchers, Mohammad Khalil |
LAK | 1 |
| 2026 | Causal Pre-training Under the Fairness Lens: An Empirical Study of TabPFNabstractFoundation models for tabular data, such as the Tabular Prior-data Fitted Network (TabPFN), are pre-trained on a massive number of synthetic datasets generated by structural causal models (SCM). They leverage in-context learning to offer high predictive accuracy in real-world tasks. However, the fairness properties of these foundational models, which incorporate ideas from causal reasoning during pre-training, remain underexplored. In this work, we conduct a comprehensive empirical evaluation of TabPFN and its fine-tuned variants, assessing predictive performance, fairness, and robustness across varying dataset sizes and distributional shifts. Our results reveal that while TabPFN achieves stronger predictive accuracy compared to baselines and exhibits robustness to spurious correlations, improvements in fairness are moderate and inconsistent, particularly under missing-not-at-random (MNAR) covariate shifts. These findings suggest that the causal pre-training in TabPFN is helpful but insufficient for algorithmic fairness, highlighting implications for deploying TabPFN (and similar) models in practice and the need for further fairness interventions. Qinyi Liu, Mohammad Khalil, Naman Goel |
WWW | 1 |
| 2025 | Creating Artificial Students that Never Existed: Leveraging Large Language Models and CTGANs for Synthetic Data Generation
Mohammad Khalil, Sam Urmian, Ronas Shakya, Qinyi Liu |
LAK | 4 |
| 2025 | Can Synthetic Data be Fair and Private? A Comparative Study of Synthetic Data Generation and Fairness Algorithms
Qinyi Liu, Oscar Blessed Deho, Sam Urmian, Mohammad Khalil, Srecko Joksimovic, George Siemens |
LAK | 1 |
| 2025 | Advancing privacy in learning analytics using differential privacyabstractThis paper addresses the challenge of balancing learner data privacy with the use of data in learning analytics (LA) by proposing a novel framework by applying Differential Privacy (DP). The need for more robust privacy protection keeps increasing, driven by evolving legal regulations and heightened privacy concerns, as well as traditional anonymization methods being insufficient for the complexities of educational data. To address this, we introduce the first DP framework specifically designed for LA and provide practical guidance for its implementation. We demonstrate the use of this framework through a LA usage scenario and validate DP in safeguarding data privacy against potential attacks through an experiment on a well-known LA dataset. Additionally, we explore the trade-offs between data privacy and utility across various DP settings. Our work contributes to the field of LA by offering a practical DP framework that can support researchers and practitioners in adopting DP in their works. Qinyi Liu, Ronas Shakya, Mohammad Khalil, Jelena Jovanovic 0001 |
LAK | 1 |
| 2024 | Explainable AI in Learning Analytics: Improving Predictive Models and Advancing Transparency TrustabstractAs online education becomes more widely available, the amount of data available and accessible has exploded. Such a wealth of educational data provides ample opportunities for the fields of learning analytics and educational data mining to expand and yield numerous benefits, such as identifying at-risk students, providing personalized feedback, and providing actionable insights. Machine learning and deep learning techniques are commonly used for these and other purposes. However, simply applying these technologies is not enough without allowing humans to understand them. Because only those who understand the logic behind the technology can cultivate trust and make the technology work to its maximum effectiveness. To address the gap, this paper focuses specifically on a case study of explainable machine learning techniques for predicting student performance in online courses, and its contribution is twofold. First, we show how the use of explainable AI techniques can inform model diagnosis and direction forward in situations where performance is less than ideal. Secondly, we show how different types of explainable AI techniques can improve transparency and trust in educational scenarios, allowing stakeholders to benefit from explainable AI. Qinyi Liu, Mohammad Khalil |
EDUCON | 1 |
| 2024 | Scaling While Privacy Preserving: A Comprehensive Synthetic Tabular Data Generation and Evaluation in Learning AnalyticsabstractPrivacy poses a significant obstacle to the progress of learning analytics (LA), presenting challenges like inadequate anonymization and data misuse that current solutions struggle to address. Synthetic data emerges as a potential remedy, offering robust privacy protection. However, prior LA research on synthetic data lacks thorough evaluation, essential for assessing the delicate balance between privacy and data utility. Synthetic data must not only enhance privacy but also remain practical for data analytics. Moreover, diverse LA scenarios come with varying privacy and utility needs, making the selection of an appropriate synthetic data approach a pressing challenge. To address these gaps, we propose a comprehensive evaluation of synthetic data, which encompasses three dimensions of synthetic data quality, namely resemblance, utility, and privacy. We apply this evaluation to three distinct LA datasets, using three different synthetic data generation methods. Our results show that synthetic data can maintain similar utility (i.e., predictive performance) as real data, while preserving privacy. Furthermore, considering different privacy and data utility requirements in different LA scenarios, we make customized recommendations for synthetic data generation. This paper not only presents a comprehensive evaluation of synthetic data but also illustrates its potential in mitigating privacy concerns within the field of LA, thus contributing to a wider application of synthetic data in LA and promoting a better practice for open science. Qinyi Liu, Mohammad Khalil, Jelena Jovanovic 0001, Ronas Shakya |
LAK | 1 |