VLDB 2026 Research / reviewers in the wild / expert
Yifan Song 0007
dblp:66/7929-7
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0004-6205-2655ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Who Plays Which Role When? Communication Role Dynamics for Peer Recognition and Team Performance PredictionabstractTeam roles offer an interpretable lens on collaboration, yet computational studies of roles often rely on domain-specific personas or datadriven clustering rather than theory-grounded taxonomies.We operationalize a taxonomy of eight communication roles grounded in education literature and annotate a corpus of 6,307 Slack messages from 55 students across 18 teams in a semester-long computer science course project.We evaluate whether LLMs can approximate expert labels, enabling scalable, taxonomy-driven role annotation.Using these role labels, we characterize role dynamics over teams' lifecycles, finding that different roles peak at different moments and that students enact a more diverse set of roles as projects progress.To evaluate the utility of our role constructs, we use them to predict peer recognition, outperforming lexical, conversational, and LLM-prompting baselines.To assess generalizability beyond the educational context, we apply the same role constructs to a public dataset (DeliData) to predict team performance improvement after deliberation, again exceeding prior performance. Yifan Song 0007, Wenxuan Wendy Shi, Brian P. Bailey, Tal August |
ACL (1) | 1 |
| 2026 | Supporting Learners' Use of Imperfect Generative Pedagogical Chatbots: The Role of Chatbot Response Uncertainty and Reduced VerbosityabstractGenerative chatbots promise to scale personalized learning. Most publicly available generative chatbots are designed to provide confident and eloquent responses by default, even when hallucinating. Prior work has observed that learners using such chatbots often engage shallowly and fail to detect chatbot errors due to overtrust, cognitive overload, and prioritization of short-term gains. To address these challenges, this work examines two chatbot design options in a STEM learning context: introducing verbal uncertainty and reducing response verbosity. Using Bayesian causal inference and thematic analysis in a quasi-experimental setting, we found that a less verbose chatbot improved detection of errors with logical fallacies, but did not increase the use of alternative resources. A chatbot that always expressed uncertainty reduced the adoption of incorrect chatbot responses, but had mixed effects on learning outcomes, suggesting the need to increase signal credibility and maintain learners’ engagement in the learning process despite chatbot disuse. Tiffany Wenting Li, Yifan Song 0007, Hari Sundaram, Karrie Karahalios |
CHI | 2 |
| 2026 | From Data to Action: Empowering Students to Assess and Improve Teamwork with Cross-Tool Log DataabstractTeamwork assessment methods, such as peer evaluations, often fail to accurately capture the collaboration processes in team projects. While log data from digital collaboration tools provide objective evidence of contributions, effectively leveraging these data remains challenging. We interviewed 10 instructors and 16 students, then surveyed additional 51 students to identify their valued teamwork behaviors to guide log data usage, and investigate their perceived benefits and concerns. Students prioritized work-related behaviors such as equitable contribution and timeliness, whereas instructors additionally valued interaction behaviors like mutual support. Both groups highlighted significant limitations in raw log data, including omission of offline contributions, insufficient representation of work quality, and misattributing collaborative work to only the person who interacts with the tool. Many students reviewed logs on their own to monitor project progress and workload equity, but lacked structured guidance to meaningfully interpret the insights. Our findings underscore the need for student-centered teamwork assessment approaches that enable students not only to annotate and contextualize their own data, but also to improve their teamwork by reflecting on a rubric of valued teamwork behaviors. Yifan Song 0007, Ritika Vithani, Wenxuan Wendy Shi, Brian P. Bailey |
SIGCSE (1) | 1 |
| 2026 | RIPEL: A Data-Augmented Peer Evaluation System for Assessing Teamwork
Wenxuan Wendy Shi, Jiaqi Linna Niu, Yifan Song 0007, Brian P. Bailey |
SIGCSE (2) | 4 |
| 2025 | Can Learners Navigate Imperfect Generative Pedagogical Chatbots? An Analysis of Chatbot Errors on LearningabstractGenerative pedagogical chatbots offer a promising solution to transform personalized learning at scale, but their benefits are at risk because of the potential of providing inaccurate information. We have a limited understanding of how effectively learners handle factual chatbot errors and how these errors affect learners with varying backgrounds. This study addresses these questions in an ecologically valid open-ended online STEM learning environment. Using Bayesian causal inference and thematic analysis on survey and interview data from a quasi-experimental setting, we found that most participants struggled to detect factual errors even with access to reading materials and the Internet. Undetected errors harmed learning outcomes and self-efficacy, underscoring the need to help learners evaluate chatbot responses. By analyzing participants' evaluation strategies, we identified challenges during error management and suggested ideas on designing effective supporting resources and learner empowerment. Finally, we revealed differential impacts of chatbot errors across learners and called for personalized support and deployment. Tiffany Wenting Li, Yifan Song 0007, Hari Sundaram, Karrie Karahalios |
L@S | 2 |
| 2024 | Programming Plagiarism Detection with Learner DataabstractCourses with programming assignments have long faced the issue of academic integrity violations (AIV) where cheating could harm the outcome of student learning. Checking code similarity in students' final submissions is a common way to mitigate this issue. But this single analysis is insufficient as 1) students can refactor their code to evade the check, 2) mere code similarity may not be strong enough evidence to support an AIV case, particularly for simpler assignments that may have similar solutions, and 3) code similarity cannot reveal much about the actual circumstances and behaviors of plagiarism. Due to the lack of supporting data or tools, many educators either abandon solving these challenges or rely on manual approaches that are not feasible at scale. In this paper, we propose a workflow to solve the above challenges for large programming classes by providing supporting evidence of cheating with additional learner data: detailed submission timelines with scores and source code. Running this workflow in a large advanced programming course over several years has helped us identify many cheating cases effectively and efficiently. Yifan Song 0007, Yuanxin Wang 0001, Marshall An, Christopher Bogart, Majd F. Sakr |
SIGCSE (2) | 1 |
| 2023 | Can Generative Pre-trained Transformers (GPT) Pass Assessments in Higher Education Programming Courses?abstractWe evaluated the capability of generative pre-trained transformers (GPT), to pass assessments in introductory and intermediate Python programming courses at the postsecondary level. Discussions of potential uses (e.g., exercise generation, code explanation) and misuses (e.g., cheating) of this emerging technology in programming education have intensified, but to date there has not been a rigorous analysis of the models' capabilities in the realistic context of a full-fledged programming course with diverse set of assessment instruments. We evaluated GPT on three Python courses that employ assessments ranging from simple multiple-choice questions (no code involved) to complex programming projects with code bases distributed into multiple files (599 exercises overall). Further, we studied if and how successfully GPT models leverage feedback provided by an auto-grader. We found that the current models are not capable of passing the full spectrum of assessments typically involved in a Python programming course (<70% on even entry-level modules). Yet, it is clear that a straightforward application of these easily accessible models could enable a learner to obtain a non-trivial portion of the overall available score (>55%) in introductory and intermediate courses alike. While the models exhibit remarkable capabilities, including correcting solutions based on auto-grader's feedback, some limitations exist (e.g., poor handling of exercises requiring complex chains of reasoning steps). These findings can be leveraged by instructors wishing to adapt their assessments so that GPT becomes a valuable assistant for a learner as opposed to an end-to-end solution. Jaromír Savelka, Arav Agarwal, Christopher Bogart, Yifan Song 0007, Majd F. Sakr |
ITiCSE (1) | 4 |
| 2022 | Cheating Detection in Online Assessments via Timeline AnalysisabstractThe potential for academic integrity violations increases in online courses and instructors must place extra attention on academic integrity, since cheating techniques and costs are different than in the physical classroom. Although students are less supervised and able to study in a self-paced mode in online learning, unauthorized collaboration is still considered to be a serious integrity violation. However, online learning platforms have the advantage that they may capture detailed timelines of student activity. Analysis of these can enable instructors to detect many patterns of collaboration, e.g., working on assessments together, or copying solutions from unauthorized web pages. In this paper, we describe detection methods for several common patterns of alignment between work timelines of pairs of students, and these patterns' relationship with corroborative evidence such as similar answers and unusually fast completion times. We describe data collection necessary to apply the timeline analysis technique to weekly quiz assessments and project submissions, and discuss the strength of evidence the technique can provide in different situations. We have been applying these techniques in an online project-based course over several years, and it has helped instructors to successfully identify potential cheating cases. Jiameng Du, Yifan Song 0007, Mingxiao An, Marshall An, Christopher Bogart, Majd F. Sakr |
SIGCSE (1) | 2 |