VLDB 2026 Research / reviewers in the wild / expert
Ruochi Li
dblp:257/1007
· DBLP profile ↗
12ranked-venue papers
1as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdaQE-CG: Adaptive Query Expansion for Web-Scale Generative AI Model and Data Card Generation
Haoxuan Zhang, Ruochi Li, Zhenni Liang, Mehri Sattari, Phat Vo, Collin Qu, Ting Xiao 0003, Junhua Ding 0001, Yang Zhang 0095, Haihua Chen 0002 |
WWW | 2 |
| 2026 | A comprehensive survey on medical concept normalization: Datasets, techniques, applications, and future directions
Haihua Chen 0002, Ruochi Li, Aryan Murthy Illa, Ana D. Cleveland, Junhua Ding 0001 |
J. Biomed. Informatics | 3 |
| 2025 | Unveiling the Merits and Defects of LLMs in Automatic Review Generation for Scientific PapersabstractThe surge in scientific submissions has placed increasing strain on the traditional peer-review process, prompting the exploration of large language models (LLMs) for automated review generation. While LLMs demonstrate competence in producing structured and coherent feedback, their capacity for critical reasoning, contextual grounding, and quality sensitivity remains limited. To systematically evaluate these aspects, we propose a comprehensive evaluation framework that integrates semantic similarity analysis and structured knowledge graph metrics to assess LLM-generated reviews against human-written counterparts. We construct a large-scale benchmark of 1,683 papers and 6,495 expert reviews from ICLR and NeurIPS in multiple years, and generate reviews using five LLMs. Our findings show that LLMs perform well in descriptive and affirmational content, capturing the main contributions and methodologies of the original work, with GPT-4o highlighted as an illustrative example, generating 15.74% more entities than human reviewers in the strengths section of good papers in ICLR 2025. However, they consistently underperform in identifying weaknesses, raising substantive questions, and adjusting feedback based on paper quality. GPT-4o produces 59.42% fewer entities than real reviewers in the weaknesses and increases node count by only 5.7% from good to weak papers, compared to 50% in human reviews. Similar trends are observed across all conferences, years, and models, providing empirical foundations for understanding the merits and defects of LLM-generated reviews and informing the development of future LLM-assisted reviewing tools. Data, code, and more detailed results are publicly available at https://github.com/RichardLRC/Peer-Review. Ruochi Li, Haoxuan Zhang, Edward F. Gehringer, Ting Xiao 0003, Junhua Ding 0001, Haihua Chen 0002 |
ICDM | 1 |
| 2025 | Enhancing data quality in medical concept normalization through large language models
Haihua Chen 0002, Ruochi Li, Ana D. Cleveland, Junhua Ding 0001 |
J. Biomed. Informatics | 2 |
| 2024 | How Much Effort Do You Need to Expend on a Technical Interview? A Study of LeetCode Problem Solving StatisticsabstractA technical interview is the culmination of the recruiting process for hiring software engineers in the tech industry. Many well-known companies, including Amazon, Meta (formerly Facebook), Alphabet (Google), and Microsoft, use it to filter candidates. However, the drawbacks of technical interviews are well-documented, including their lack of real-world relevance, bias towards newer developers, demanding time commitment, and potential to induce unnecessary anxiety and frustration. De-spite these criticisms, there is no clear indication that the industry will alter the format of technical interviews in the near future. To assist student developers in preparing for these challenges, we conducted a quantitative analysis using over 300,000 user profiles from LeetCode, arguably the most popular online platform for preparing software development candidates for interviews. Our analysis aims to provide developers with insights into the effort required to prepare for technical interviews, especially in terms of solving programming questions, to secure a position at a renowned company. Jialin Cui, Runqiu Zhang, Fangtong Zhou, Ruochi Li, Yang Song 0019, Edward F. Gehringer |
CSEE&T | 4 |
| 2024 | LLM-generated Feedback in Real Classes and Beyond: Perspectives from Students and Instructors
Qinjin Jia, Jialin Cui, Haoze Du, M. Parvez Rashid, Ruijie Xi, Ruochi Li, Edward F. Gehringer |
EDM | 6 |
| 2024 | On Assessing the Faithfulness of LLM-generated Feedback on Student Assignments
Qinjin Jia, Jialin Cui, Ruijie Xi, M. Parvez Rashid, Ruochi Li, Edward F. Gehringer |
EDM | 6 |
| 2024 | A Statistical Study of Female Students in a Software Engineering Class: Preparedness, Performance, and ContributionabstractThis is a research-to-practice full paper. Several research studies indicate that women who have opted into a computing career path must regularly contend with negative stereotypes about their technical abilities. These stereotypes are often cited as contributing factors to the underrepresentation of women in computing. To counter these stereotypes and enhance female participation in computer science, numerous interventions have been designed. However, most existing research tends to rely on anecdotal evidence and questionnaires to study these stereotypes. In contrast, our study collected data from over 900 students over a span of eight years and adopted a comprehensive quantitative approach to examine these stereotypes about female students. We utilized pre-class GitHub contribution metrics to evaluate students' programming experience and an array of in-class grading items to measure students' performance. Additionally, we mined the project repositories' git logs to gain insights into students' contributions to team projects. Our investigation began by probing whether there was a notable difference in the technical backgrounds or preparedness between female and male students. The results indicated that males tended to be better prepared. Next, we explored potential disparities in class performance between the two genders. Our findings revealed that males and females each excelled in different areas. We were also interested in discerning if female and male students contributed equally to team projects; our analysis affirmed that the contributions were comparable between the two groups. If allowed to choose their teammates, we examined whether they showed a preference for single-gender teams or mixed-gender teams. Our conclusions indicated no marked preference. This paper aims to augment the body of research on computing education by assisting educators in gaining a better understanding of female students in the class. Moreover, it tests the stereotypes by comparing them with empirical results. Jialin Cui, Runqiu Zhang, Qinjin Jia, Fangtong Zhou, Ruochi Li, Edward F. Gehringer |
FIE | 5 |
| 2024 | A Comparative Analysis of GitHub Contributions Before and After An OSS Based Software Engineering ClassabstractThis study presents a comparative analysis of contributions to GitHub by students before and after participating in a Software Engineering class based on Open Source Software (OSS). The primary objective is to understand the influence of formal software engineering education on students' engagement in OSS projects, as reflected in their GitHub activities. The research addresses two key questions. Firstly, it examines how GitHub contributions change before and after the class. The corresponding hypothesis posits that students' average GitHub contributions will exhibit a distinct pattern post-class compared to pre-class. Additionally, the study explores the potential association between students' academic performance in the class and their level of GitHub contributions after the class. The strength and direction of the potential association are quantified using the Spearman correlation coefficient, considering the potential non-linear nature of the data. This analysis uses data from over 1000 students across more than 10 years, encompassing their GitHub contribution data over multiple timeframes and their grades in the class. The study employs a combination of statistical methods, including paired tests and correlation analysis, to explore these dynamics. While causality cannot be established due to the absence of a control group, the findings offer valuable insights into the correlation between academic engagement and practical contributions in the realm of OSS development. This research contributes to the understanding of how theoretical software engineering education might relate to practical application and engagement in real-world projects. Jialin Cui, Runqiu Zhang, Ruochi Li, Fangtong Zhou, Yang Song 0019, Edward F. Gehringer |
ITiCSE (1) | 3 |
| 2024 | How Pre-class Programming Experience Influences Students' Contribution to Their Team Project: A Statistical StudyabstractGroup or team projects are an essential component of the software engineering curriculum. Earlier studies have explored how prior programming experience influences students' team project performance and overall class performance in software engineering. However, few studies address the impact of prior programming experience on students' contributions to team projects. Previous work has varied in its definitions of prior programming experience or skill, leading to inconsistent findings. In this study, we collected pre-class GitHub contribution metrics from 237 students (forming 79 teams of three) across two academic years to measure their prior programming experience and skills. We also mined students' project repositories' git logs to collect individual student contributions. A central question revolved around whether students with more substantial prior programming experience were indeed more active contributors to their project teams. Interestingly, our data indicated a positive correlation between prior programming experience and contributions to team projects. We further delved into team dynamics. Specifically, we questioned if teams made up of members with comparable skill levels exhibited a more even distribution of contributions. Contrary to expectations, our findings revealed no association between these two variables. Moreover, we investigated the team configurations that might encourage the rise of "free riders"-students who contributed only minimally. This paper seeks to augment the body of research on computing education and assist educators in understanding how prior programming experience impacts students' contributions in team projects. Jialin Cui, Runqiu Zhang, Ruochi Li, Fangtong Zhou, Yang Song 0019, Edward F. Gehringer |
SIGCSE (1) | 3 |
| 2023 | Predicting Students' Software Engineering Class Performance with Machine Learning and Pre-Class GitHub MetricsabstractResearch into predicting students' performance in computer science classes has been conducted globally for over five decades. Numerous metrics, including performance in prior courses, demographic information, and programming experience, have been used to predict success in computer science. Various analytical methods, such as linear regression, decision trees, ensemble methods, and even neural networks, have also been explored. In this study, we investigate whether pre-class GitHub contribution metrics, combined with machine learning techniques, can forecast student performance in a software engineering class. We address two research questions in this paper. Firstly, can pre-class GitHub contribution metrics predict students' performance? Secondly, which machine learning technique is most effective in predicting student performance? We collected data from 802 students over five years and 11 semesters, including pre-class GitHub contribution stats, students' exam grades, project grades, documentation grades, and review writing grades. Eight different machine learning methods were then tested to predict in-class performance using pre-class GitHub contributions. Our results indicate that exam performance can be relatively accurately predicted by machine learning methods. Ensemble methods such as Random Forest, AdaBoost, and XGBoost performed better than other methods. This suggests that pre-class GitHub contribution metrics can be a useful tool for predicting students' performance in software engineering classes, carrying significant implications for educators. This approach can help educators identify at-risk students at the earliest point in the class, enabling early intervention strategies to prevent failure. Our study uniquely utilizes pre-class GitHub contributions, providing a preliminary indication of a student's familiarity with the course material. While prior research has focused on using in-class data to predict student performance, our approach identifies struggling students from the very beginning. We believe this can provide the most beneficial support for students. Jialin Cui, Fangtong Zhou, Runqiu Zhang, Ruochi Li, Edward F. Gehringer |
FIE | 4 |
| 2023 | Correlating Students' Class Performance Based on GitHub Metrics: A Statistical StudyabstractWhat skills does a student need to succeed in a programming class? Ostensibly, previous programming experience may affect a student's performance. Most past studies on this topic use self-reporting questionnaires to query students about their programming experience. This paper presents a novel, unified, and replicable way to measure previous programming experience using students' pre-class GitHub contributions. To our knowledge, we are the first to use GitHub contributions in this way. We conducted a comprehensive statistical study of students in an object-oriented design and development class from 2017 to 2022 (n = 751) to explore the relationships between GitHub contributions (commits, comments, pull requests, etc.) and students' performance on exams, projects, designs, etc. in the class. Several kinds of contributions were shown to have statistically significant correlations with performance in the class. A set of two-samplet -tests demonstrate statistical significance of the difference between the means of some contributions from the high-performing and low-performing groups. Jialin Cui, Runqiu Zhang, Ruochi Li, Yang Song 0019, Fangtong Zhou, Edward F. Gehringer |
ITiCSE (1) | 3 |