Yang Song 0019

dblp:24/4470-19 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
6since 2021 · last 2024
0000-0001-5297-3072ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 14 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2024 How Much Effort Do You Need to Expend on a Technical Interview? A Study of LeetCode Problem Solving Statistics
abstract
A technical interview is the culmination of the recruiting process for hiring software engineers in the tech industry. Many well-known companies, including Amazon, Meta (formerly Facebook), Alphabet (Google), and Microsoft, use it to filter candidates. However, the drawbacks of technical interviews are well-documented, including their lack of real-world relevance, bias towards newer developers, demanding time commitment, and potential to induce unnecessary anxiety and frustration. De-spite these criticisms, there is no clear indication that the industry will alter the format of technical interviews in the near future. To assist student developers in preparing for these challenges, we conducted a quantitative analysis using over 300,000 user profiles from LeetCode, arguably the most popular online platform for preparing software development candidates for interviews. Our analysis aims to provide developers with insights into the effort required to prepare for technical interviews, especially in terms of solving programming questions, to secure a position at a renowned company.
Jialin Cui, Runqiu Zhang, Fangtong Zhou, Ruochi Li, Yang Song 0019, Edward F. Gehringer
CSEE&T5
2024 Utilizing the Constrained K-Means Algorithm and Pre-Class GitHub Contribution Statistics for Forming Student Teams
abstract
In modern software engineering education, team formation is crucial for mimicking real-world collaborative scenarios and boosting project-based learning outcomes. This paper introduces a simple, innovative, and universally adaptable method for forming student teams within a software engineering class. We utilize publicly available pre-class GitHub metrics as our input variables (e.g., number of commits, pull requests, code size, etc.). For team formation, the constrained k-means algorithm is employed. This algorithm embraces domain-specific constraints, ensuring the resulting teams not only resonate with the inherent data clusters but also meet educational requirements. Preliminary results suggest that our methodology yields teams with a harmonious blend of skills, experiences, and collaborative potentials, thereby setting the stage for enhanced project success and enriched learning experiences. Quantitative analyses show that teams formed via our approach outperform both randomly assembled teams and student self-selected teams concerning project grades. Moreover, teams created using our method also display a reduced standard deviation in grades, suggesting a more consistent performance across the board.
Jialin Cui, Fangtong Zhou, Qinjin Jia, Yang Song 0019, Edward F. Gehringer
ITiCSE (1)5
2024 A Comparative Analysis of GitHub Contributions Before and After An OSS Based Software Engineering Class
abstract
This study presents a comparative analysis of contributions to GitHub by students before and after participating in a Software Engineering class based on Open Source Software (OSS). The primary objective is to understand the influence of formal software engineering education on students' engagement in OSS projects, as reflected in their GitHub activities. The research addresses two key questions. Firstly, it examines how GitHub contributions change before and after the class. The corresponding hypothesis posits that students' average GitHub contributions will exhibit a distinct pattern post-class compared to pre-class. Additionally, the study explores the potential association between students' academic performance in the class and their level of GitHub contributions after the class. The strength and direction of the potential association are quantified using the Spearman correlation coefficient, considering the potential non-linear nature of the data. This analysis uses data from over 1000 students across more than 10 years, encompassing their GitHub contribution data over multiple timeframes and their grades in the class. The study employs a combination of statistical methods, including paired tests and correlation analysis, to explore these dynamics. While causality cannot be established due to the absence of a control group, the findings offer valuable insights into the correlation between academic engagement and practical contributions in the realm of OSS development. This research contributes to the understanding of how theoretical software engineering education might relate to practical application and engagement in real-world projects.
Jialin Cui, Runqiu Zhang, Ruochi Li, Fangtong Zhou, Yang Song 0019, Edward F. Gehringer
ITiCSE (1)5
2024 How Pre-class Programming Experience Influences Students' Contribution to Their Team Project: A Statistical Study
abstract
Group or team projects are an essential component of the software engineering curriculum. Earlier studies have explored how prior programming experience influences students' team project performance and overall class performance in software engineering. However, few studies address the impact of prior programming experience on students' contributions to team projects. Previous work has varied in its definitions of prior programming experience or skill, leading to inconsistent findings. In this study, we collected pre-class GitHub contribution metrics from 237 students (forming 79 teams of three) across two academic years to measure their prior programming experience and skills. We also mined students' project repositories' git logs to collect individual student contributions. A central question revolved around whether students with more substantial prior programming experience were indeed more active contributors to their project teams. Interestingly, our data indicated a positive correlation between prior programming experience and contributions to team projects. We further delved into team dynamics. Specifically, we questioned if teams made up of members with comparable skill levels exhibited a more even distribution of contributions. Contrary to expectations, our findings revealed no association between these two variables. Moreover, we investigated the team configurations that might encourage the rise of "free riders"-students who contributed only minimally. This paper seeks to augment the body of research on computing education and assist educators in understanding how prior programming experience impacts students' contributions in team projects.
Jialin Cui, Runqiu Zhang, Ruochi Li, Fangtong Zhou, Yang Song 0019, Edward F. Gehringer
SIGCSE (1)5
2023 Correlating Students' Class Performance Based on GitHub Metrics: A Statistical Study
abstract
What skills does a student need to succeed in a programming class? Ostensibly, previous programming experience may affect a student's performance. Most past studies on this topic use self-reporting questionnaires to query students about their programming experience. This paper presents a novel, unified, and replicable way to measure previous programming experience using students' pre-class GitHub contributions. To our knowledge, we are the first to use GitHub contributions in this way. We conducted a comprehensive statistical study of students in an object-oriented design and development class from 2017 to 2022 (n = 751) to explore the relationships between GitHub contributions (commits, comments, pull requests, etc.) and students' performance on exams, projects, designs, etc. in the class. Several kinds of contributions were shown to have statistically significant correlations with performance in the class. A set of two-samplet -tests demonstrate statistical significance of the difference between the means of some contributions from the high-performing and low-performing groups.
Jialin Cui, Runqiu Zhang, Ruochi Li, Yang Song 0019, Fangtong Zhou, Edward F. Gehringer
ITiCSE (1)4
2021 PyRoboCar: A Low-cost Deep Neural Network-based Autonomous Car
abstract
In this innovative practice WIP paper, we present PyRoboCar, a low-cost deep neural network-based autonomous car project. PyRoboCar is a small-scale replication of a real self-driving car using a deep convolutional neural network (CNN), which takes images from a front fisheye camera as input and produces car steering angles as output. PyRoboCar uses a similar network architecture as industry-level autonomous cars and can drive itself in real-time using a camera, an additional Tensor Processing Unit, and a Raspberry Pi 4 platform. We have also made this project open online, including the code and instructions.
Gordon Hendry, Seth Harris, Ryan Goodwin, Leonid Neverov, Yunkai Xiao, David Tian, Yang Song 0019
FIE7
2020 EDM and Privacy: Ethics and Legalities of Data Collection, Usage, and Storage
Mark Klose, Vasvi Desai, Yang Song 0019, Edward F. Gehringer
EDM3
2020 Detecting Problem Statements in Peer Assessments
Yunkai Xiao, Gabriel Zingle, Qinjin Jia, Harsh R. Shah, Mohsin Karovaliya, Weixiang Zhao, Yang Song 0019, Ashwin Balasubramaniam, Harshit Patel, Priyankha Bhalasubbramanian, Vikram Patel, Edward F. Gehringer
EDM9
2019 A Test-Driven Approach to Improving Student Contributions to Open-Source Projects
abstract
Test-driven development (TDD) promises to help students write high-quality code with fewer defects. Although many studies of TDD usage have been conducted in entry-level computer science courses, few have looked at more advanced students doing projects of larger scope, such as contributing to open-source software (OSS). To test the performance of the test-driven approach on OSS-based course projects, we conducted a quasi-experimental controlled study, which lasted for more than one month. Thirty-five masters students participated in our study. They worked on course projects in teams, half of which were assigned to the TDD group (using a test-driven approach), and the rest of which were assigned to the non-TDD group (using the traditional test-last approach). We found that students in the TDD group were able to apply test-driven techniques pragmatically-spending more than 20% of their time on average complying with the test-driven process-throughout the whole project. There were no major differences in the quality of source-code modifications and newly added tests between the TDD group and the non-TDD group; however, the TDD group wrote more tests and achieved significantly higher (12% more) statement coverage.
Zhewei Hu, Yang Song 0019, Edward F. Gehringer
FIE2
2018 Early Detection on Students' Failing Open-Source based Course Projects using Machine Learning Approaches: (Abstract Only)
abstract
Open-source course projects offer students a glimpse of real-world projects and opportunities to learn about architectural design and coding style. While students often have more difficulties with these projects than with traditional "toy" projects, instructors are also spending excessive time on grading miscellaneous projects. There is an improvising need for means to help students and instructors with their difficulties. This poster presents our work on predicting which course projects are likely to fail at an early stage with machine learning approaches. We collected metadata from 247 course projects in a graduate-level Object-Oriented Design and Development course over the past 5 years, built models to fit the course projects and use the classifier to help instructors to identify potential failing projects, thus to help students to salvage their works. By assuming that the project acceptances are related to the working patterns of project teams, we made innovations of adding temporal-based patterns into the training data, and achieved 86.36% classification accuracy with the addition of those features. We also proved several observations, such as most of the rejected projects are those begun relatively late during the project period, and the projects which modified more files/code does not result in better possibility of being accepted. By contrast, accepted projects tend to deliver a volume of code that is neither very small nor very large, compared to rejected ones. Our results also suggest that setting milestone checkpoints at roughly a week before the submission deadlines would enable more students to succeed in their OSS projects.
Yang Song 0019, Edward F. Gehringer
SIGCSE2
2017 Collusion in educational peer assessment: How much do we need to worry about it?
abstract
Several decades of research have shown peer assessment to be an effective pedagogical approach. Researchers have shown that peer assessment has the potential to provide students more copious, timely and helpful feedback, and also helping reviewers to learn as well. In recent decades, peer assessments in educational settings have increasingly been facilitated by online tools. Some MOOC platforms also rely on aggregated peer-assessment scores to assign grades for each artifact. However, in peer assessment, students can potentially game this process, and thereby harm the validity and reliability of the aggregated scores. Most of instructors assume that the majority of the students' peer assessments are honest since most of the peer assessment is done in double-blind fashion. This assumption only holds in the absence of organized collusion - when no more than a small number of students game the peer assessment and give each other very high scores. This paper identifies two types of collusion that we have observed. They are small-circle collusion and pervasive collusion. Small-circle collusion refers to the behaviors of students who form small circles and give higher peer review grades to each other. Pervasive collusion refers to students assigning top grades to all the submissions they review. We also present our algorithms for detecting these two types of colluders. Our experiments are based on a peer-assessment dataset shared by multiple peer-assessment systems. By removing these colluders' peer assessments, we are able to estimate how much inflation is brought by colluders in educational peer assessment.
Yang Song 0019, Zhewei Hu, Edward F. Gehringer
FIE1
2016 Five years of extra credit in a studio-based course: An effort to incentivize socially useful behavior
abstract
In studio-based education, students collaboratively work on projects that allow them to learn aspects of design through experience on authentic artifacts, often for outside clients. In our Object-Oriented Design and Development class, this means work on open-source projects, such as Mozilla Servo, OpenMRS, Sahana Eden, and Apache Ambari, as well as our own Expertiza project. This paper recounts five years of experience awarding extra credit for activity that students engaged in to help fellow students. These activities included doing extra peer reviews of classmates' work, helping other teams with their projects, writing various types of quiz questions, and answering questions on the Piazza message board. The incentives often led to “gaming” behavior, where students engaged in activities simply to earn points, with very little benefit to their fellow students. Consequently, the scoring policy has been changed several times during the five-year period. We show how student behavior changed as the rules changed, how we have enhanced review quality, improved responses and response time on Piazza, and transitioned students away from merely helping with project setup, and toward helping other students with the substance of their projects.
Edward F. Gehringer, Zhewei Hu, Yang Song 0019
FIE3
2016 A quantitative case study on students' strategy for using authorized cheat-sheets
abstract
Traditional formal tests are usually given in a time-limited, closed book/notes style because instructors believe this approach measures students' learning. Nonetheless, other researchers argue that there are better alternatives and allowing students to use authorized cheat-sheets is one of them. The most important reason for allowing cheat-sheets is to help students to focus more on greater understanding and deeper learning. However, some researchers have questioned the efficacy of authorized cheat-sheets in formal exams. While it might be true, in some cases, that authorized cheat-sheets function like “crutches” in exams, we should not ignore that they are also learning tools. To make the cheat-sheets, students need to read the class material, process information actively, and select, organized, prioritize the content for the cheat-sheets. Our earlier research showed that in an undergraduate engineering course, the quality of students' cheat-sheet is strongly positively correlated with the students' performance on the exams. Moreover, the quality of the cheat-sheets improved through the sequence of exams during the semester and the students' grades did likewise. In this research, we collected more than 300 cheat-sheets from two sections of a graduate course on the same subject as our previous research. The class design of the graduate level course is similar to the undergraduate one but it covers more conceptually difficult topics at a greater depth including theory and proofs. We use the same cheat-sheet rating scheme and compare the quality of graduate students' cheat-sheets with the undergraduate students' on multiple dimensions including density, organization, number of sample answers, number of formulas and number of graph representations. We discovered significant differences between the graduate and undergraduate students' use of and how to best create authorized cheat-sheets and formulate useful directives.
Yang Song 0019, David Thuente
FIE1
2016 An experiment with separate formative and summative rubrics in educational peer assessment
abstract
Educational peer assessment has proven to be a powerful approach for providing students timely feedback and allowing them to help and learn from each other. In an educational setting, most peer assessment consists of a single round. The problem with this setting is that, either the authors do not have a chance to update their work, which makes the suggestions from their peers useless, or the author can make changes after receiving the peer reviews, which forecloses using peer review to help assign grades. To address these issues, in our classes we now use two rounds of online review, with a different rubric for each. Our Expertiza peer-review system allows the evaluation rubric to vary by rounds. In the first review round, we present a formative review rubric to the peer reviewers. In the formative rubric, we try to encourage student reviewers to look into details, point out the problems they can find in the author's work, and offer insightful suggestions. After the formative review round, authors have the opportunity to submit an updated version of their artifacts. Next comes a summative peer-review round, using a summative rubric. A summative peer-review rubric focused more on evaluating the quality of the artifact by comparing it against specific benchmarks. In this paper, we discuss the design of the two-round peer-review assignments in a computer-science course and present our observations on student peer-review activity. An analysis of students' peer-assessment responses confirms the effectiveness of this design of peer-review activity.
Yang Song 0019, Zhewei Hu, Edward F. Gehringer
FIE1
2016 A markup language for building a data warehouse for educational peer-assessment research
abstract
Peer assessment has proved to be a useful technique in all levels of education. The process of giving and receiving comments can encourage critical thinking and help students learn both from reviewing and being reviewed. Peer assessment generates a large volume of data, especially if done online. Online peer-assessment systems are designed differently and use different schema for their data, which complicates the work of comparing different designs. For example, some systems are based on ranking - reviewers rank the artifacts they are asked to assess, while other systems use rating - reviewers assess a single artifact at a time and score it on various criteria. Comparing these two types of systems, e.g. on rating accuracy, or usefulness of formative feedback, can be challenging because researchers need to learn the design and terminology of each system before analyzing the data. We introduce a Peer-Review Markup Language to provide a common definition of terminology across multiple systems. We are using this markup language to build a data warehouse for data from different systems. We discuss issues raised during this process and our approach to solving them.
Yang Song 0019, Ferry Pramudianto, Edward F. Gehringer
FIE1
2015 Pluggable reputation systems for peer review: A web-service approach
abstract
Peer review has long been used in education to provide students more timely feedback and allow them to learn from each other's work. In large courses and MOOCs, there is also interest in having students determine, or help determine, their classmates' grades. This requires a way to tell which peer reviewers' scores are credible. This can be done by comparing scores assigned by different reviewers with each other, and with scores that the instructor would have assigned. For this reason, several reputation systems have been designed; but until now, they have not been compared with each other, so we have no information about which performs best. To make the reputation algorithms pluggable for different peer-review system, we are carrying out a project to develop a reputation web service. This paper compares two reputation algorithms, each of which has two versions, and reports on our efforts to make them “pluggable,” so they can easily be adopted by different peer-review systems. Toward this end, we have defined a Peer-Review Markup Language (PRML), which is a generic schema for data sharing among different peer-review systems.
Yang Song 0019, Zhewei Hu, Edward F. Gehringer
FIE1