VLDB 2026 Research / reviewers in the wild / expert
Binglin Chen
dblp:198/7452
· DBLP profile ↗
16ranked-venue papers
9as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 7 first-author · 1 since 2021Artificial intelligence and machine learning · 7 · 5 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 first-author · 4 since 2021Systems, architecture and hardware · 5 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cross-modal Causal Relation Alignment for Video Question GroundingabstractVideo question grounding (VideoQG) requires models to answer the questions and simultaneously infer the relevant video segments to support the answers. However, existing VideoQG methods usually suffer from spurious cross-modal correlations, leading to a failure to identify the dominant visual scenes that align with the intended question. Moreover, vision-language models exhibit unfaithful generalization performance and lack robustness on challenging downstream tasks such as VideoQG. In this work, we propose a novel VideoQG framework named Cross-modal Causal Relation Alignment (CRA), to eliminate spurious correlations and improve the causal consistency between question-answering and video temporal grounding. Our CRA involves three essential components: i) Gaussian Smoothing Grounding (GSG) module for estimating the time interval via cross-modal attention, which is de-noised by an adaptive Gaussian filter, ii) Cross-Modal Alignment (CMA) enhances the performance of weakly supervised VideoQG by leveraging bidirectional contrastive learning between estimated video segments and QA features, iii) Explicit Causal Intervention (ECI) module for multimodal deconfounding, which involves front-door intervention for vision and backdoor intervention for language. Extensive experiments on two VideoQG datasets demonstrate the superiority of our CRA in discovering visually grounded content and achieving robust question reasoning. Codes are available at https://github.com/WissingChen/CRA-GQA. Yang Liu 0084, Binglin Chen, Jiandong Su, Yongsen Zheng, Liang Lin 0004 |
CVPR | 3 |
| 2025 | Evaluating AI Models for Autograding Explain in Plain English Questions: Challenges and ConsiderationsabstractCode-reading ability has traditionally been under-emphasized in assessments as it is difficult to assess at scale. Prior research has shown that code-reading and code-writing are closely related skills; thus being able to assess and train code reading skills may be necessary for student learning. One way to assess code-reading ability is using Explain in Plain English (EiPE) questions, which ask students to describe what a piece of code does with natural language. Previous research deployed a binary (correct/incorrect) autograder using bigram models that performed comparably with human teaching assistants on student responses. With a dataset of 3,064 student responses from 17 EiPE questions, we investigated multiple autograders for EiPE questions. We evaluated methods as simple as logistic regression trained on bigram features, to more complicated Support Vector Machines (SVMs) trained on embeddings from Large Language Models (LLMs) to GPT-4. We found multiple useful autograders, most with accuracies in the \(86\!\!-\!\!88\%\) range, with different advantages. SVMs trained on LLM embeddings had the highest accuracy; few-shot chat completion with GPT-4 required minimal human effort; pipelines with multiple autograders for specific dimensions (what we call 3D autograders) can provide fine-grained feedback; and code generation with GPT-4 to leverage automatic code testing as a grading mechanism in exchange for slightly more lenient grading standards. While piloting these autograders in a non-major introductory Python course, students had largely similar views of all autograders, although they more often found the GPT-based grader and code-generation graders more helpful and liked the code-generation grader the most. Maxwell Fowler, Chinedu Emeka, Binglin Chen, David H. Smith IV, Matthew West 0001, Craig B. Zilles |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2025 | ODMixer: Fine-Grained Spatial-Temporal MLP for Metro Origin-Destination PredictionabstractMetro Origin-Destination (OD) prediction is a crucial yet challenging spatial-temporal prediction task in urban computing, which aims to accurately forecast cross-station ridership for optimizing metro scheduling and enhancing overall transport efficiency. Analyzing fine-grained and comprehensive relations among stations effectively is imperative for metro OD prediction. However, existing metro OD models either mix information from multiple OD pairs from the station's perspective or exclusively focus on a subset of OD pairs. These approaches may overlook fine-grained relations among OD pairs, leading to difficulties in predicting potential anomalous conditions. To address these challenges, we learn traffic evolution from the perspective of all OD pairs and propose a fine-grained spatialtemporal MLP architecture for metro OD prediction, namely ODMixer. Specifically, our ODMixer has double-branch structure and involves the Channel Mixer, the Multi-view Mixer, and the Bidirectional Trend Learner. The Channel Mixer aims to capture short-term temporal relations among OD pairs, the Multi-view Mixer concentrates on capturing spatial relations from both origin and destination perspectives. To model long-term temporal relations, we introduce the Bidirectional Trend Learner. Extensive experiments on two large-scale metro OD prediction datasets HZMOD and SHMO demonstrate the advantages of our ODMixer. Our code is available at https://github.com/KLatitude/ODMixer Yang Liu 0084, Binglin Chen, Yongsen Zheng, Lechao Cheng, Guanbin Li, Liang Lin 0004 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Plagiarism in the Age of Generative AI: Cheating Method Change and Learning Loss in an Intro to CS CourseabstractBackground: ChatGPT became widespread in early 2023 and enabled the broader public to use powerful generative AI, creating a new means for students to complete course assessments. Binglin Chen, Colleen M. Lewis, Matthew West 0001, Craig B. Zilles |
L@S | 1 |
| 2023 | "\"I Don't Gamble To Make My Livelihood\": Understanding the Incentives ForabstractBackground: Prior work has primarily been concerned with identifying: (1) how Open Education Resources (OERs) can be used to increase the availability of educational materials, (2) what motivations are behind their adoption and usage in classrooms, and (3) what barriers impede said adoption. However, there is relatively little work investigating the motives and barriers to contribution in OER. Maxwell Fowler, David H. Smith IV, Binglin Chen, Craig B. Zilles |
ICER (1) | 3 |
| 2022 | Peer-grading "Explain in Plain English": A Bayesian Calibration Method for Categorical Answersabstract"Explain in plain English'' (EipE) questions have been proposed as an important activity and assessment for studying novice programmers' grasp of programming knowledge and their ability to communicate their understanding. However, EipE questions aren't widely used in introductory programming courses in part because of the large grading effort required. In this paper, we present our experience of using peer grading for EipE questions in a large-enrollment introductory programming course, where students were asked to categorize other students' responses. We developed a novel Bayesian algorithm for performing calibrated peer grading on categorical data, and we used a heuristic grade assignment method based on the Bayesian estimates. The peer-grading exercises served both as a way to coach students on what is expected from EipE questions and as a way to alleviate the grading load for the course staff. Based on four rounds of peer-grading activities, we found that students are generally capable of categorizing responses to EiPE questions and that our proposed Bayesian method is more robust than unweighted voting. Binglin Chen, Matthew West 0001, Craig B. Zilles |
SIGCSE (1) | 1 |
| 2021 | How should we 'Explain in plain English'? Voices from the Communityabstract“Explain in plain English” (EipE) questions are seen as an important developmental activity and assessment tool in the research community studying how people learn to program, but they aren’t widely used in practice because of difficulty of grading and workload issues. In this paper, we interviewed eleven members of the introductory programming education research community about their thoughts on EipE questions as a whole and how individual borderline student answers should be graded. Through inductive coding of the interview transcripts, we identify: (1) themes relating to how EipE questions should be used in class, (2) the importance of training students to complete EipE questions, (3) standards for the selection and presentation of code in EipE questions, (4) the theoretical and practical considerations relating to grading EipE questions, and (5) English as a second language (ESL) concerns. In addition, we attempt to extrapolate from our observations what the underlying grading process is that faculty are using to grade EipE questions. Maxwell Fowler, Binglin Chen, Craig B. Zilles |
ICER | 2 |
| 2021 | AutogradingabstractPrevious research suggests that "Explain in Plain English" (EiPE) code reading activities could play an important role in the development of novice programmers, but EiPE questions aren't heavily used in introductory programming courses because they (traditionally) required manual grading. We present what we believe to be the first automatic grader for EiPE questions and its deployment in a large-enrollment introductory programming course. Based on a set of questions deployed on a computer-based exam, we find that our implementation has an accuracy of 87-89%, which is similar in performance to course teaching assistants trained to perform this task and compares favorably to automatic short answer grading algorithms developed for other domains. In addition, we briefly characterize the kinds of answers that the current autograder fails to score correctly and the kinds of errors made by students. Maxwell Fowler, Binglin Chen, Sushmita Azad, Matthew West 0001, Craig B. Zilles |
SIGCSE | 2 |
| 2020 | Strategies for Deploying Unreliable AI Graders in High-Transparency High-Stakes Exams
Sushmita Azad, Binglin Chen, Maxwell Fowler, Matthew West 0001, Craig B. Zilles |
AIED (1) | 2 |
| 2020 | Learning to Cheat: Quantifying Changes in Score Advantage of Unproctored Assessments Over TimeabstractProctoring educational assessments (e.g., quizzes and exams) has a cost, be it in faculty (and/or course staff) time or in money to pay for proctoring services. Previous estimates of the utility of proctoring (generally by estimating the score advantage of taking an exam without proctoring) vary widely and have mostly been implemented using an across subjects experimental designs and sometimes with low statistical power. Binglin Chen, Sushmita Azad, Maxwell Fowler, Matthew West 0001, Craig B. Zilles |
L@S | 1 |
| 2020 | A Validated Scoring Rubric for Explain-in-Plain-English QuestionsabstractPrevious research has identified the ability to read code and understand its high-level purpose as an important developmental skill that is harder to do (for a given piece of code) than executing code in one's head for a given input ("code tracing"), but easier to do than writing the code. Prior work involving code reading ("Explain in plain English") problems, have used a scoring rubric inspired by the SOLO taxonomy, but we found it difficult to employ because it didn't adequately handle the three dimensions of answer quality: correctness, level of abstraction, and ambiguity. In this paper, we describe a 7-point rubric that we developed for scoring student responses to "Explain in plain English'' questions, and we validate this rubric through four means. First, we find that the scale can be reliably applied with with a median Krippendorff's alpha (inter-rater reliability) of 0.775. Second, we report on an experiment to assess the validity of our scale. Third, we find that a survey consisting of 12 code reading questions had a high internal consistency (Cronbach's alpha = 0.954). Last, we find that our scores for code reading questions in a large enrollment (N = 452) data structures course are correlated (Pearson's R = 0.555) to code writing performance to a similar degree as found in previous work. Binglin Chen, Sushmita Azad, Rajarshi Haldar, Matthew West 0001, Craig B. Zilles |
SIGCSE | 1 |
| 2019 | Effect of Discrete and Continuous Parameter Variation on Difficulty in Automatic Item Generation
Binglin Chen, Craig B. Zilles, Matthew West 0001, Timothy Bretl |
AIED (1) | 1 |
| 2019 | Predicting the difficulty of automatic item generators on exams from their difficulty on homeworksabstractTo design good assessments, it is useful to have an estimate of the difficulty of a novel exam question before running an exam. In this paper, we study a collection of a few hundred automatic item generators (short computer programs that generate a variety of unique item instances) and show that their exam difficulty can be roughly predicted from student performance on the same generator during pre-exam practice. Specifically, we show that the rate that students correctly respond to a generator on an exam is on average within 5% of the correct rate for those students on their last practice attempt. This study is conducted with data from introductory undergraduate Computer Science and Mechanical Engineering courses. Binglin Chen, Matthew West 0001, Craig B. Zilles |
L@S | 1 |
| 2018 | Towards a Model-Free Estimate of the Limits to Student Modeling Accuracy
Binglin Chen, Matthew West 0001, Craig B. Zilles |
EDM | 1 |
| 2018 | How much randomization is needed to deter collaborative cheating on asynchronous exams?abstractThis paper investigates randomization on asynchronous exams as a defense against collaborative cheating. Asynchronous exams are those for which students take the exam at different times, potentially across a multi-day exam period. Collaborative cheating occurs when one student (the information producer) takes the exam early and passes information about the exam to other students (the information consumers) that are taking the exam later. Using a dataset of computerized exam and homework problems in a single course with 425 students, we identified 5.5% of students (on average) as information consumers by their disproportionate studying of problems that were on the exam. These information consumers ("cheaters") had a significant advantage (13 percentage points on average) when every student was given the same exam problem (even when the parameters are randomized for each student), but that advantage dropped to almost negligible levels (2--3 percentage points) when students were given a random problem from a pool of two or four problems. We conclude that randomization with pools of four (or even three) problems, which also contain randomized parameters, is an effective mitigation for collaborative cheating. Our analysis suggests that this mitigation is in part explained by cheating students having less complete information about larger pools. Binglin Chen, Matthew West 0001, Craig B. Zilles |
L@S | 1 |
| 2017 | Do Performance Trends Suggest Wide-spread Collaborative Cheating on Asynchronous Exams?abstractUsing a data set from 29,492 asynchronous exams in an on-campus proctored computer-based testing facility (CBTF), we observed correlations between when a student chooses to take their exam within the exam period and their score on the exam. Somewhat surprisingly, instead of increasing throughout the exam period, which might be indicative of widespread collaborative cheating, we find that exam scores decrease throughout the exam period. While this could be attributed to weaker students putting off exams, this effect holds even when accounting for student ability as measured by a synchronous exam taken during the same semester. This suggests that precautions can be taken by a CBTF to maintain cheating at a low level (e.g., the level of proctored synchronous exams), in spite of the fact that students are taking their exams over a multi-day period. Binglin Chen, Matthew West 0001, Craig B. Zilles |
L@S | 1 |