Chunmei Shi

dblp:21/2870 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Dynamic Grade Prediction in Programming Education Using Time-Series XGBoost and SHAP Analysis
abstract
This study presents a time-series-based machine learning approach to predict final grades in a C programming course, leveraging temporal and behavioral features to support early intervention. We evaluated classification and regression models using three-class (Needs Improvement, Average, Excellent) and five-class (Fail, Poor, Average, Good, Excellent) grading schemes, addressing class imbalance with RandomOverSampler, SMOTE and ADASYN. Advanced sampling strategies, particularly SMOTE, enhanced minority class prediction in the three-class scheme, with XGBoost achieving superior performance. The five-class scheme offered finer granularity, revealing nuanced patterns in mid-tier performance through practice-related features, but faced challenges from increased class imbalance. Regression models, while suitable for continuous prediction, underperformed due to thresholding biases. SHAP analysis identified historical average score and difficulty-adjusted score as key predictors, providing actionable insights for educators. These findings highlight the trade-offs between broad and fine-grained prediction, with the three-class scheme supporting robust interventions and the five-class scheme enabling nuanced feedback. Future work includes incorporating qualitative features and hybrid approaches to improve fine-grained prediction and generalizability across educational contexts.
Jing Qiu 0003, Chunmei Shi
Int. J. Softw. Eng. Knowl. Eng.2
2024 PanDepth, an ultrafast and efficient genomic tool for coverage calculation
abstract
Coverage quantification is required in many sequencing datasets within the field of genomics research. However, most existing tools fail to provide comprehensive statistical results and exhibit limited performance gains from multithreading. Here, we present PanDepth, an ultra-fast and efficient tool for calculating coverage and depth from sequencing alignments. PanDepth outperforms other tools in computation time and memory efficiency for both BAM and CRAM-format alignment files from sequencing data, regardless of read length. It employs chromosome parallel computation and optimized data structures, resulting in ultrafast computation speeds and memory efficiency. It accepts sorted or unsorted BAM and CRAM-format alignment files as well as GTF, GFF and BED-formatted interval files or a specific window size. When provided with a reference genome sequence and the option to enable GC content calculation, PanDepth includes GC content statistics, enhancing the accuracy and reliability of copy number variation analysis. Overall, PanDepth is a powerful tool that accelerates scientific discovery in genomics research.
Huiyang Yu, Chunmei Shi, Weiming He, Bo Ouyang
Briefings Bioinform.2
2023 Three Approaches for Detecting Direct Output Cheating in Program Online Judge Systems
abstract
Program online judge (POJ) systems allow students to view questions, submit solution code, and receive scores automatically via the web. Most POJs use test cases for scoring. When a POJ is scored by test case pass rate or a problem that has only one test case, students can usually score by providing the direct output of the test cases (direct output cheating). Currently, there is only one work on detecting such cheating. However, its precision is very low. To solve this problem, three novel approaches are proposed to detect direct output cheating: (i) Line Statistics, which computes the proportion of output calls against other statements; (ii) the control flow graph (CFG) Search computes the maximum similarity between the CFG of a program and that of known samples; (iii) abstract syntax tree (AST) Search identifies cheating by matching rules that are summarized from ASTs of previously detected cheating attempts. A student’s code is marked as cheating if the similarity exceeds a predefined threshold; and a program is detected as cheating if the proportion exceeds a predefined threshold. The proposed approaches and three well-known code plagiarism detection tools (JPlag, Sherlock, and SIM) were evaluated using 100,000 submissions for 1153 problems from a POJ based on the C programming language. The F1 scores of these approaches were determined as 0.9752 (AST Search), 0.9440 (CFG Search), 0.7405 (Line Statistics), 0.6446 (JPlag), 0.1587 (Sherlock), and 0.0076 (SIM), respectively. The result indicates that (i) AST Search is most suitable for the detection of direct output cheating; (ii) traditional code search or plagiarism detection methods based on similarity calculations are not effective for complex cheat detection because these cheats are highly similar to normal code.
Jing Qiu 0003, Chunmei Shi, Yuehua Lv
Int. J. Softw. Eng. Knowl. Eng.2
2018 The Delta Generalized Labeled Multi-Bernoulli Filter for Cell Tracking
abstract
Cell tracking automatically in time-lapse image sequences is important for understanding the dynamic pattern of micro-cell. In this paper, we present a novel method for tracking cell with shape feature based on the delta generalized labeled multi-Bernoulli (delta-GLMB) filter which is of great research significance. The delta-GLMB filter with cell shape parameters can improve the tracking accuracy. This approach is evaluated and compared with raw detection using the generalized optimal sub-pattern assignment (GOSPA) metric on real N2DH-SIM cell sequences. Experiment results show that the delta-GLMB filter can provide the shape information as well as the better estimation than raw detection and KTH method.
Chunmei Shi, Junjie Wang 0005, Lingling Zhao, Xiaohong Su, Guangshun Jiang
BIBE1
2014 Multi-object Tracking Based on Particle Probability Hypothesis Density Tracker in Microscopic Video
abstract
Research on biological objects requires tracking hundreds of micro-objects from the microscopy video. We propose an automated tracking framework to extract trajectories of micro-objects. This framework uses a particle probability hypothesis density (PF-PHD) tracker to implement a recursive Bayesian state estimation and trajectories association. In the framework, an ellipse target model is presented to describe the micro-objects with shape parameters instead of point-like targets. Furthermore, an orientation and positional constraint model is developed to deal with the data association of crossing trajectories in multitarget tracking. Using this framework, a significantly larger number of tracks are obtained than manual tracking. The experiments on simulated image sequences of microtubule movement are performed in order to evaluate the proposed PF-PHD tracking method.
Chunmei Shi, Lingling Zhao, Peijun Ma, Xiaohong Su, Junjie Wang 0005, Chiping Zhang
BIBE1
2009 Object tracking using SIFT features and mean shift
Huiyu Zhou 0001, Yuan Yuan 0001, Chunmei Shi
Comput. Vis. Image Underst.3
2009 Non-rigid object tracking in complex scenes
Huiyu Zhou 0001, Yuan Yuan 0001, Chunmei Shi
Pattern Recognit. Lett.4