VLDB 2026 Research / reviewers in the wild / expert
Jing Qiu 0003
dblp:20/1461-3
· DBLP profile ↗
10ranked-venue papers
7as first author
4since 2021 · last 2025
0000-0003-3264-1681ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 6 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2Security and privacy · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dynamic Grade Prediction in Programming Education Using Time-Series XGBoost and SHAP AnalysisabstractThis study presents a time-series-based machine learning approach to predict final grades in a C programming course, leveraging temporal and behavioral features to support early intervention. We evaluated classification and regression models using three-class (Needs Improvement, Average, Excellent) and five-class (Fail, Poor, Average, Good, Excellent) grading schemes, addressing class imbalance with RandomOverSampler, SMOTE and ADASYN. Advanced sampling strategies, particularly SMOTE, enhanced minority class prediction in the three-class scheme, with XGBoost achieving superior performance. The five-class scheme offered finer granularity, revealing nuanced patterns in mid-tier performance through practice-related features, but faced challenges from increased class imbalance. Regression models, while suitable for continuous prediction, underperformed due to thresholding biases. SHAP analysis identified historical average score and difficulty-adjusted score as key predictors, providing actionable insights for educators. These findings highlight the trade-offs between broad and fine-grained prediction, with the three-class scheme supporting robust interventions and the five-class scheme enabling nuanced feedback. Future work includes incorporating qualitative features and hybrid approaches to improve fine-grained prediction and generalizability across educational contexts. Jing Qiu 0003, Chunmei Shi |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2023 | Three Approaches for Detecting Direct Output Cheating in Program Online Judge SystemsabstractProgram online judge (POJ) systems allow students to view questions, submit solution code, and receive scores automatically via the web. Most POJs use test cases for scoring. When a POJ is scored by test case pass rate or a problem that has only one test case, students can usually score by providing the direct output of the test cases (direct output cheating). Currently, there is only one work on detecting such cheating. However, its precision is very low. To solve this problem, three novel approaches are proposed to detect direct output cheating: (i) Line Statistics, which computes the proportion of output calls against other statements; (ii) the control flow graph (CFG) Search computes the maximum similarity between the CFG of a program and that of known samples; (iii) abstract syntax tree (AST) Search identifies cheating by matching rules that are summarized from ASTs of previously detected cheating attempts. A student’s code is marked as cheating if the similarity exceeds a predefined threshold; and a program is detected as cheating if the proportion exceeds a predefined threshold. The proposed approaches and three well-known code plagiarism detection tools (JPlag, Sherlock, and SIM) were evaluated using 100,000 submissions for 1153 problems from a POJ based on the C programming language. The F1 scores of these approaches were determined as 0.9752 (AST Search), 0.9440 (CFG Search), 0.7405 (Line Statistics), 0.6446 (JPlag), 0.1587 (Sherlock), and 0.0076 (SIM), respectively. The result indicates that (i) AST Search is most suitable for the detection of direct output cheating; (ii) traditional code search or plagiarism detection methods based on similarity calculations are not effective for complex cheat detection because these cheats are highly similar to normal code. Jing Qiu 0003, Chunmei Shi, Yuehua Lv |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2023 | End-to-End TCP Congestion Control as a Classification ProblemabstractThe traditional rule-based congestion control algorithms cannot set congestion window size flexibly, resulting in the inadaptation of the dynamic networks. This article presents a method to model end-to-end TCP congestion control problem using classification techniques. The network status parameters as the input and the type of network status as the output are defined through the analysis of some existing congestion control algorithms. NewReno, CUBIC, and Compound are used as feedback to produce the training data of the XGBoost classifier. The experimental results show that the classifier effectively shapes the strategies of three outstanding congestion control algorithms and almost achieves the same throughput, delay, and fairness. The proposed method makes the congestion control algorithm able to learn from data produced by network. Guanglu Sun, Jing Qiu 0003 |
IEEE Trans. Reliab. | 5 |
| 2022 | Return Instruction Classification in Binary Code Using Machine LearningabstractBinary code analysis is vital in source code unavailable cases, such as malware analysis and software vulnerability mining. Its first step could be function identification. Most function identification methods are based on function prologs/epilogs. However, functions may not have standard prologs/epilogs. To identify these functions, we need to use other methods. One approach is to identify return instructions first and then identify the start of a function. Currently, the multi-layer perceptron model is exploited to identify and validate a return instruction at a specific location. On this basis, a new approach is proposed to improve accuracy and provide more details. Specifically, a return instruction is classified into three classes: (1) false return instruction, (2) true return instruction inner a function but not the last instruction, and (3) true return instruction at the end of a function. The evaluation is performed on 5782 real-world binaries. Meanwhile, common classifiers including fully connected neural network, Two-layer Bidirectional Recurrent Neural Network (TBRNN), Two-layer Bidirectional Gate Recurrent Unit (TBGRU), Two-layer Bidirectional Long Short-term Memory Network (TBLSTM), Decision Tree, Random Forest, XGBoost, and Support Vector Machine (SVM) are evaluated on the same data set. The result shows that TBLSTM achieves an accuracy of 99.78%, which is higher than that of other classifiers in the evaluation, including the state-of-the-art tool IDA Pro 7.7. Jing Qiu 0003, Xiaoxu Geng |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2016 | Optimization and improvements of a Moodle-Based online learning system for C programmingabstractIt is important for students to solve problems with specific requirements in the programming teaching. Our teaching system is a Moodle-based interactive teaching platform for C programming. Its online judging system can grade students code automatically. It plays an extremely important role in programming language teaching. This paper is devoted to optimizing and improving the system. We firstly analyze the five problems in the system according to the feedback from the teachers and students: 1) logical errors in students programs cannot be located; 2) a cheating that directly outputs answers cannot be detected; 3) evaluation results lack statistics and visualization; 4) the code submitting procedure is complicated; 5) the feedback of incorrect answers is not detailed. In order to solve these problems, we employ a fault localization algorithm, revise the evaluation logic, introduce third-party visualization plug-ins and refactor the system, respectively. Detailed and exact solutions are given also. After optimizing and improving the system, the user experience is significantly improved. It is convenient for the student to find and correct errors in their programs. Also, it is easier for teachers to acquire valuable feedback and master students' learning situation. Xiaohong Su, Jing Qiu 0003, Tiantian Wang 0001, Lingling Zhao |
FIE | 2 |
| 2016 | Using Reduced Execution Flow Graph to Identify Library Functions in Binary CodeabstractDiscontinuity and polymorphism of a library function create two challenges for library function identification, which is a key technique in reverse engineering. A new hybrid representation of dependence graph and control flow graph called Execution Flow Graph (EFG) is introduced to describe the semantics of binary code. Library function identification turns to be a subgraph isomorphism testing problem since the EFG of a library function instance is isomorphic to the sub-EFG of this library function. Subgraph isomorphism detection is time-consuming. Thus, we introduce a new representation called Reduced Execution Flow Graph (REFG) based on EFG to speed up the isomorphism testing. We have proved that EFGs are subgraph isomorphic as long as their corresponding REFGs are subgraph isomorphic. The high efficiency of the REFG approach in subgraph isomorphism detection comes from fewer nodes and edges in REFGs and new lossless filters for excluding the unmatched subgraphs before detection. Experimental results show that precisions of both the EFG and REFG approaches are higher than the state-of-the-art tool and the REFG approach sharply decreases the processing time of the EFG approach with consistent precision and recall. Jing Qiu 0003, Xiaohong Su, Peijun Ma |
IEEE Trans. Software Eng. | 1 |
| 2015 | Identifying and Understanding Self-Checksumming Defenses in SoftwareabstractSoftware self-checksumming is widely used as an anti-tampering mechanism for protecting intellectual property and deterring piracy. This makes it important to understand the strengths and weaknesses of various approaches to self-checksumming. This paper describes a dynamic information-flow-based attack that aims to identify and understand self-checksumming behavior in software. Our approach is applicable to a wide class of self chesumming defenses and the information obtained can be used to determine how the checksumming defenses may be bypassed. Experiments using a prototype implementation of our ideas indicate that our approach can successfully identify self-checksumming behavior in (our implementations of) proposals from the research literature. Jing Qiu 0003, Babak Yadegari, Brian Johannesmeyer, Saumya K. Debray, Xiaohong Su |
CODASPY | 1 |
| 2015 | Motivating students with new mechanisms of online assignments and examination to meet the MOOC challenges for programmingabstractThe advent of massive open online courses (MOOC) poses challenges for teaching and learning programming. This paper has analyzed these challenges and thereby proposed a self-motivating learning platform for students in the introductory programming course. Novel mechanisms of online assignments and examination have been introduced. Our platform provides functions for self-motivating learning and practicing in MOOC, which makes it distinguish from the others. For example, self-paced timetable with supervision, self-motivated exercise contents, exercise market, and relative ranking. The automatic grading approach is also a highlight. Programs even with syntactic or semantic errors can be automatically graded. Our platform gains popularity among both students and teachers. The platform has been used together with a programming MOOC. This course is ranked as the third most popular courses among over 500 courses. The platform has also been used by more than 100 other universities. The application of the platform in both MOOC and the traditional classroom courses has shown that students' self-motivation in learning programming has been greatly promoted, and their practical skills have also been significantly improved. Xiaohong Su, Tiantian Wang 0001, Jing Qiu 0003, Lingling Zhao |
FIE | 3 |
| 2015 | Library functions identification in binary code by using graph isomorphism testingsabstractLibrary functions identification is a key technique in reverse engineering. Discontinuity and polymorphism of inline and optimized library functions in binary code create a difficult challenge for library functions identification. To solve this problem, a novel approach is developed to identify library functions. First, we introduce execution dependence graphs (EDGs) to describe the behavior characteristics of binary code. Then, by finding similar EDG subgraphs in target functions, we identify both full and inline library functions. Experimental results from the prototype tool show that the proposed method is not only capable of identifying inline functions but is also more efficient and precise than the current methods for identifying full library functions. Jing Qiu 0003, Xiaohong Su, Peijun Ma |
SANER | 1 |
| 2015 | Identifying functions in binary code with reverse extended control flow graphsabstractAbstract In binary code analysis, current function identification approaches are challenged by functions without explicit call sites and handcrafted assembly without standard prologues/epilogues. We propose a new function representation called a reverse extended control flow graph (RECFG) and a RECFG‐based method for identifying functions in stripped binary code. A function has at least one return instruction (an instruction that makes the control flow leave a function). Therefore, return instructions are more reliable than the function prologues and epilogues used by traditional methods. We first build RECFGs from any values that can be interpreted as return instructions in a code range. Then, for each independent RECFG, the multiple‐decision method chooses a subgraph as the control flow graph of a function. A prototype tool is developed for evaluation on seven open source applications, 138 binaries in MASM32 code examples, and 292 binaries in Windows XP SP3. Experimental results show that the proposed method can identify functions that cannot be identified by current methods with high precision and stable recall. Copyright © 2015 John Wiley & Sons, Ltd. Jing Qiu 0003, Xiaohong Su, Peijun Ma |
J. Softw. Evol. Process. | 1 |