VLDB 2026 Research / reviewers in the wild / expert
Xiaopeng Gao
dblp:51/1201
· DBLP profile ↗
22ranked-venue papers
0as first author
4since 2021 · last 2024
0000-0003-3442-4373ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 9 · 3 since 2021Systems, architecture and hardware · 4Software engineering, systems software and programming languages · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4Artificial intelligence and machine learning · 1Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Peer Grading Approach for Open-ended Programming Projects Based on Binary System and Swiss SystemabstractPeer grading is widely used in high education as effective active learning but still faces challenges. We present the peer grading approach for Open-ended Programming Projects based on the binary and Swiss systems. First, we design a grading specification to improve the accuracy of scoring. Second, to make grading easier for inexperienced students, we utilize a pairwise comparison system based on the binary system. Third, we propose a score calculation algorithm based on Technique for Order Preference by Similarity to an Ideal Solution (TOPSIS) to improve grading accuracy. We developed an online peer review tool called Peer Review Studio (PRS) based on the approach. We carry out the method in the undergraduate programming course of 2023. We collect and analyze the learning data between 2022 and 2023. When measured by Krippendorff's alpha, the inter-rater reliability between instructor and peer grading is in good agreement. When measured by Kruskal-Wallis, students' project performance and learning engagement significantly improve in the first year of peer grading. The course questionnaire 2023 reveals that most students hold a positive attitude toward peer grading and have benefited significantly from this approach. Liang Zhang 0044, Xiaopeng Gao |
SIGCSE (1) | 4 |
| 2023 | A study on the impact of pre-trained model on Just-In-Time defect predictionabstractPrevious researchers conducting Just-In-Time (JIT) defect prediction tasks have primarily focused on the performance of individual pre-trained models, without exploring the relationship between different pre-trained models as backbones. In this study, we build six models: RoBERTaJIT, CodeBERTJIT, BARTJIT, PLBARTJIT, GPT2JIT, and CodeGPTJIT, each with a distinct pre-trained model as its backbone. We systematically explore the differences and connections between these models. Specifically, we investigate the performance of the models when using Commit code and Commit message as inputs, as well as the relationship between training efficiency and model distribution among these six models. Additionally, we conduct an ablation experiment to explore the sensitivity of each model to inputs. Furthermore, we investigate how the models perform in zero-shot and few-shot scenarios. Our findings indicate that each model based on different backbones shows improvements, and when the backbone’s pre-training model is similar, the training resources that need to be consumed are closer. We also observe that Commit code plays a significant role in defect detection, and different pre-trained models demonstrate better defect detection ability with a balanced dataset under few-shot scenarios. These results provide new insights for optimizing JIT defect prediction tasks using pre-trained models and highlight the factors that require more attention when constructing such models. Additionally, CodeGPTJIT and GPT2JIT achieved better performance than DeepJIT and CC2Vec on the two datasets respectively under 2000 training samples. These findings emphasize the effectiveness of transformer-based pre-trained models in JIT defect prediction tasks, especially in scenarios with limited training data. Yuxiang Guo 0004, Xiaopeng Gao, Zhenyu Zhang 0004, Wing Kwong Chan, Bo Jiang 0001 |
QRS | 2 |
| 2022 | Programming Hints Generation based on Abstract Syntax Tree RetrievalabstractThis paper presents research that Works in Progress (WIP). Small private online courses (SPOCs) have recently received extensive attention in computing education. In SPOCs, programming exercises are frequently included to train students’ programming skills. Abstract Syntax Tree Retrieval (ASTR) is a system that can help students solve Python problems by inferring the coding goals. However, the coding goal retrieved by ASTR gives students little information about what to do next. In response to this limitation, this work focuses on generating modification hints for students based on the coding goal. In addition, this paper reports on an effort to translate this idea over to Verilog-HDL programming problems. Without any programmed expert knowledge, the final results demonstrate that our system is generally accurate for 1 out of 2 submissions to give hints at a minimum. And for some favorable problems, it potentially performs much better. Furthermore, the results indicate that in the process of retrieval, weighted tree edit distance calculations resulted in improved accuracy over metric tree edit distance calculations. Han Wan, Hongzhen Luo, Zihao Zhong, Xiaopeng Gao |
FIE | 4 |
| 2021 | Investigating Learners' Behaviors and Implementing Intervention in a SPOCabstractThis Work-In-Progress paper is in the Innovative Practice category. In the MOOC-related research field, many researchers analyzed students' learning behavior based on the logging data to predict students' performance and improve the course design. Nowadays, Small Private Online Courses (SPOC) are favored in college education, especially in computing education. This hybrid teaching model allows courses to be conducted through Internet, which enables teachers and students to access the course anytime, anywhere. Besides, multimedia resources, including images, videos, and audio could be contained in course materials to strengthen the expressiveness of SPOC. On the other hand, the online learning management system (LMS) collects all the students' interactions with it. But how could we extract meaningful information from them? And how could we improve the learning outcomes of a SPOC? In this study, we analyzed LMS data from a sophomore Computer Structure course. We applied several data mining techniques and conducted an intervention using several visualization techniques. Features were selected according to Spearman's rank correlation coefficient with grades. The correlation coefficient of these selected features ranged from 0.42 to 0.84. Course data were further processed to predict students' performance. The predicted grade was processed in the form of heatmaps to illustrate students' learning behavior. Besides, we further designed an overall view for teachers' perspective, which contains data of all the students in each heatmap. The predicting models were evaluated by ROC-AVC values. Several hyperparameters were tuned in order to pursue better predict performance. The best ROC-AVC value could reach 97.44%. Han Wan, Zihao Zhong, Lina Tang, Xiaopeng Gao |
FIE | 4 |
| 2020 | Wasserstein Distance Regularized Sequence Representation for Text Matching in Asymmetrical DomainsabstractOne approach to matching texts from asymmetrical domains is projecting the input sequences into a common semantic space as feature vectors upon which the matching function can be readily defined and learned.In realworld matching practices, it is often observed that with the training goes on, the feature vectors projected from different domains tend to be indistinguishable.The phenomenon, however, is often overlooked in existing matching models.As a result, the feature vectors are constructed without any regularization, which inevitably increases the difficulty of learning the downstream matching functions.In this paper, we propose a novel match method tailored for text matching in asymmetrical domains, called WD-Match.In WD-Match, a Wasserstein distance-based regularizer is defined to regularize the features vectors projected from different domains.As a result, the method enforces the feature projection function to generate vectors such that those correspond to different domains cannot be easily discriminated.The training process of WD-Match amounts to a game that minimizes the matching loss regularized by the Wasserstein distance.WD-Match can be used to improve different text matching methods, by using the method as its underlying matching model.Four popular text matching methods have been exploited in the paper.Experimental results based on four publicly available benchmarks showed that WD-Match consistently outperformed the underlying methods and the baselines. Weijie Yu 0003, Chen Xu 0010, Jun Xu 0001, Liang Pang 0001, Xiaopeng Gao, Xiaozhao Wang, Ji-Rong Wen |
EMNLP (1) | 5 |
| 2020 | Exploring the Relationship between SPOC Forum Behaviors and Learning Outcomes Based on Social Network AnalysisabstractThis paper presents research that works in progress. Small private online courses (SPOCs) have received widespread attention for their adaptability to blended teaching in higher education. As an interactive tool, the SPOC discussion forum generates a large amount of data every day, including learning contents discussion, questions raising and feedback. In this paper, the computer structure course served as the research object, which is a SPOC for sophomores. Social network analysis (SNA) methods were utilized to explore the network extracted from the discussion forum. The results show that learners' three measures of centrality are significantly positively related to learning outcomes, and learners who play different roles in the discussion forum have a significant difference in their final grades. Our results can enable faculties to improve the curriculum and use online learning forums more effectively, such as increasing the number of teacher assistants (TAs) participating in the discussion forum, encouraging students to check the discussion forum regularly and post their learning feelings or questions. Han Wan, Lina Tang, Kangxu Liu, Xiaopeng Gao |
FIE | 4 |
| 2020 | Towards Systems Education for Artificial Intelligence: A Course Practice in Intelligent Computing ArchitecturesabstractWith the rapid development of artificial intelligence (AI) community, education in AI is receiving more and more attentions. There have been many AI related courses in the respects of algorithms and applications, while not many courses in system level are seriously taken into considerations. In order to bridge the gap between AI and computing systems, we are trying to explore how to conduct AI education from the perspective of computing systems. In this paper, a course practice in intelligent computing architectures are provided to demonstrate the system education in AI era. The motivation for this course practice is first introduced as well as the learning orientations. The main goal of this course aims to teach students for designing AI accelerators on FPGA platforms. The elaborated course contents include lecture notes and related technical materials. Especially several practical labs and projects are detailed illustrated. Finally, some teaching experiences and effects are discussed as well as some potential improvements in the future. Jianlei Yang 0001, Xiaopeng Gao, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2019 | A Web-based Remote FPGA Laboratory for Computer Organization CourseabstractLearning in digital systems could be enhanced by applying a learn-by-doing mechanism. In this paper the implementation of a web-based remote FPGA laboratory for Computer Organization course is proposed. The projects created for this course are designed towards entry-level users thus producing a trouble-free virtualization of experimental equipment that in turn maximizes educational gain. The network laboratory had made learners complete all the experimental activities through the Internet without the limitation of time, space and resources. We show also some preliminary results compared between Fall 2016 and Fall 2017 semester, which highlight improved student engagement, learning and achievement. Han Wan, Kangxu Liu, Jiazhen Lin, Xiaopeng Gao |
ACM Great Lakes Symposium on VLSI | 4 |
| 2018 | Improving Blended Learning Outcomes Through Academic Social MediaabstractSocial media had played an increasingly crucial role in young lives. In this paper, the most popular mobile social media application in China - WeChat has been used to develop an academic social media platform in order to improve students' blended learning outcomes of a SPOC (Small Private Online Course). We divided our workflow into four major parts: (a) pushing course updates directly to students; (b) reminding inactive learners during the self-paced tutorials learning phase; (c) promoting students to participate in discussion forum; (d) providing self-service inquiry for course progress and seating chart. Split-tests between two groups (each group has 219 on-campus students) were conducted in the Fall 2017 semester. We analyzed the behavior statistics and the long-term influence between the controlled group and the experimental group. The results showed that the social tool has a positive impact in the promotion of on-campus students' learning. Our work shows the value of leaving a door open for SPOC researchers, properly identifying participants who are at-risk and developing personalized recommending systems could help in the improving learning outcomes. Han Wan, Kangxu Liu, Qiaoye Yu, Xiaopeng Gao |
COMPSAC (1) | 5 |
| 2018 | Token-based Approach for Real-time Plagiarism Detection in Digital DesignsabstractThis Research to Practice Work in Progress Paper presents a token-based approach to detecting plagiarism in university courses with hardware programming assignments. Detecting plagiarism manually is a difficult and time-consuming work. In the last two decades, various of plagiarism detection tools have been developed. These techniques could be mainly divided into the following categories: Textual Match, Program Dependence Graph Comparison, Abstract Syntax Tree Analysis and Low-Level Form Code Comparison. Although there had been a lot of researches on detecting code clones in software programming languages (e.g. Basic, C/C++, Java, Python, etc.), research that focused on hardware description languages is still lacking. Based on the effective of the locality sensitive hash function (simhash), which was usually used in detecting near duplicates for web crawling, we proposed an improved real-time plagiarism detection approach for Verilog HDL (hardware description language) programming assignments. The core detecting steps are extracting weighted tokens from source code as high-dimensional feature, and mapping it to a f-bit fingerprints with simhash technique. On account of the syntax characteristics of Verilog HDL, a token extraction strategy was designed to maximize the valid information that a fixed length hash value could represent. Experiments over real course data sets were conducted to evaluate the performance of token-based approach comparing with an existing plagiarism detection tool (Moss). The result shows that our token-based approach does qualify the plagiarism detecting job for both online-query and batch-query in digital designs. Furthermore, token-based plagiarism detection approach could enable conduct incremental plagiarism detection for a single submission without excessive overhead. Finally, we also give a discussion of current way limitations and future research directions. Han Wan, Kangxu Liu, Xiaopeng Gao |
FIE | 3 |
| 2017 | Dropout Prediction in MOOCs using Learners' Study Habits Features
Han Wan, Xiaopeng Gao, David E. Pritchard |
EDM | 3 |
| 2017 | Predicting Performance in a Small Private Online Course
Han Wan, Xiaopeng Gao, Qiaoye Yu, Kangxu Liu |
EDM | 3 |
| 2017 | VREX: Virtual reality education expansion could help to improve the class experience (VREX platform and community for VR based education)abstractThis paper proposed an innovative education platform-VREX (Virtual Reality based Education eXpansion), with combination of online and offline, to improve the curriculum building and teaching experience. VREX is based on Virtual Reality (VR) and we believe VR can revolutionize the education ecosystem. With some trials, we found VR can be used to promote curriculum effectiveness in an immersive environment so that students can have intuitive sense to understand some abstract knowledge, which is always hard for teachers to describe. We have tried to transfer slides into VR scenes, for the students to learn knowledge in a rather real but totally virtual world. The main contributions were made: (1) VREX build an open and immersion virtual O2O classroom with internet and VR devices so that real classrooms might be used in a different way in the future. (2) VREX provides a distributed mode for students to experience an interactive learning process at anytime, anywhere and any-frequency. (3) VREX can be used to support education in different disciplines, from K-12 to Universities, and we provided some practical cases, like `Marine Life' to show creatures in deep sea, which provides immersive experience to makes students feel they were there. Finally, the feasibility and advantage of VREX are proved by the actual statistical data in the 3rdseason 2017. Jingchun Wang, Xiaopeng Gao |
FIE | 5 |
| 2017 | Supporting quality teaching using educational data mining based on OpenEdX platformabstractOur lab-based small private online course (SPOC) combined online resources and technology with engagement between faculty and students based on OpenEdX platform. It worked with an auto-grading submission system which could reduce the instructors' burden of evaluation and provide better learners' experience. Different study behaviors were observed from the system tracking logs. Identifying at-risk students becomes timely important in SPOC, and the early prediction can help instructors provide proper supports. In this paper, we focused on extracting features from students' learning activities and study habits for building machine learning models to predict students' performance. We conducted experiments to compare feature importance, and the results showed that study habits related features had played more important role in predicting students' performance. 34 predictive features extracted from Computer Structure Course in Fall 2016, and our model achieved an ROC (Receiver Operating Characteristic Curve)-AUC (area under the curve) in the range of 0.927-0.984 when predicting the performance. Our evaluation showed that data mining is useful in education especially when examining students' learning behavior in online environment, and could support quality teaching. In the next course iteration, we will do A/B testing to determine efficacy for subsequent interventions in a SPOC. Han Wan, Xiaopeng Gao, Kangxu Liu |
FIE | 3 |
| 2017 | Introducing parallel computing concepts in computer system related coursesabstractAll semiconductor market domains are converging to concurrent platforms. This trend has certainly led real challenge to develop applications software that effectively uses these concurrent processors to achieve efficiency and performance goals. This paper argues that the Computer System related courses are natural places to introduce the parallelism, and the earlier to parallel computing concepts will have a wide reach. This paper showed how digital logic classes can motivate topics from parallel computing through common logic structures. We provided an alternative view of the digital logic topics, including: binary representation of integers using decision tree with recursive thinking; use a carry look-ahead adder to show how sequential operations can be parallelized. Another part to introduce parallel concepts focused on write high performance code with specific emphasis on graphic processing unit (GPU). In our teaching experience, parallel pattern teaching has been confirmed to be a useful pedagogical method for teaching parallel concepts. Finally, we report course experience in teaching parallel computing injected course, which resulted in positive student feedback. Han Wan, Xiaopeng Gao, Xiang Long, Bo Jiang 0001 |
FIE | 2 |
| 2016 | Hybrid teaching mode for laboratory-based remote education of computer structure courseabstractWe describe an Open edX-based blended course developed for a reformed Computer Structure course at Beihang University. In three iterations of this laboratory-based course, we dive into key issues that impact students' learning, and then redesign our curriculum, which integrated with virtual laboratory technique into the MOOC platform. We show how certain course design aspects affect students' learning in the hybrid teaching mode: (a) Kung Fu style competency education with online-support laboratory system prompts students to own their learning as the pace and/or the path of learning, which is dictated by mastery instead of the time/space; (b) strengthen the use of learning-aid tools empowered teachers with the skills and information to define standard for each learner in each stage; (c) automated test technology make this blended learning possible at scale and also financially sustainable; (d) using discussion forum to build the lesson about `what to do' when learners get stuck helped in overcoming challenges. Han Wan, Xiaopeng Gao |
FIE | 2 |
| 2012 | Using Basic Block Based Instruction Prefetching to Optimize WCET Analysis for Real-Time ApplicationsabstractCache is an important component existing in modern computer system to bridge the performance gap between the fast CPU and the slow memory system. A variety of cache optimization technologies and mechanisms are proposed to improve the cache performance, such as instruction cache prefetching. Most instruction prefetching mechanisms existing are proposed to improve the average-case cache performance. However, real-time systems care more about the worst-case performance, and the worst-case execution time (WCET) analysis of real-time applications is critical for schedulability analysis of real-time systems. Due to its unpredictable behaviour, cache disastrously complicates the WCET analysis of real-time applications. In this paper, we proposed a basic block based instruction prefetching (BBIP) mechanism to improve both the average-case cache performance and the tightness of the WCET analysis of real-time applications. Measurements on typical real-time benchmarks show that BBIP can not only eliminate most of the instruction access misses, but also result in lower WCET estimations. To discuss the effectiveness of BBIP, we measured the WCET of the benchmarks for three processor configurations with and without BBIP: 1) processor with in-order pipeline and perfect branch prediction, 2) processor with out-of-order pipeline and perfect branch prediction, and 3) processor with out-of-order pipeline and 2-level branch prediction. The results show that BBIP can provide notable improvements in the tightness of WCET estimation, with the WCET values being 30.4% to 97.7% of the original ones. Our simulation results also reveal that 70% to 80% instruction access misses are eliminated with BBIP. Fan Ni, Xiang Long, Han Wan, Xiaopeng Gao |
PDCAT | 4 |
| 2012 | Complex networks properties analysis for mobile ad hoc networksabstractRecently, research on complex network theory and applications draws a lot of attention in both academy and industry. In mobile ad hoc networks (MANETs) area of research, a critical issue is to design the most effective topology for given problems. It is natural and significant to consider complex networks topology when optimising the MANET topology. Current works usually transform MANET or sensor network topologies into either small-world or scale-free. However, some fundamental problems remain unsolved. Specifically, what are the average shortest path length, degree distribution and clustering characteristics of MANETs? Do MANETs have small-world effect and scale-free property? In this work, the authors introduce complex networks theory into the context of MANET topology and study complex network properties of the MANETs to answer the above questions. The authors have theoretically analysed the degree distribution and clustering coefficient of MANETs and proposed approach to computing them. The degree distribution and clustering coefficient of MANETs are theoretically deduced from node space probability distribution on different mobility models (including but not limited to random waypoint model). Simulation results on average shortest path length, clustering coefficient and degree distribution show that in most cases MANETs do not have the small-world effect and scale-free property. Chao Tong 0001, Jianwei Niu 0002, Guangzhi Qu, Xiang Long, Xiaopeng Gao |
IET Commun. | 5 |
| 2010 | Testing in Parallel - A Need for Practical Regression Testing
Zhenyu Zhang 0004, Zijian Tong, Xiaopeng Gao |
ICSOFT (2) | 3 |
| 2009 | GCSim: A GPU-Based Trace-Driven Simulator for Multi-level Cache
Han Wan, Xiaopeng Gao, Xiang Long |
APPT | 2 |
| 2009 | Using GPU to Accelerate Cache SimulationabstractCaches play a major role in the performance of high-speed computer systems. Trace-driven simulator is the most widely used method to evaluate cache architectures. However, as the cache design moves to more complicated architectures, along with the size of the trace is becoming larger and larger. Traditional simulation methods are no longer practical due to their long simulation cycles. Several techniques have been proposed to reduce the simulation time of sequential trace-driven simulation. This paper considers the use of generic GPU to accelerate cache simulation which exploits set-partitioning as the main source of parallelism. We develop more efficient parallel simulation techniques by introducing more knowledge into the Compute Unified Device Architecture (CUDA) on the GPU. Our experimental result shows that the new algorithm gains 2.76x performance improvement compared to traditional CPU-based sequential algorithm. Han Wan, Xiaopeng Gao |
ISPA | 2 |
| 2007 | SQS: A Secure and QoS Guaranteed Solution for Mobile ServiceabstractMobile IP protocol, a standard proposed by the Internet Engineering Task Force, was designed to support IP mobility. On the basis of analysis of Mobile IP protocol, it can be concluded that some problems (e.g., security flaws, limited agents deployments, triangular routing and just a little supports by operating systems, etc.) remain to be solved. In this paper, we present a novel mobility supporting scheme based on embedded technology which is called Secure and QoS Guaranteed Solution for Mobile Service (SQS). It implements mobility management for handling seamless handoffs with Embedded Mobile Agent (EmMA) and Mobility Management Server (MMS), which allows nodes to continue to receive datagrams when the nodes change their points of attachment to the Internet. It reduces the dependence on network infrastructure and provides nodes with transparent mobile service. Experimental results show that SQS effectively solves the problems mentioned above and improves Mobile IP protocol. Xiaopeng Gao, Xiang Long |
COMPSAC (1) | 2 |