Han Wan

dblp:75/5153 · DBLP profile ↗
← Back
29ranked-venue papers
18as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 10 · 8 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 first-author · 4 since 2021Systems, architecture and hardware · 6 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PIMRL: Physics-Informed Multi-Scale Recurrent Learning for Burst-Sampled Spatiotemporal Dynamics
abstract
Deep learning has shown strong potential in modeling complex spatiotemporal dynamics. However, most existing methods depend on densely and uniformly sampled data, which is often unavailable in practice due to sensor and cost limitations. In many real-world settings, such as mobile sensing and physical experiments, data are burst-sampled with short high-frequency segments followed by long gaps, making it difficult to learn accurate dynamics from sparse observations. To address this issue, we propose Physics-Informed Multi-Scale Recurrent Learning (PIMRL), a novel framework specifically designed for burst-sampled spatiotemporal data. PIMRL combines macro-scale latent dynamics inference with micro-scale adaptive refinement guided by incomplete prior information from partial differential equations (PDEs). It further introduces a temporal message-passing mechanism to effectively propagate information across burst intervals. This multi-scale architecture enables PIMRL to model complex systems accurately even under severe data scarcity. We evaluate our approach on five benchmark datasets involving 1D to 3D multi-scale PDEs. The results show that PIMRL consistently outperforms state-of-the-art baselines, achieving substantial improvements and reducing errors by up to 80\% in the most challenging settings, which demonstrates the clear advantage of our model. Our work demonstrates the effectiveness of physics-informed recurrent learning for accurate and efficient modeling of sparse spatiotemporal systems.
Han Wan, Qi Wang 0123, Yuan Mi, Rui Zhang 0052, Hao Sun 0002
AAAI1
2026 L2V-CoT: Cross-Modal Transfer of Chain-of-Thought Reasoning via Latent Intervention
abstract
Recently, Chain-of-Thought (CoT) reasoning has significantly enhanced the capabilities of large language models (LLMs), but Vision–Language Models (VLMs) still struggle with multi-step reasoning tasks due to limited multimodal reasoning data. To bridge this gap, researchers have explored methods to transfer CoT reasoning from LLMs to VLMs. However, existing approaches either need high training costs or require architectural alignment. In this paper, we use Linear Artificial Tomography (LAT) to empirically show that LLMs and VLMs share similar low-frequency latent representations of CoT reasoning despite architectural differences. Based on this insight, we propose L2V-CoT, a novel training-free latent intervention approach that transfers CoT reasoning from LLMs to VLMs. L2V-CoT extracts and resamples low-frequency CoT representations from LLMs in the frequency domain, enabling dimension matching and latent injection into VLMs during inference to enhance reasoning capabilities. Extensive experiments demonstrate that our approach consistently outperforms training-free baselines and even surpasses supervised methods.
Yuliang Zhan, Xinyu Tang 0004, Han Wan, Jian Li 0064, Ji-Rong Wen, Hao Sun 0002
AAAI3
2026 Generative spatial downscaling of global ocean wind speed profiles via a diffusion model
Anyuan Xiong, Lijuan Cao, Lifan Chen, Rui Zhang 0052, Qi Wang 0123, Zhihong Liao, Han Wan, Bocheng Zeng, Chongxuan Li, Hao Sun 0002
Neurocomputing8
2026 TinyFormer: Efficient Sparse Transformer Design and Deployment on Tiny Devices
abstract
Developing deep learning models on tiny devices (e.g. Microcontroller units, MCUs) has attracted much attention in various embedded IoT applications. However, it is challenging to efficiently design and deploy recent advanced models (e.g. transformers) on tiny devices due to their severe hardware resource constraints. In this work, we proposeTinyFormer, a framework specifically designed to develop and deploy resource-efficient transformer models on MCUs. TinyFormer consists ofSuperNAS,SparseNAS, andSparseEngine. Separately, SuperNAS aims to search for an appropriate supernet from a vast search space. SparseNAS evaluates the best sparse single-path transformer model from the identified supernet. Finally, SparseEngine efficiently deploys the searched sparse models onto MCUs. To the best of our knowledge, SparseEngine is the first deployment framework capable of performing inference of sparse transformer models on MCUs. Evaluation results on the CIFAR-10 dataset demonstrate that TinyFormer can design efficient transformers with an accuracy of 96.1% while adhering to hardware constraints of 1MB storage and 320KB memory. Additionally, TinyFormer achieves significant speedups in sparse inference, up to$12.2\times $comparing to the CMSIS-NN library. TinyFormer is believed to bring powerful transformers into TinyML scenarios and to greatly expand the scope of deep learning applications.
Jianlei Yang 0001, Jiacheng Liao, Fanding Lei, Meichen Liu, Lingkun Long, Han Wan, Bei Yu 0001, Weisheng Zhao 0001
IEEE Trans. Circuits Syst. I Regul. Pap.7
2025 Towards Affordable, Adaptive and Automatic GNN Training on CPU-GPU Heterogeneous Platforms
abstract
Graph Neural Networks (GNNs) have been widely adopted due to their strong performance. However, GNN training often relies on expensive, high-performance computing platforms, limiting accessibility for many tasks. Profiling of representative GNN workloads indicates that substantial efficiency gains are possible on resource-constrained devices by fully exploiting available resources. This paper introduces$\mathrm{A}^{3} \text{GNN}$, a framework for Affordable, Adaptive, and Automatic GNN training on heterogeneous CPU-GPU platforms. It improves resource usage through locality-aware sampling and fine-grained parallelism scheduling. Moreover, it leverages reinforcement learning to explore the design space and achieve pareto-optimal trade-offs among throughput, memory footprint, and accuracy. Experiments show that$\mathrm{A}^{3}$GNN can bridge the performance gap, allowing seven Nvidia 2080Ti GPUs to outperform two A100 GPUs by up to$1.8 \times$in throughput with minimal accuracy loss.
Yingjie Qi, Yiou Wang, Han Wan, Jianlei Yang 0001, Chunming Hu
ICCD5
2025 PeSANet: Physics-encoded Spectral Attention Network for Simulating PDE-Governed Complex Systems
abstract
Accurately modeling and forecasting complex systems governed by partial differential equations (PDEs) is crucial in various scientific and engineering domains. However, traditional numerical methods struggle in real-world scenarios due to incomplete or unknown physical laws. Meanwhile, machine learning approaches often fail to generalize effectively when faced with scarce observational data and the challenge of capturing local and global features. To this end, we propose the Physics-encoded Spectral Attention Network (PeSANet), which integrates local and global information to forecast complex systems with limited data and incomplete physical priors. The model consists of two key components: a physics-encoded block that uses hard constraints to approximate local differential operators from limited data, and a spectral-enhanced block that captures long-range global dependencies in the frequency domain. Specifically, we introduce a novel spectral attention mechanism to model inter-spectrum relationships and learn long-range spatial features. Experimental results demonstrate that PeSANet outperforms existing methods across all metrics, particularly in long-term forecasting accuracy, providing a promising solution for simulating complex systems with limited data and incomplete physics.
Han Wan, Rui Zhang 0052, Qi Wang 0123, Yang Aron Liu, Hao Sun 0002
IJCAI1
2025 SlotPi: Physics-informed Object-centric Reasoning Models
abstract
Understanding and reasoning about dynamics governed by physical laws through visual observation, akin to human capabilities in the real world, poses significant challenges. Currently, object-centric dynamic simulation methods, which emulate human behavior, have achieved notable progress but overlook two critical aspects: 1) the integration of physical knowledge into models. Humans gain physical insights by observing the world and apply this knowledge to accurately reason about various dynamic scenarios; 2) the validation of model adaptability across diverse scenarios. Real-world dynamics, especially those involving fluids and objects, demand models that not only capture object interactions but also simulate fluid flow characteristics. To address these gaps, we introduce SlotPi, a slot-based physics-informed object-centric reasoning model. SlotPi integrates a physical module based on Hamiltonian principles with a spatio-temporal prediction module for dynamic forecasting. Our experiments highlight the model's strengths in tasks such as prediction and Visual Question Answering (VQA) on benchmark and fluid datasets. Furthermore, we have created a real-world dataset encompassing object interactions, fluid dynamics, and fluid-object interactions, on which we validated our model's capabilities. The model's robust performance across all datasets underscores its strong adaptability, laying a foundation for developing more advanced world models.
Jian Li 0064, Han Wan, Ning Lin, Yuliang Zhan, Ruizhi Chengze, Yi Zhang 0164, Hongsheng Liu 0002, Zidong Wang 0010, Fan Yu 0004, Hao Sun 0002
KDD (2)2
2024 Fault Localization for Novice Programs Combining Static Analysis and Dynamic Detection
abstract
In programming teaching, teachers or teaching assistants often need to spend a lot of energy helping students solve the problems they face when doing programming. It will be helpful to provide students with valuable programming feedback, such as information on faulty lines. However, most existing algorithms do not perform well on novice programs. Therefore, considering the background in programming teaching, we proposed a novel approach combining static analysis with dynamic detection by using the correct programs submitted by previous students and coverage information for the incorrect program. In particular, the core of the static analysis module is to locate specific faulty lines through syntax tree difference comparison, which includes matching similar programs, variable mapping and replacement, and fault localization based on abstract syntax tree differences. The core of the module on dynamic detection is to perform traditional Spectrum-based Fault Localization. To evaluate the effectiveness of our proposed approach, we conducted some empirical studies on 223 student-failure programs in the real world. The experimental results indicate that our approach outperforms other baselines regarding TOP-l and TOP-3. Furthermore, we analyzed the performance of our method on different categories of programming problems as well as the effectiveness of the combination of static analysis and dynamic detection.
Han Wan, Wenhao Nie, Shiyang Yue, Xiaoyan Luo
COMPSAC1
2023 Early Prediction of Student Performance with LSTM-Based Deep Neural Network
Han Wan, Mengying Li, Zihao Zhong, Xiaoyan Luo
COMPSAC1
2023 Learning Path Recommendation Based on Knowledge Tracing and Reinforcement Learning
abstract
Adaptive and intelligent web-based educational systems are made to automate the adaptation of the system to the learners' behaviors and needs. Personalized e-learning platforms should make adaptive adjustments according to the individual students' interactions and their knowledge states (KS). This study proposes a more effective personalized learning path recommendation algorithm to promote the individualized development of students. First, the Dynamic Key-Value Memory Network (DKVMN) is enhanced by integrating a learning behavior module, which is used to trace student knowledge states. Then, the proposed knowledge tracing model is used to simulate virtual students and train recommendation policy based on reinforcement learning (RL). The experimental results show that our personalized learning path recommendation algorithm increases the average knowledge state of students by 12.11% and 5.38% on two different data sets, respectively.
Han Wan, Baoliang Che, Hongzhen Luo, Xiaoyan Luo
ICALT1
2023 RSCanner: rapid assessment and visualization of RNA structure content
abstract
MOTIVATION: The increasing availability of RNA structural information that spans many kilobases of transcript sequence imposes a need for tools that can rapidly screen, identify, and prioritize structural modules of interest. RESULTS: We describe RNA Structural Content Scanner (RSCanner), an automated tool that scans RNA transcripts for regions that contain high levels of secondary structure and then classifies each region for its relative propensity to adopt stable or dynamic structures. RSCanner then generates an intuitive heatmap enabling users to rapidly pinpoint regions likely to contain a high or low density of discrete RNA structures, thereby informing downstream functional or structural investigation. AVAILABILITY AND IMPLEMENTATION: RSCanner is freely available as both R script and R Markdown files, along with full documentation and test data (https://github.com/pylelab/RSCanner).
Gandhar Mahadeshwar, Rafael de Cesaris Araujo Tavares, Han Wan, Zion R. Perry, Anna Marie Pyle
Bioinform.3
2022 Programming Hints Generation based on Abstract Syntax Tree Retrieval
abstract
This paper presents research that Works in Progress (WIP). Small private online courses (SPOCs) have recently received extensive attention in computing education. In SPOCs, programming exercises are frequently included to train students’ programming skills. Abstract Syntax Tree Retrieval (ASTR) is a system that can help students solve Python problems by inferring the coding goals. However, the coding goal retrieved by ASTR gives students little information about what to do next. In response to this limitation, this work focuses on generating modification hints for students based on the coding goal. In addition, this paper reports on an effort to translate this idea over to Verilog-HDL programming problems. Without any programmed expert knowledge, the final results demonstrate that our system is generally accurate for 1 out of 2 submissions to give hints at a minimum. And for some favorable problems, it potentially performs much better. Furthermore, the results indicate that in the process of retrieval, weighted tree edit distance calculations resulted in improved accuracy over metric tree edit distance calculations.
Han Wan, Hongzhen Luo, Zihao Zhong, Xiaopeng Gao
FIE1
2021 Investigating Learners' Behaviors and Implementing Intervention in a SPOC
abstract
This Work-In-Progress paper is in the Innovative Practice category. In the MOOC-related research field, many researchers analyzed students' learning behavior based on the logging data to predict students' performance and improve the course design. Nowadays, Small Private Online Courses (SPOC) are favored in college education, especially in computing education. This hybrid teaching model allows courses to be conducted through Internet, which enables teachers and students to access the course anytime, anywhere. Besides, multimedia resources, including images, videos, and audio could be contained in course materials to strengthen the expressiveness of SPOC. On the other hand, the online learning management system (LMS) collects all the students' interactions with it. But how could we extract meaningful information from them? And how could we improve the learning outcomes of a SPOC? In this study, we analyzed LMS data from a sophomore Computer Structure course. We applied several data mining techniques and conducted an intervention using several visualization techniques. Features were selected according to Spearman's rank correlation coefficient with grades. The correlation coefficient of these selected features ranged from 0.42 to 0.84. Course data were further processed to predict students' performance. The predicted grade was processed in the form of heatmaps to illustrate students' learning behavior. Besides, we further designed an overall view for teachers' perspective, which contains data of all the students in each heatmap. The predicting models were evaluated by ROC-AVC values. Several hyperparameters were tuned in order to pursue better predict performance. The best ROC-AVC value could reach 97.44%.
Han Wan, Zihao Zhong, Lina Tang, Xiaopeng Gao
FIE1
2020 Phone Keypad Voice Recognition (PKVR): An Integrated Experiment for Digital Signal Processing Education
abstract
This Innovative Practice Work-In-Progress presents an integrated signal processing experiment, which can cover most knowledge points of digital signal processing (DSP) course. Since the DSP course focuses on one dimension signal processing, voice signal has a great advantage. We provide an integrated voice signal processing experiment named as Phone Keypad Voice Recognition (PKVR), including the following parts: phone keypad voice collection, Discrete Fourier Transform (DFT) and analysis, filter design, digital query table establishment, number recognition of any keypad voice. Through the improvement of the practice training, the classroom teaching theory can be better understood in an interesting way for our students.
Xiaoyan Luo, Han Wan, Fugen Zhou
FIE3
2020 Exploring the Relationship between SPOC Forum Behaviors and Learning Outcomes Based on Social Network Analysis
abstract
This paper presents research that works in progress. Small private online courses (SPOCs) have received widespread attention for their adaptability to blended teaching in higher education. As an interactive tool, the SPOC discussion forum generates a large amount of data every day, including learning contents discussion, questions raising and feedback. In this paper, the computer structure course served as the research object, which is a SPOC for sophomores. Social network analysis (SNA) methods were utilized to explore the network extracted from the discussion forum. The results show that learners' three measures of centrality are significantly positively related to learning outcomes, and learners who play different roles in the discussion forum have a significant difference in their final grades. Our results can enable faculties to improve the curriculum and use online learning forums more effectively, such as increasing the number of teacher assistants (TAs) participating in the discussion forum, encouraging students to check the discussion forum regularly and post their learning feelings or questions.
Han Wan, Lina Tang, Kangxu Liu, Xiaopeng Gao
FIE1
2020 A Comprehensive Experiment to Enhance Multidisciplinary Engineering Ability via UAVs Visual Navigation
abstract
This Research to Practice WIP presents a UAVs visual navigation based comprehensive experiment to enhance multidisciplinary engineering ability in Aerospace engineering education. In traditional courses, aerospace-related disciplines are independently distributed in different courses, and there is rarely a hands-on platform which includes signal processing, control theory, and artificial intelligence into Aerospace engineering. Facing this problem, this paper designs a multidisciplinary comprehensive experiment, aiming to provide a hand-on platform and flexible project-based program to students of aerospace engineering professions. First of all, in order to let the students understand actual aerospace problems, a multidisciplinary simulation platform containing UAVs and remote objects scenarios is constructed for them to explore in the experiments. Second, the content of the experiment is designed into three stages including data acquisition and processing, conceptual design and simulation, in-flight validation, during which the multidisciplinary engineering ability runs through the whole process of the activities. Finally, Project Oriented Design Based Learning is also introduced here to combine engineering design education with innovation and creativity. Through the project demonstration and presentation at the end of the experiment, the multidisciplinary engineering ability of each student can be effectively evaluated. The UVN comprehensive experiment enables students to work on real-world aerospace engineering problems through a hardware-software integration framework, which may greatly stimulate their curiosity and interest in autonomously learning. It also provides students unprecedented opportunities to immerse themselves in projects that cross disciplinary boundaries, improve their professional ability and enhance their exploration competence in aerospace areas.
Xiaoyan Luo, Han Wan, Chengxi Wu, Yu Zheng 0017, Fugen Zhou
FIE3
2020 Decomposition-Based Real-Time Scheduling of Parallel Tasks on Multicores Platforms
abstract
Multicore processors have become mainstream computation platforms not only for general and high-performance computers but also for real-time embedded systems. To fully utilize the computation power of multicores, software must be parallelized. Recently, there has been a rapidly increasing interest in real-time scheduling of parallel real-time tasks, but the field is still much less mature than traditional real-time scheduling of sequential tasks. In this article, we study the real-time scheduling and techniques for parallel real-time tasks based on decomposition, where a task graph is transferred to a set of independent sporadic tasks. In particular, we propose new decomposition strategies that better explore the structure feature of each task to improve schedulability. We develop schedulability tests for the global earliest deadline first (EDF) scheduling algorithm based on decomposition and three types of its variants, with their own pros and cons in different aspects. We conduct experiments to evaluate the real-time performance of our proposed scheduling algorithms against the state-of-the-art scheduling and analysis methods of different types.
Xu Jiang 0004, Nan Guan, Xiang Long, Han Wan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2019 A Web-based Remote FPGA Laboratory for Computer Organization Course
abstract
Learning in digital systems could be enhanced by applying a learn-by-doing mechanism. In this paper the implementation of a web-based remote FPGA laboratory for Computer Organization course is proposed. The projects created for this course are designed towards entry-level users thus producing a trouble-free virtualization of experimental equipment that in turn maximizes educational gain. The network laboratory had made learners complete all the experimental activities through the Internet without the limitation of time, space and resources. We show also some preliminary results compared between Fall 2016 and Fall 2017 semester, which highlight improved student engagement, learning and achievement.
Han Wan, Kangxu Liu, Jiazhen Lin, Xiaopeng Gao
ACM Great Lakes Symposium on VLSI1
2018 Improving Blended Learning Outcomes Through Academic Social Media
abstract
Social media had played an increasingly crucial role in young lives. In this paper, the most popular mobile social media application in China - WeChat has been used to develop an academic social media platform in order to improve students' blended learning outcomes of a SPOC (Small Private Online Course). We divided our workflow into four major parts: (a) pushing course updates directly to students; (b) reminding inactive learners during the self-paced tutorials learning phase; (c) promoting students to participate in discussion forum; (d) providing self-service inquiry for course progress and seating chart. Split-tests between two groups (each group has 219 on-campus students) were conducted in the Fall 2017 semester. We analyzed the behavior statistics and the long-term influence between the controlled group and the experimental group. The results showed that the social tool has a positive impact in the promotion of on-campus students' learning. Our work shows the value of leaving a door open for SPOC researchers, properly identifying participants who are at-risk and developing personalized recommending systems could help in the improving learning outcomes.
Han Wan, Kangxu Liu, Qiaoye Yu, Xiaopeng Gao
COMPSAC (1)1
2018 Token-based Approach for Real-time Plagiarism Detection in Digital Designs
abstract
This Research to Practice Work in Progress Paper presents a token-based approach to detecting plagiarism in university courses with hardware programming assignments. Detecting plagiarism manually is a difficult and time-consuming work. In the last two decades, various of plagiarism detection tools have been developed. These techniques could be mainly divided into the following categories: Textual Match, Program Dependence Graph Comparison, Abstract Syntax Tree Analysis and Low-Level Form Code Comparison. Although there had been a lot of researches on detecting code clones in software programming languages (e.g. Basic, C/C++, Java, Python, etc.), research that focused on hardware description languages is still lacking. Based on the effective of the locality sensitive hash function (simhash), which was usually used in detecting near duplicates for web crawling, we proposed an improved real-time plagiarism detection approach for Verilog HDL (hardware description language) programming assignments. The core detecting steps are extracting weighted tokens from source code as high-dimensional feature, and mapping it to a f-bit fingerprints with simhash technique. On account of the syntax characteristics of Verilog HDL, a token extraction strategy was designed to maximize the valid information that a fixed length hash value could represent. Experiments over real course data sets were conducted to evaluate the performance of token-based approach comparing with an existing plagiarism detection tool (Moss). The result shows that our token-based approach does qualify the plagiarism detecting job for both online-query and batch-query in digital designs. Furthermore, token-based plagiarism detection approach could enable conduct incremental plagiarism detection for a single submission without excessive overhead. Finally, we also give a discussion of current way limitations and future research directions.
Han Wan, Kangxu Liu, Xiaopeng Gao
FIE1
2017 Dropout Prediction in MOOCs using Learners' Study Habits Features
Han Wan, Xiaopeng Gao, David E. Pritchard
EDM1
2017 Predicting Performance in a Small Private Online Course
Han Wan, Xiaopeng Gao, Qiaoye Yu, Kangxu Liu
EDM1
2017 Supporting quality teaching using educational data mining based on OpenEdX platform
abstract
Our lab-based small private online course (SPOC) combined online resources and technology with engagement between faculty and students based on OpenEdX platform. It worked with an auto-grading submission system which could reduce the instructors' burden of evaluation and provide better learners' experience. Different study behaviors were observed from the system tracking logs. Identifying at-risk students becomes timely important in SPOC, and the early prediction can help instructors provide proper supports. In this paper, we focused on extracting features from students' learning activities and study habits for building machine learning models to predict students' performance. We conducted experiments to compare feature importance, and the results showed that study habits related features had played more important role in predicting students' performance. 34 predictive features extracted from Computer Structure Course in Fall 2016, and our model achieved an ROC (Receiver Operating Characteristic Curve)-AUC (area under the curve) in the range of 0.927-0.984 when predicting the performance. Our evaluation showed that data mining is useful in education especially when examining students' learning behavior in online environment, and could support quality teaching. In the next course iteration, we will do A/B testing to determine efficacy for subsequent interventions in a SPOC.
Han Wan, Xiaopeng Gao, Kangxu Liu
FIE1
2017 Introducing parallel computing concepts in computer system related courses
abstract
All semiconductor market domains are converging to concurrent platforms. This trend has certainly led real challenge to develop applications software that effectively uses these concurrent processors to achieve efficiency and performance goals. This paper argues that the Computer System related courses are natural places to introduce the parallelism, and the earlier to parallel computing concepts will have a wide reach. This paper showed how digital logic classes can motivate topics from parallel computing through common logic structures. We provided an alternative view of the digital logic topics, including: binary representation of integers using decision tree with recursive thinking; use a carry look-ahead adder to show how sequential operations can be parallelized. Another part to introduce parallel concepts focused on write high performance code with specific emphasis on graphic processing unit (GPU). In our teaching experience, parallel pattern teaching has been confirmed to be a useful pedagogical method for teaching parallel concepts. Finally, we report course experience in teaching parallel computing injected course, which resulted in positive student feedback.
Han Wan, Xiaopeng Gao, Xiang Long, Bo Jiang 0001
FIE1
2016 Hybrid teaching mode for laboratory-based remote education of computer structure course
abstract
We describe an Open edX-based blended course developed for a reformed Computer Structure course at Beihang University. In three iterations of this laboratory-based course, we dive into key issues that impact students' learning, and then redesign our curriculum, which integrated with virtual laboratory technique into the MOOC platform. We show how certain course design aspects affect students' learning in the hybrid teaching mode: (a) Kung Fu style competency education with online-support laboratory system prompts students to own their learning as the pace and/or the path of learning, which is dictated by mastery instead of the time/space; (b) strengthen the use of learning-aid tools empowered teachers with the skills and information to define standard for each learner in each stage; (c) automated test technology make this blended learning possible at scale and also financially sustainable; (d) using discussion forum to build the lesson about `what to do' when learners get stuck helped in overcoming challenges.
Han Wan, Xiaopeng Gao
FIE1
2016 On the Decomposition-Based Global EDF Scheduling of Parallel Real-Time Tasks
abstract
Real-time systems are shifting from single-core to multi-core processors, on which software must be parallelized to fully utilize the additional computation power. Recently different types of scheduling algorithms and analysis techniques have been proposed for parallel real-time tasks modeled as directed acyclic graphs (DAG). However, this field is still much less mature than traditional real-time scheduling of sequential tasks. In this paper, we study the decomposition-based scheduling for parallel real-time tasks, where a task graph is transferred to a set of independent sporadic tasks. In particular, we proposed a new decomposition strategy that better explores the feature of each task, represented by its structure characteristic value, to improve schedulability. The structure characteristic values do not only provide a clear guidance in task decomposition, but also can be directly used for schedulability tests, as well as to quantify the suboptimality of our scheduling algorithm in terms of capacity augmentation bounds. We conduct comprehensive experiments to evaluate the real-time performance of our proposed scheduling algorithm, against the state-of-the-art scheduling and analysis methods of different types. Experiment results show that our method consistently outperforms all of the previous methods under different parameter settings.
Xu Jiang 0004, Xiang Long, Nan Guan, Han Wan
RTSS4
2012 Using Basic Block Based Instruction Prefetching to Optimize WCET Analysis for Real-Time Applications
abstract
Cache is an important component existing in modern computer system to bridge the performance gap between the fast CPU and the slow memory system. A variety of cache optimization technologies and mechanisms are proposed to improve the cache performance, such as instruction cache prefetching. Most instruction prefetching mechanisms existing are proposed to improve the average-case cache performance. However, real-time systems care more about the worst-case performance, and the worst-case execution time (WCET) analysis of real-time applications is critical for schedulability analysis of real-time systems. Due to its unpredictable behaviour, cache disastrously complicates the WCET analysis of real-time applications. In this paper, we proposed a basic block based instruction prefetching (BBIP) mechanism to improve both the average-case cache performance and the tightness of the WCET analysis of real-time applications. Measurements on typical real-time benchmarks show that BBIP can not only eliminate most of the instruction access misses, but also result in lower WCET estimations. To discuss the effectiveness of BBIP, we measured the WCET of the benchmarks for three processor configurations with and without BBIP: 1) processor with in-order pipeline and perfect branch prediction, 2) processor with out-of-order pipeline and perfect branch prediction, and 3) processor with out-of-order pipeline and 2-level branch prediction. The results show that BBIP can provide notable improvements in the tightness of WCET estimation, with the WCET values being 30.4% to 97.7% of the original ones. Our simulation results also reveal that 70% to 80% instruction access misses are eliminated with BBIP.
Fan Ni, Xiang Long, Han Wan, Xiaopeng Gao
PDCAT3
2009 GCSim: A GPU-Based Trace-Driven Simulator for Multi-level Cache
Han Wan, Xiaopeng Gao, Xiang Long
APPT1
2009 Using GPU to Accelerate Cache Simulation
abstract
Caches play a major role in the performance of high-speed computer systems. Trace-driven simulator is the most widely used method to evaluate cache architectures. However, as the cache design moves to more complicated architectures, along with the size of the trace is becoming larger and larger. Traditional simulation methods are no longer practical due to their long simulation cycles. Several techniques have been proposed to reduce the simulation time of sequential trace-driven simulation. This paper considers the use of generic GPU to accelerate cache simulation which exploits set-partitioning as the main source of parallelism. We develop more efficient parallel simulation techniques by introducing more knowledge into the Compute Unified Device Architecture (CUDA) on the GPU. Our experimental result shows that the new algorithm gains 2.76x performance improvement compared to traditional CPU-based sequential algorithm.
Han Wan, Xiaopeng Gao
ISPA1