VLDB 2026 Research / reviewers in the wild / expert
Lei Qiao 0002
dblp:54/3557-2
· DBLP profile ↗
43ranked-venue papers
1as first author
34since 2021 · last 2026
0000-0002-2637-9683ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 15 since 2021Software engineering, systems software and programming languages · 11 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 9 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Theory of computation · 2 · 1 first-authorComputer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RefineDedup: efficient deduplication for mobile systems via application-wise learning
Wei Li 0322, Xianzhang Chen, Xingjie Zhou, Duo Liu 0002, Yujuan Tan, Ao Ren, Kan Zhong, Lei Qiao 0002 |
Sci. China Inf. Sci. | 8 |
| 2026 | Orchestrating optimization passes of machine learning compiler for reducing memory footprints of computation graphs
Qianwei Yu, Pengbo Nie, Chengcheng Wan 0001, He Jiang 0001, Jianjun Zhao 0001, Lei Qiao 0002, Yuting Chen 0001 |
J. Syst. Archit. | 8 |
| 2025 | MIFS: A low overhead and efficient mixture file index management method in flash file system
Jingjing Jiang, Mengfei Yang, Lei Qiao 0002, Tingyu Wang 0003 |
J. Syst. Archit. | 3 |
| 2025 | GNNBoost: Accelerating sampling-based GNN training on large scale graph by optimizing data preparation
Yujuan Tan, Yan Gan, Zhaoyang Zeng, Zhuoxin Bai, Lei Qiao 0002, Duo Liu 0002, Kan Zhong, Ao Ren |
J. Syst. Archit. | 5 |
| 2025 | TRACED: A Temporal Graph Neural Networks-based Model for Data PrefetchingabstractIn modern microarchitectures, machine-learning-based prefetchers use past memory requests to learn access patterns and predict memory addresses, thereby prefetching data into the cache to mitigate the processor-memory speed gap. However, they face two key challenges in capturing irregular access patterns generated by complex data structures and algorithms. One is data dispersion: the disorderliness of memory addresses makes it difficult for prefetchers to extract meaningful data features. The other is temporal and spatial complexity: existing prefetchers fail to effectively learn temporal and spatial characteristics, and thus are unable to explore more complex access patterns. To resolve these challenges, we propose TRACED, a novel temporal graph neural network-based prefetcher aimed at learning access patterns of memory addresses. TRACED consists of two key components: a dynamic clustering component and a temporal graph neural network component. In the dynamic clustering component, we introduce a similarity function to quantify the similarity of memory addresses. Based on the quantified similarity, we dynamically group unordered memory addresses into different clusters. This ensures that the memory addresses in each cluster are ordered and change smoothly, thus resolving the first challenge. The temporal graph neural network component constructs a spatiotemporal graph to represent relationships among memory addresses. This helps capture temporal and spatial characteristics both across and within clusters, thus resolving the second challenge. This article demonstrates the effectiveness of the proposed prefetcher through experiments. Specifically, in terms of accuracy, TRACED outperforms BO, SPP, DOMINO, Delta-LSTM, and VOYAGER by 2.29%–40.83% on average. Furthermore, TRACED attains remarkable coverage of 55.67% and IPC of 43.75%, outperforming all competing approaches in both metrics. He Jiang 0001, Liuwei Fu, Dong Liu 0025, Zhilei Ren, Yuting Chen 0001, Lei Qiao 0002 |
ACM Trans. Archit. Code Optim. | 6 |
| 2024 | Control Flow Divergence Optimization by Exploiting Tensor CoresabstractKernels are scheduled on Graphics Processing Units (GPUs) in the granularity of GPU warp, which is a bunch of threads that must be scheduled together. When executing kernels with conditional branches, the threads within a warp may execute different branches sequentially, resulting in a considerable utilization loss and unpredictable execution time. This problem is known as the control flow divergence. In this work, we propose a novel method to predict threads' execution path before the launch of the kernel by deploying a branch prediction network on the GPU's tensor cores, which can efficiently parallel run with the kernels on CUDA cores, so that the divergence problem can be eased in a large extent with the lowest overhead. Combined with a well-designed thread data reorganization algorithm, this solution can better mitigate GPUs' control flow divergence problem. Weiguang Pang, Xu Jiang 0004, Songran Liu, Lei Qiao 0002, Kexue Fu 0001, Longxiang Gao, Wang Yi 0001 |
DAC | 4 |
| 2024 | An Adaptive Real-Time Garbage Collection Method Based on File Write Prediction
Jingjing Jiang, Mengfei Yang, Lei Qiao 0002, Tingyu Wang 0003, Shenghui Zhu |
TASE | 3 |
| 2024 | An efficient schedulability analysis based on worst-case interference time for real-time systems
Hongbiao Liu, Mengfei Yang, Lei Qiao 0002 |
Sci. China Inf. Sci. | 3 |
| 2024 | MTPS: A Multi-Task Perceiving and Scheduling Framework Across Multiple Mobile DevicesabstractThe prevalence of cross-device resource sharing enables users to utilize various device resources of the connected mobile devices seamlessly. Since there are often numerous connected mobile devices under the same network, cross-device tasks are often executed concurrently. However, the existing resource sharing schemes suffer from significant performance degradation for the parallel cross-device tasks due to competition for limited system resources (e.g., network and CPU). This paper first analyzes the performance penalty in parallel execution of the cross-device resource sharing tasks. Then, a novel multi-task perceiving and scheduling framework (MTPS) is proposed to guarantee the quality of service of the parallel tasks. The basic idea of MTPS is to first build a master-slave system model to reorganize mobile devices under the same network. Then, MTPS perceives the running cross-device resource sharing tasks and schedules the parallel execution of multiple tasks to avoid mutual interference. Experimental results on real devices show that MTPS can reduce the average completion time of file sharing by 63.5%, and maintain at least 24 frames per second for screen casting at optimal levels in the presence of other tasks. Wentong Li 0002, Lei Qiao 0002, Liang Shi 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | What's Wrong With Low-Code Development Platforms? An Empirical Study of Low-Code Development Platform BugsabstractLow-code development platforms (LCDPs) are increasingly being introduced and leveraged by major IT enterprises to lower the threshold and promote the efficiency of software development. Like other software systems, LCDPs are also inevitable to have bugs. The bugs in LCDPs may cause unpredictable consequences as they pose risks to all the downstream software products. However, to the best of our knowledge, there exist no studies that ever consider the bugs caused by LCDPs. To handle the LCDP bugs better, in this article, we conduct an empirical study of the characteristics of LCDP bugs by examining 974 confirmed bugs of four dominant LCDPs (i.e., OutSystems, Mendix, Appsmith, and Budibase) from both commercial and open-source domains. These bugs are analyzed from three perspectives, including bug root causes, bug symptoms, and the affected stages of LCDPs. Based on the analysis, we obtain a series of valuable findings. For example, around 60% of the bugs reside in the stage of designing and specifying the developed applications. Over 37% of the bugs lead LCDPs to behave unexpectedly but without showing explicit signs. Moreover, the bugs relevant to the incorrect graphics of user interfaces are significant due to the characteristics of LCDPs. These findings point out the guidelines, challenges, and future directions to address LCDP bugs. Dong Liu 0025, He Jiang 0001, Shikai Guo, Yuting Chen 0001, Lei Qiao 0002 |
IEEE Trans. Reliab. | 5 |
| 2023 | Efficient CUDA stream management for multi-DNN real-time inference on embedded GPUs
Weiguang Pang, Xiantong Luo, Kailun Chen, Dong Ji, Lei Qiao 0002, Wang Yi 0001 |
J. Syst. Archit. | 5 |
| 2023 | Scanner++: Enhanced Vulnerability Detection of Web Applications with Attack Intent SynchronizationabstractScanners are commonly applied for detecting vulnerabilities in web applications. Various scanners with different strategies are widely in use, but their performance is challenged by the increasing diversity of target applications that have more complex attack surfaces (i.e., website paths) and covert vulnerabilities that can only be exploited by more sophisticated attack vectors (i.e., payloads). In this paper, we propose Scanner++, a framework that improves web vulnerability detection of existing scanners through combining their capabilities with attack intent synchronization. We design Scanner++ as a proxy-based architecture while using a package-based intent synchronization approach. Scanner++ first uses a purification mechanism to aggregate and refine attack intents, consisting of attack surfaces and attack vectors extracted from the base scanners’ request packets. Then, Scanner++ uses a runtime intent synchronization mechanism to select relevant attack intents according to the scanners’ detection spots to guide their scanning process. Consequently, base scanners can expand their attack surfaces, generate more diverse attack vectors and achieve better vulnerability detection performance. For evaluation, we implemented and integrated Scanner++ together with four widely used scanners, BurpSuite, AWVS, Arachni, and ZAP, testing it on ten benchmark web applications and three well-tested real-world web applications of a critical financial platform from our industry partner. Working under the Scanner++ framework helps BurpSuite, AWVS, Arachni, and ZAP cover 15.26%, 37.14%, 59.21%, 68.54% more pages, construct 12.95×, 1.13×, 15.03×, 52.66× more attack packets, and discover 77, 55, 77, 176 more bugs, respectively. Furthermore, Scanner++ detected eight serious previously unknown vulnerabilities on real-world applications, while the base scanners only found three of them. Zijing Yin, Fuchen Ma, Haohao Gao, Lei Qiao 0002, Yu Jiang 0001 |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2023 | Scheduling Parallel Real-Time Tasks on Virtual ProcessorsabstractIn many popular parallel programming models, e.g., OpenMP (OpenMP, 2013), applications are usually dispatched into several dedicated scheduling entities (named ”threads” in common) for which the processor time of physical platform is provided through the OS schedulers. This behavior requires for a hierarchical scheduling framework, considering each thread as a virtual processor (VP). Moreover, hierarchical scheduling allow separate applications to execute together on a common hardware platform, with each application having the “illusion” of executing on a dedicated component. However, the problem for scheduling parallel real-time tasks on virtual multiprocessor platform has not been addressed yet. An analogous approach to virtual scheduling for parallel real-time tasks is federeted scheudling, where each task exclusively executes on a set of dedicated physical processors. However, federated scheduling suffers significant resource wasting. In this article, we study the scheduling of real-time parallel task on virtual multiprocessors. As a physical processor is shared by virtual processors, tasks effectively share processors with each other. We conduct comprehensive performance evaluation to compare our proposed approach with existing methods of different types. Experiment results show that our approach consistently outperforms existing methods to a considerable extent under a wide range of parameter settings. Xu Jiang 0004, Haochun Liang, Nan Guan, Yue Tang 0001, Lei Qiao 0002, Wang Yi 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2023 | Detecting C++ Compiler Front-End Bugs via Grammar Mutation and Differential TestingabstractC++ is a widely used programming language and the C++ front-end is a critical part of a C++ compiler. Although many techniques have been proposed to test compilers, few studies are devoted to detecting bugs in C++ compiler. In this study, we take the first step to detect bugs in C++ compiler front-ends. To do so, two main challenges need to be addressed, namely, the acquisition of test programs that are more likely to trigger bugs in compiler front-ends and the bug identification from complicated compiler outputs. In this article, we propose a novel framework namedCcoftto detect bugs in C++ compiler front-ends. To address the first challenge,Ccoftimplements a practical program generator. The generator first transforms C++ grammars into a flexible structured format and then utilizes an equal-chance selection (ECS) strategy to conduct structure-aware grammar mutation to generate diverse C++ programs. Next,Ccoftemploys a set of differential testing strategies to identify various kinds of bugs in C++ compiler front-ends by comparing complex outputs emitted by C++ compilers, thus tackling the second challenge. Empirical evaluation results over two mainstream compilers (i.e., GCC and Clang) show thatCcoftgreatly improves two state-of-the-art approaches (i.e., Dharma and Grammarinator) by 135% and 111% in terms of the numbers of detected bugs, respectively. By runningCcoftfor three months, we have successfully reported 136 bugs for two C++ compilers, of which 78 (57 confirmed, assigned, or fixed) for GCC and 58 (10 confirmed or fixed) for Clang. Haoxin Tu, He Jiang 0001, Zhide Zhou, Zhilei Ren, Lei Qiao 0002, Lingxiao Jiang |
IEEE Trans. Reliab. | 6 |
| 2022 | Hierarchical memory-constrained operator scheduling of neural architecture search networksabstractNeural Architecture Search (NAS) is widely used in industry, searching for neural networks meeting task requirements. Meanwhile, it faces a challenge in scheduling networks satisfying memory constraints. This paper proposes HMCOS that performs hierarchical memory-constrained operator scheduling of NAS networks: given a network, HMCOS constructs a hierarchical computation graph and employs an iterative scheduling algorithm to progressively reduce peak memory footprints. We evaluate HMCOS against RPO and Serenity (two popular scheduling techniques). The results show that HMCOS outperforms existing techniques in supporting more NAS networks, reducing 8.7~42.4% of peak memory footprints, and achieving 137--283x of speedups in scheduling. Chengcheng Wan 0001, Yuting Chen 0001, He Jiang 0001, Lei Qiao 0002 |
DAC | 6 |
| 2022 | Optimizing CoW-based File Systems on Open-Channel SSDs with Persistent MemoryabstractBlock-based file systems, such as Btrfs, utilize the copy-on-write (CoW) mechanism to guarantee data consistency on solid-state drives (SSDs). Open-channel SSD provides opportunities for in-depth optimization of block-based file systems. However, existing systems fail to co-design the two-layer semantics and cannot take full advantage of the open-channel characteristics. Specifically, synchronizing an overwrite in Btrfs will copy-on-write all pages in the update path and induce severe write amplification. In this paper, we propose a hybrid fine-grained copy-on-write and journaling mechanism (HyFiM) to address these problems. We first utilize persistent memories to preserve the address mapping table of open-channel SSD. Then, we design an intra-FTL copy-on-write mechanism (IFCoW) that eliminates the recursive updates caused by overwrites. Finally, we devise fine-grained metadata journals (FGMJ) to guarantee the consistency of metadata with minimum overhead. We prototype HyFiM based on Btrfs in the Linux kernel. Comprehensive evaluations demonstrate that HyFiM can outperform over Btrfs by 30.77% and 33.82% for sequential and random overwrites, respectively. Runyu Zhang 0002, Duo Liu 0002, Chaoshu Yang, Xianzhang Chen, Lei Qiao 0002, Yujuan Tan |
DATE | 5 |
| 2022 | Surrogate-Assisted Multi-objective Optimization for Compiler Optimization Sequence Selection
Guojun Gao, Lei Qiao 0002, Dong Liu 0025, Shifei Chen, He Jiang 0001 |
PPSN (2) | 2 |
| 2022 | Detecting Compiler Bugs Via a Deep Learning-Based FrameworkabstractCompiler testing is the most widely used way to assure compiler quality. However, since compilers require a large number of sophisticated test programs as inputs, the existing approaches in compiler testing still have a limited capability in generating both syntactically valid and diverse test programs. In this paper, we propose DeepGen, a deep learning-based approach to support compiler testing through the inference of a generative model for compiler inputs. First, DeepGen trains a Transformer-XL model based on a large corpus of seed programs, and uses the trained model to generate syntactically valid programs. Then, DeepGen adopts a sampling strategy in the inference phase to generate diverse test programs. Finally, DeepGen leverages differential testing on the generated programs to discover compiler bugs. We have evaluated DeepGen over two popular C++ compilers GCC and LLVM, and the results confirm the effectiveness of our approach. DeepGen detects 35.29%, 53.33%, and 187.50% more bugs than three existing approaches, i.e. DeepSmith, DeepFuzz, and Csmith, respectively. In addition, 30.43% bugs detected by DeepGen are not detected by other approaches. Furthermore, DeepGen has successfully detected 38 bugs in the latest development versions of GCC and LLVM; 21 of them have been confirmed/fixed by the developers. Zhilei Ren, He Jiang 0001, Lei Qiao 0002, Dong Liu 0025, Zhide Zhou, Weiqiang Kong |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2022 | Automatically repairing tensor shape faults in deep learning programs
Dangwei Wu, Beijun Shen, Yuting Chen 0001, He Jiang 0001, Lei Qiao 0002 |
Inf. Softw. Technol. | 5 |
| 2022 | CoDiscard: A revenue model based cross-layer cooperative discarding mechanism for flash memory devices
Xiaoliu Feng, Xianzhang Chen, Ruolan Li, Chunlin Song, Duo Liu 0002, Yujuan Tan, Lei Qiao 0002 |
J. Syst. Archit. | 8 |
| 2022 | ELOFS: An Extensible Low-Overhead Flash File System for Resource-Scarce Embedded DevicesabstractEmerging applications like machine learning in embedded devices (e.g., satellites and vehicles) require huge storage space, which recently stimulates the widespread deployment of large-scale flash memory in IoT devices. However, existing embedded file systems fall short in managing large-capacity storage efficiently for two reasons. First, prior arts store data structures of file systems either in flash or in main memory, which severely magnifies the scarcity of computing and memory resources. Moreover, the fine-grained metadata management in the existing embedded file systems induces significant energy consumption for large-capacity storage. In this paper, we propose a novel embedded file system, ELOFS, to tackle the above issues and manage large-capacity NAND flash on resource-scarce devices. ELOFS is made efficient through three novel techniques. First, we redefine the space management granularity and streamline the metadata to speed up the mounting performance. In addition, we design hybrid file structures to adapt dissimilar access patterns of embedded devices. Furthermore, ELOFS provides opportunities for in-depth cooperation with application-specific systems. We implement ELOFS with Memory Technology Device (MTD) interfaces, and the experimental results show that ELOFS outperforms YAFFS and UBIFS in terms of write, read, and deletions with orders of magnitude reductions on memory footprint and mounting time. Runyu Zhang 0002, Duo Liu 0002, Xianzhang Chen, Xiongxiong She, Chaoshu Yang, Yujuan Tan, Zhaoyan Shen, Zili Shao, Lei Qiao 0002 |
IEEE Trans. Computers | 9 |
| 2022 | Horae: A Hybrid I/O Request Scheduling Technique for Near-Data Processing-Based SSDabstractNear-data processing (NDP) architecture is promised to break the bottleneck of data movement in many scenarios (e.g., databases and recommendation systems), which limits the efficiency of data processing. Different from traditional SSD, NDP-based SSD not only needs to handle normal I/Os (e.g., read and write), but also needs to handle NDP requests that contain data processing operations. NDP and normal I/O requests share some function units of NDP-based SSD, such as flash chips and embedded processors. However, existing works ignore the resource competition between normal I/Os and NDP requests, which drastically degrades the performance. In this article, we propose a novel scheduling technique called Horae, which can efficiently schedule hybrid NDP-normal I/O requests in NDP-based SSD to improve performance. Horae exploits the critical paths on critical resources to maximize the parallelism of multiple stages of requests. The experimental results on typical workloads show that Horae can significantly improve the performance of hybrid NDP-normal I/O requests over the state-of-the-art scheduling algorithms of NDP-based SSDs. Xianzhang Chen, Duo Liu 0002, Jiapin Wang, Zhaoyang Zeng, Yujuan Tan, Lei Qiao 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2022 | LocSeq: Automated Localization for Compiler Optimization Sequence Bugs of LLVMabstractCompiler bugs may be triggered when programs are optimized with optimization sequences. However, diagnosing compiler optimization sequence bugs is difficult due to limited debugging information. Although some techniques (e.g., DiWi and RecBi) have been proposed to automatically localize compiler bugs, no systematic work has been conducted to automatically localize compiler optimization sequence bugs. In this article, we propose LocSeq, a novel technique to automatically localize compiler optimization sequence bugs of LLVM. The core insight of LocSeq is based on the fact that the behaviors of optimizations may be influenced by each other, and thus, the innocent files may be excluded by constructing bug-free optimization sequences. First, given a buggy optimization sequence that triggers a compiler bug, in LocSeq, we transform the problem of the localization for a compiler optimization sequence bug to the problem of the construction for bug-free optimization sequences, which are helpful to localize buggy compiler files. Then, a constrained genetic algorithm is presented in LocSeq to generate a set of bug-free optimization sequences that share similar compiler execution traces with the buggy optimization sequence. Finally, LocSeq leverages a spectrum-based bug localization technique to localize the compiler optimization sequence bug by comparing the execution traces between bug-free optimization sequences and the buggy optimization sequence. To evaluate the effectiveness of LocSeq, we build a benchmark, including 60 optimization sequence bugs of LLVM, and compare LocSeq with the state-of-the-art techniques DiWi and RecBi. The experimental results show that LocSeq significantly outperforms DiWi and RecBi by up to 366.66%/72.27% and 250.00%/56.00% for localizing optimization sequence bugs within Top-1/5 files, respectively. Zhide Zhou, He Jiang 0001, Zhilei Ren, Yuting Chen 0001, Lei Qiao 0002 |
IEEE Trans. Reliab. | 5 |
| 2022 | DPWord2Vec: Better Representation of Design Patterns in SemanticsabstractWith the plain text descriptions of design patterns, developers could better learn and understand the definitions and usage scenarios of design patterns. To facilitate the automatic usage of these descriptions, e.g., recommending design patterns by free-text queries, design patterns and natural languages should be adequately associated. Existing studies usually use texts in design pattern books as the representations of design patterns to calculate similarities with the queries. However, this way is problematic. Lots of information of design patterns may be absent from design pattern books and many words would be out of vocabulary due to the content limitation of these books. To overcome these issues, a more comprehensive method should be constructed to estimate the relatedness between design patterns and natural language words. Motivated by Word2Vec, in this study, we propose DPWord2Vec that embeds design patterns and natural language words into vectors simultaneously. We first build a corpus containing more than 400 thousand documents extracted from design pattern books, Wikipedia, and Stack Overflow. Next, we redefine the concept of context window to associate design patterns with words. Then, the design pattern and word vector representations are learnt by leveraging an advanced word embedding method. The learnt design pattern and word vectors can be universally used in textual description based design pattern tasks. An evaluation shows that DPWord2Vec outperforms the baseline algorithms by 24.2-120.9 percent in measuring the similarities between design patterns and words in terms of Spearman’s rank correlation coefficient. Moreover, we adopt DPWord2Vec on two typical design pattern tasks. In the design pattern tag recommendation task, the DPWord2Vec-based method outperforms two state-of-the-art algorithms by 6.6 and 32.7 percent respectively when considering$Recall@10$. In the design pattern selection task, DPWord2Vec improves the existing methods by 6.5-70.7 percent in terms of MRR. Dong Liu 0025, He Jiang 0001, Zhilei Ren, Lei Qiao 0002, Zuohua Ding |
IEEE Trans. Software Eng. | 5 |
| 2022 | Pluto: Exposing Vulnerabilities in Inter-Contract ScenariosabstractAttacks on smart contracts have caused considerable losses to digital assets. Many techniques based on symbolic execution, fuzzing, and static analysis are used to detect contract vulnerabilities. Most of the current analyzers only consider vulnerability detection intra-contract scenarios. However, Ethereum contracts usually interact with others by calling their functions. A bug hidden in a path that depends on information from external contract calls is defined as an inter-contract vulnerability. Failure to deal with this kind of bug can result in potential false negatives and false positives. In this work, we propose Pluto, which supports vulnerability detection in inter-contract scenarios. It first builds an Inter-contract Control Flow Graph (ICFG) to extract semantic information among contract calls. Afterward, it symbolically explores the ICFG and deduces Inter-Contract Path Constraints (ICPC) to check the reachability of execution paths more accurately. Finally, Pluto detects whether there is a vulnerability based on some predefined rules. For evaluation, we compare Pluto with five state-of-the-art tools, including Oyente, Mythril, Securify, ILF, and Clairvoyance on a labeled benchmark and 39,443 real-world Ethereum smart contracts. The result shows that other tools can only detect 10% of the inter-contract vulnerabilities, while Pluto can detect 80% of them on the labeled dataset. Beyond that, Pluto has detected 451 confirmed vulnerabilities on real-world contracts, including 36 vulnerabilities in inter-contract scenarios. Two bugs have been assigned with unique CVE identifiers by the US National Vulnerability Database (NVD). On average, Pluto costs 16.9 seconds to analyze a contract, which is as fast as the state-of-the-art tools. Fuchen Ma, Zijing Yin, Yuanliang Chen, Lei Qiao 0002, Bin Gu 0006, Huizhong Li, Yu Jiang 0001, Jia-Guang Sun 0001 |
IEEE Trans. Software Eng. | 6 |
| 2021 | ORBBuf: A Robust Buffering Method for Remote Visual SLAMabstractThe data loss caused by unreliable network seriously impacts the results of remote visual SLAM systems. From our experiment, a loss of less than 1 second of data can cause a visual SLAM algorithm to lose tracking. We present a novel buffering method, ORBBuf, to reduce the impact of data loss on remote visual SLAM systems. We model the buffering problem as an optimization problem by introducing a similarity metric between frames. To solve the buffering problem, we present an efficient greedy algorithm to discard the frames that have the least impact on the quality of SLAM results. We implement our ORBBuf method on ROS, a widely used middleware framework. Through an extensive evaluation on real-world scenarios and tens of gigabytes of datasets, we demonstrate that our ORBBuf method can be applied to different state-estimation algorithms (DSO and VINS-Fusion), different sensor data (both monocular images and stereo images), different scenes (both indoor and outdoor), and different network environments (both WiFi networks and 4G networks). Our experimental results indicate that the network losses indeed affect the SLAM results, and our ORBBuf method can reduce the RMSE up to 50 times comparing with the Drop-Oldest and Random buffering methods. Yu-Ping Wang 0001, Zixin Zou, Cong Wang 0045, Yue-Jiang Dong, Lei Qiao 0002, Dinesh Manocha |
IROS | 5 |
| 2021 | Tensfa: Detecting and Repairing Tensor Shape Faults in Deep Learning SystemsabstractSoftware developers frequently invoke deep learning (DL) APIs to incorporate learning solutions into software systems. However, misuses of these APIs can cause various DL faults, such as tensor shape faults. Tensor shape faults occur when restriction conditions of operations are not met; they are prevalent in practice, leading to many system crashes. Meanwhile, researchers and engineers still face a strong challenge in detecting tensor shape faults ─ static techniques incur heavy overheads in defining detection rules, and the only dynamic technique requires human engineers to rewrite APIs for tracking shape changes. To address the above challenge, we conduct a deep empirical study on crashing tensor shape faults (i.e., those causing programs to crash), categorizing them into four types and revealing twelve repair patterns. We then propose and implement Tensfa, an approach to detecting and repairing crashing tensor shape faults. Tensfa takes a machine learning method to learn from crash messages and employs decision trees in detecting tensor shape faults. Tensfa also provides the first automated solution to repairing the detected faults: it tracks shape properties by a customized Python debugger, analyzes their data dependences, and uses the twelve patterns to generate patches. We construct SFData, a set of 146 buggy programs with crashing tensor shape faults. Our Tensfa has been implemented and evaluated on SFData and IslamData (another dataset of tensor shape faults). The results clearly show the effectiveness of Tensfa. In particular, Tensfa achieves the state-of-the-art results: it reaches an F1-score of 96.88% in detecting the faults and repairs 80 out of 146 buggy programs in SFData. Dangwei Wu, Beijun Shen, Yuting Chen 0001, He Jiang 0001, Lei Qiao 0002 |
ISSRE | 5 |
| 2021 | Virtually-Federated Scheduling of Parallel Real-Time TasksabstractFederated scheduling is a promising approach to schedule parallel real-time tasks, where each task exclusively executes on a set of dedicated processors. However, federated scheduling suffers significant resource wasting since a task typically only uses part of the processing capacity allocated to it, while the unused part cannot be shared with other tasks. To solve this problem, we present a virtually-federated scheduling approach, which both enjoys the good analyzability of federated scheduling and allows tasks to efficiently share processors with others. The main idea is to construct virtual processors on physical processors, and let a task exclusively execute on a set of virtual processors. As a physical processor is shared by virtual processors, tasks effectively share processors with each other. On the other hand, as each task exclusively executes on its own virtual processor set, the good analyzability of federated scheduling can be carried into to our virtually-federated scheduling approach. We conduct comprehensive performance evaluation to compare our proposed approach with existing methods of different types. Experiment results show that our approach consistently outperforms existing methods to a considerable extent under a wide range of parameter settings. Xu Jiang 0004, Nan Guan, Haochun Liang, Yue Tang 0001, Lei Qiao 0002, Wang Yi 0001 |
RTSS | 5 |
| 2021 | Sound and efficient concurrency bug predictionabstractConcurrency bugs are extremely difficult to detect. Recently, several dynamic techniques achieve sound analysis. M2 is even complete for two threads. It is designed to decide whether two events can occur consecutively. However, real-world concurrency bugs can involve more events and threads. Some can occur when the order of two or more events can be exchanged even if they occur not consecutively. We propose a new technique SeqCheck to soundly decide whether a sequence of events can occur in a specified order. The ordered sequence represents a potential concurrency bug. And several known forms of concurrency bugs can be easily encoded into event sequences where each represents a way that the bug can occur. To achieve it, SeqCheck explicitly analyzes branch events and includes a set of efficient algorithms. We show that SeqCheck is sound; and it is also complete on traces of two threads. Yan Cai 0001, Hao Yun, Jinqiu Wang, Lei Qiao 0002, Jens Palsberg |
ESEC/SIGSOFT FSE | 4 |
| 2021 | STUaNet: Understanding Uncertainty in Spatiotemporal Collective Human MobilityabstractThe high dynamics and heterogeneous interactions in the complicated urban systems have raised the issue of uncertainty quantification in spatiotemporal human mobility, to support critical decision-makings in risk-aware web applications such as urban event prediction where fluctuations are of significant interests. Given the fact that uncertainty quantifies the potential variations around prediction results, traditional learning schemes always lack uncertainty labels, and conventional uncertainty quantification approaches mostly rely upon statistical estimations with Bayesian Neural Networks or ensemble methods. However, they have never involved any spatiotemporal evolution of uncertainties under various contexts, and also have kept suffering from the poor efficiency of statistical uncertainty estimation while training models with multiple times. To provide high-quality uncertainty quantification for spatiotemporal forecasting, we propose an uncertainty learning mechanism to simultaneously estimate internal data quality and quantify external uncertainty regarding various contextual interactions. To address the issue of lacking labels of uncertainty, we propose a hierarchical data turbulence scheme where we can actively inject controllable uncertainty for guidance, and hence provide insights to both uncertainty quantification and weak supervised learning. Finally, we re-calibrate and boost the prediction performance by devising a gated-based bridge to adaptively leverage the learned uncertainty into predictions. Extensive experiments on three real-world spatiotemporal mobility sets have corroborated the superiority of our proposed model in terms of both forecasting and uncertainty quantification. Zhengyang Zhou, Yang Wang 0015, Xike Xie, Lei Qiao 0002, Yuantao Li |
WWW | 4 |
| 2021 | Verification of Real Time Operating System Exception Management Based on SPARCv8
Lei Qiao 0002, Mengfei Yang, Jin-Kun Zhang |
J. Comput. Sci. Technol. | 2 |
| 2021 | Blocking analysis of suspension-based protocols for parallel real-time tasks under global fixed-priority scheduling
Ze-Wei Chen, Maolin Yang 0004, Lei Qiao 0002 |
J. Syst. Archit. | 5 |
| 2021 | A Hierarchical Hybrid Locking Protocol for Parallel Real-Time TasksabstractParallel tasks have been paid growing attention in recent years, and the scheduling with shared resources is of significant importance to real-time systems. As an efficient mechanism to provide mutual exclusion for parallel processing, spin-locks are ubiquitous in multi-processor real-time systems. However, the spin-locks suffer the scalability problem, and the intra-task parallelism further exacerbates the analytical pessimism. To overcome such deficiencies, we propose a Hierarchical Hybrid Locking Protocol (H2LP) under federated scheduling. The proposed H2LP integrates the classical Multiprocessor Stack Resource Policy (MSRP) and uses a token mechanism to reduce global contentions. We provide a complete analysis framework supporting both heavy and light tasks under federated scheduling and develop a blocking analysis with the state-of-the-art linear optimization technique. Empirical evaluations showed that the H2LP outperformed the other state-of-the-art locking protocols in at least configurations when considering exclusive clustering. Furthermore, our partitioned approach for light tasks can substantially improve schedulability by mitigating the over-provisioning problem. Zewei Chen, Maolin Yang 0004, Lei Qiao 0002 |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2021 | Memory State Verification Based on Inductive and Deductive ReasoningabstractMemory allocation and deallocation are the fundamental operations of embedded operating systems, which have been extensively used in many safety critical systems. The correctness of the operations is of paramount importance because their failure could incur severe consequences. While the system is running, the memory state can easily grow to a gigantic amount, which means that it is impossible to verify the huge memory states one by one. Therefore, it is a challenge how to verify the correctness of running memory state of the system. In this article, we propose a novel memory state verification method based on inductive and deductive reasoning. First, we abstract the memory state as a list of memory blocks, which will transform in memory operations. Second, we construct the generic model based on the transition function of the memory management and summarize the invariant properties of the memory state. Third, we use the inductive method to calculate the changes between the memory states, and verify that the memory state of the system always satisfy the global properties. All the proofs are implemented in the interactive theorem prover Coq. On the basis of our proposed model, we verify the correctness of a two-level segregated fit (TLSF) algorithm through some extensions, and we also apply this method to verify the correctness of the memory state of the embedded system at runtime. Lei Qiao 0002, Mengfei Yang |
IEEE Trans. Reliab. | 2 |
| 2020 | Modular Verification of SPARCv8 Code
Junpeng Zha, Xinyu Feng 0001, Lei Qiao 0002 |
J. Comput. Sci. Technol. | 3 |
| 2020 | Formalizing SPARCv8 instruction set architecture in Coq
Ming Fu, Lei Qiao 0002, Xinyu Feng 0001 |
Sci. Comput. Program. | 3 |
| 2019 | Tumbler: Energy Efficient Task Scheduling for Dual-Channel Solar-Powered Sensor NodesabstractEnergy harvesting technology has been popularly adopted in embedded systems. However, unstable energy source results in unsteady operation. In this paper, we devise a long-term energy efficient task scheduling targeting for solar-powered sensor nodes. The proposed method exploits a reinforcement learning with a solar energy prediction method to maximize the energy efficiency, which finally enhances the long-term quality of services (QoS) of the sensor nodes. Experimental results show that the proposed scheduling improves the energy efficiency by 6.0%, on average and achieves the better QoS level by 54.0%, compared with a state-of-the-art task scheduling algorithm. Hyung Gyu Lee, Yujuan Tan, Yu Wu 0016, Xianzhang Chen, Liang Liang 0002, Lei Qiao 0002, Duo Liu 0002 |
DAC | 7 |
| 2019 | Astraea: Self-Balancing Federated Learning for Improving Classification Accuracy of Mobile Deep Learning ApplicationsabstractFederated learning (FL) is a distributed deep learning method which enables multiple participants, such as mobile phones and IoT devices, to contribute a neural network model while their private training data remains in local devices. This distributed approach is promising in the edge computing system where have a large corpus of decentralized data and require high privacy. However, unlike the common training dataset, the data distribution of the edge computing system is imbalanced which will introduce biases in the model training and cause a decrease in accuracy of federated learning applications. In this paper, we demonstrate that the imbalanced distributed training data will cause accuracy degradation in FL. To counter this problem, we build a self-balancing federated learning framework call Astraea, which alleviates the imbalances by 1) Global data distribution based data augmentation, and 2) Mediator based multi-client rescheduling. The proposed framework relieves global imbalance by runtime data augmentation, and for averaging the local imbalance, it creates the mediator to reschedule the training of clients based on Kullback-Leibler divergence (KLD) of their data distribution. Compared with FedAvg, the state-of-the-art FL algorithm, Astraea shows +5.59% and +5.89% improvement of top-1 accuracy on the imbalanced EMNIST and imbalanced CINIC-10 datasets, respectively. Meanwhile, the communication traffic of Astraea can be 92% lower than that of FedAvg. Moming Duan, Duo Liu 0002, Xianzhang Chen, Yujuan Tan, Jinting Ren, Lei Qiao 0002, Liang Liang 0002 |
ICCD | 6 |
| 2019 | Archivist: A Machine Learning Assisted Data Placement Mechanism for Hybrid Storage SystemsabstractWith the rapid growth of edge-cloud computing, emerging applications pose higher performance demand on the storage system for storing massive data that are generated from various sources. The multi-sourced data shows different properties in size, retention time, and read/write frequency. Hybrid storage system is promised to efficiently handle the data in edge-cloud computing environment satisfying different data demands. The key problem is how to place the data on the hybrid storage system according to the run-time status and the properties of both data and the storage systems. In this paper, we propose Archivist - a machine learning assisted data placement mechanism for hybrid storage systems to reduce file access latency. We first design a machine learning based approach for predicting the access patterns of the incoming data. Then, we present a data placement algorithm to optimize the data on the hybrid storage mediums by matching the properties of data and the features of storage mediums. Extensive experimental results show that Archivist can achieve up to 49% improvement of system performance for file accesses compared with baseline. Jinting Ren, Xianzhang Chen, Yujuan Tan, Duo Liu 0002, Moming Duan, Liang Liang 0002, Lei Qiao 0002 |
ICCD | 7 |
| 2019 | A Formal Modeling and Verification Framework for Flash Translation Layer Algorithms
Lei Qiao 0002, Mengfei Yang |
SETTA | 1 |
| 2018 | Modular Verification of SPARCv8 Code
Junpeng Zha, Xinyu Feng 0001, Lei Qiao 0002 |
APLAS | 3 |
| 2018 | Formal modelling of list based dynamic memory allocators
Bin Fang 0004, Mihaela Sighireanu, Geguang Pu, Jean-Raymond Abrial, Mengfei Yang, Lei Qiao 0002 |
Sci. China Inf. Sci. | 7 |
| 2017 | Formalizing SPARCv8 Instruction Set Architecture in Coq
Ming Fu, Lei Qiao 0002, Xinyu Feng 0001 |
SETTA | 3 |