Mingyuan Xia 0001

dblp:88/611 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0001-5899-0295ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 5 since 2021Systems, architecture and hardware · 6 · 3 first-author · 1 since 2021Computer networks · 4 · 1 since 2021Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 RealSanitizer: Detecting Floating-Point Errors via a Dynamic Precision Oracle
Yuantao Hu, Shizhong Zhao, Yun Wang 0039, Mingyuan Xia 0001
TASE4
2025 DevTrace: Lightweight Plug-In Design for PCIe Transaction Tracing in Edge Intelligence Workloads
abstract
The complexity of host-peripheral interactions during high-load tasks poses significant challenges for system optimization, with existing tracing tools degrading performance by up to 5.39×. We introduce DevTrace, a novel low-overhead tracing framework for peripheral interactions. Its modular architecture separates data collection from kernel-level operations, enabling lightweight tracing with minimal driver modifications across entire classes of devices. By eliminating heavy kernel tracing interrupts, DevTrace reduces overhead to negligible levels while maintaining data accuracy. In edge-based intelligence deployments, DevTrace achieves a 128× reduction in memory usage and approximately 10× lower CPU overhead compared to page-fault-based solutions. It significantly reduces data loss and performance degradation under high-load conditions, establishing it as a reliable tool for analyzing host-peripheral interactions and optimizing performance in resource-constrained environments. We also discuss potential extensions to eBPF to further decouple tracing from driver frameworks.
Zhibai Huang, Kailiang Xu, Zhixiang Wei, Yinghao Deng, Chen Chen 0067, Yun Wang 0039, Fangxin Liu, Mingyuan Xia 0001, Zhengwei Qi
ICCAD8
2024 Foliage: Nourishing Evolving Software by Characterizing and Clustering Field Bugs
abstract
Modern programs, characterized by their complex functionalities, high integration, and rapid iteration cycles, are prone to errors. This complexity poses challenges in program analysis and software testing, making it difficult to achieve comprehensive bug coverage during the development phase. As a result, many bugs are only discovered during the software’s production phase. Tracking and understanding these field bugs is essential but challenging: the uploaded field error reports are extensive, and trivial yet high-frequency bugs can overshadow important low-frequency bugs. Additionally, application codebases evolve rapidly, causing a single bug to produce varied exceptions and stack traces across different code releases. In this paper, we introduce Foliage, a bug tracking and clustering toolchain designed to trace and characterize field bugs in JavaScript applications, aiding developers in locating and fixing these bugs. To address the challenges of efficiently tracking and analyzing the dynamic and complex nature of software bugs, Foliage proposes an error message enhancement technique. Foliage also introduces the verbal-characteristic-based clustering technique, along with three evaluation metrics for bug clustering: V-measure, cardinality bias, and hit rate. The results show that Foliage’s verbal-characteristic-based bug clustering outperforms previous bug clustering approaches by an average of 31.1% across these three metrics. We present an empirical study of Foliage applied to a complex real-world application over a two-year production period, capturing over 250,000 error reports and clustering them into 132 unique bugs. Finally, we open-source a bug dataset consisting of real and labeled error reports, which can be used to benchmark bug clustering techniques.
Zhanyao Lei, Yixiong Chen, Mingyuan Xia 0001, Zhengwei Qi
ISSTA3
2023 Bootstrapping Automated Testing for RESTful Web Services
abstract
Modern RESTful services expose RESTful APIs to integrate with diversified applications. Most RESTful API parameters are weakly typed, which greatly increases the possible input value space. Weakly-typed parameters pose difficulties for automated testing tools to generate effective test cases to reveal web service defects related to parameter validation. We call this phenomenon the type collapse problem. To remedy this problem, we introduce FET (Format-encoded Type) techniques, including the FET, the FET lattice, and the FET inference to model fine-grained information for API parameters. Inferred FET can enhance parameter validation, such as generating a parameter validator for a certain RESTful server. Enhanced by FET techniques, automated testing tools can generate targeted test cases. We demonstrate Leif, a trace-driven fuzzing tool, as a proof-of-concept implementation of FET techniques. Experiment results on 27 commercial services show that FET inference precisely captures documented parameter definitions, which helps Leif discover 11 new bugs and reduce$72\% - 86\%$fuzzing time compared to state-of-the-art fuzzers. Leveraged by the inter-parameter dependency inference, Leif saves$15\%$fuzzing time.
Zhanyao Lei, Yixiong Chen, Mingyuan Xia 0001, Zhengwei Qi
IEEE Trans. Software Eng.4
2022 AppSPIN: reconfiguration-based responsiveness testing and diagnosing for Android Apps
Zhanyao Lei, Wenhua Zhao, Zhenkai Ding, Mingyuan Xia 0001, Zhengwei Qi
Autom. Softw. Eng.4
2021 Bootstrapping Automated Testing for RESTful Web Services
abstract
Abstract Modern RESTful services expose RESTful APIs to integrate with diversified applications. Most RESTful API parameters are weakly typed, which greatly increases the possible input value space. This poses difficulties for automated testing tools to generate effective test cases to reveal web service defects related to parameter validation. We call this phenomenon the type collapse problem. To remedy this problem, we introduce FET (Format-encoded Type) techniques, including the FET, the FET lattice, and the FET inference to model fine-grained information for API parameters. Enhanced by FET techniques, automated testing tools can generate targeted test cases. We demonstrate Leif, a trace-driven fuzzing tool, as a proof-of-concept implementation of FET techniques. Experiment results on 27 commercial services show that FET inference precisely captures documented parameter definitions, which helps Leif to discover 11 new bugs and reduce $$72\% \sim 86\%$$ 72 % ∼ 86 % fuzzing time as compared to state-of-the-art fuzzers.
Yixiong Chen, Zhanyao Lei, Mingyuan Xia 0001, Zhengwei Qi
FASE4
2021 AdSherlock: Efficient and Deployable Click Fraud Detection for Mobile Applications
abstract
Mobile advertising plays a vital role in the mobile app ecosystem. A major threat to the sustainability of this ecosystem is click fraud, i.e., ad clicks performed by malicious code or automatic bot problems. Existing click fraud detection approaches focus on analyzing the ad requests at the server side. However, such approaches may suffer from high false negatives since the detection can be easily circumvented, e.g., when the clicks are behind proxies or globally distributed. In this paper, we present AdSherlock, an efficient and deployable click fraud detection approach at the client side (inside the application) for mobile apps. AdSherlock splits the computation-intensive operations of click request identification into an offline procedure and an online procedure. In the offline procedure, AdSherlock generates both exact patterns and probabilistic patterns based on URL (Uniform Resource Locator) tokenization. These patterns are used in the online procedure for click request identification and further used for click fraud detection together with an ad request tree model. We implement a prototype of AdSherlock and evaluate its performance using real apps. The online detector is injected into the app executable archive through binary instrumentation. Results show that AdSherlock achieves higher click fraud detection accuracy compared with state of the art, with negligible runtime overhead.
Chenhong Cao, Yi Gao 0001, Mingyuan Xia 0001, Wei Dong 0001, Chun Chen 0001, Xue (Steve) Liu
IEEE Trans. Mob. Comput.4
2020 TimelyRep: Timing deterministic replay for Android web applications
abstract
Summary With the constantly growing and changing requirements of app users, web techniques are used in mobile application development for better cross‐platform compatibility and online update. As the embedded web contents gain complexity, debugging web apps become a critical demand. Web replay tools can record program inputs and reproduce the same execution for debugging and performance tuning. However, traditional replay approaches are largely intended for apps with desktop interaction methods (keyboard, mouse) and require modification to the browser, which limits their applicability in mobile platforms. In this paper, we develop TimelyRep, which provides deterministic record‐and‐replay as a software library, running on commodity Android. TimelyRep can be used for app development with unmodified Android devices and for production to collect faulty execution from users. Also, we propose an efficient replay timing control mechanism and achieve higher timing precision as facing higher event rate on touchscreen devices. TimelyRep also supports cross‐device replay and can replay logged event traces on different devices, which is useful for developers to reproduce user inputs on their own devices. We evaluate TimelyRep with real‐world web applications. The results show that TimelyRep is useful for recreating program bugs and maintaining low delays for touch‐intensive web games.
Yanqiang Liu, Fangge Yan, Mingyuan Xia 0001, Zhengwei Qi, Xue (Steve) Liu
Softw. Test. Verification Reliab.3
2019 Systematically Testing and Diagnosing Responsiveness for Android Apps
abstract
App responsiveness is the most intuitive interpretation of app performance from user's perspective. Traditional performance profilers only focus on one kind of program activities (e.g., CPU profiling), while the cause for slow responsiveness is diverse or even due to the joint effect of multiple kinds. Also, various test configurations, such as device hardware and wireless connectivity can have dramatic impact on particular program activities and indirectly affect app responsiveness. Conventional mobile testing lacks mechanisms to reveal configuration-sensitive bugs. In this paper, we propose AppSPIN, a tool to automatically diagnose app responsiveness bugs and systematically explore configuration-sensitive bugs. AppSPIN instruments the app to collect program events and UI responsiveness. The instrumented app is exercised with automated monkey testers and AppSPIN correlates excessive and lengthy program events with bad responsiveness detected at runtime. The diagnosis process also synthesizes the major resource bottleneck for the app. After one test run, AppSPIN automatically alters the test configuration to with most bottlenecked resource to further explore responsiveness bugs happened only with particular test configurations. Our preliminary experiments with 30 real-world apps show that AppSPIN can detect 123 responsiveness bugs and successfully diagnose the cause for 87% cases, within an average of 15-minute test time. Also with altered test configurations, AppSPIN uncovers a notable number of new bugs within four extra test runs.
Wenhua Zhao, Zhenkai Ding, Mingyuan Xia 0001, Zhengwei Qi
ICSME3
2018 Reproducible Interference-Aware Mobile Testing
abstract
Mobile apps are born to work in an environment with ever-changing network connectivity, random hardware interruption, unanticipated task switches, etc. However, such interference cases are often oblivious in traditional mobile testing but happen frequently and sophisticatedly in the field, causing various robustness, responsiveness and consistency problems. In this paper, we propose JazzDroid to introduce interference to mobile testing. JazzDroid adopts a gray-box approach to instrument apps at binary level such that interference logic is inlined with app execution and can be triggered to effectively affect normal execution. Then, JazzDroid repeatedly orchestrates the instrumented app through app developers' existing tests and continuously randomizes interference on the fly to reveal possible faulty executions. Upon discovering problems, JazzDroid generates a test script with the user inputs from developers' tests and the interference injected for developers to reproduce the problems. At a high level, JazzDroid can be seamlessly integrated into app developers' testing procedures, detecting more problems from existing tests. We implement JazzDroid to function on unmodified apps directly from app markets and interface with de facto industrial testing toolchain. JazzDroid improves mobile testing by discovering 6x more problems, including crashes, functional bugs, UI consistency issues and common bug patterns that fail numerous apps.
Weilun Xiong, Shihao Chen, Mingyuan Xia 0001, Zhengwei Qi
ICSME4
2017 Every pixel counts: Fine-grained UI rendering analysis for mobile applications
abstract
For mobile apps, user-perceived delays are critical for user satisfaction. According to our measurement, long delays are commonly caused by network and storage I/O operations while short delays are mainly caused by UI rendering. Short delays are not uncommon, which account for 55.3% in our measurement cases. Previous app performance studies have largely focused on I/O operations but the understanding of UI rendering impact is limited. In this work, we propose DRAW, a system that performs two UI rendering analyses to help app developers pinpoint rendering problems and resolve short delays. The first analysis outlines the wasted rendering time on invisible or covered UI components, namely the overdraw problem. The second analysis is to identify the responsible UI components and rendering operations that cause overall low rendering efficiency. We implement DRAW on Android and apply it to study 1,158 real-world Android apps. Results show that DRAW is helpful as it can pinpoint the responsible UI components and specific rendering operations. Four concrete case studies of real-world apps are further presented to show how DRAW can help developers improve the UI rendering performance of their apps.
Yi Gao 0001, Haocheng Huang, Wei Dong 0001, Mingyuan Xia 0001, Xue (Steve) Liu, Jiajun Bu
INFOCOM6
2017 Whom to Blame? Automatic Diagnosis of Performance Bottlenecks on Smartphones
abstract
The past decade has witnessed a tremendous growth in the variety and complexity of mobile applications (apps). Although considerable amount of efforts have been spent to improve app performance, smartphones nowadays still face many performance challenges. We discover that the resource contention of multiple running apps, caused by resource bottleneck(s), is a key factor that affects the smartphone performance. In this paper, we present APB, an Automatic tool that detects Performance issues caused by resource Bottleneck(s) on commodity Android smartphones. APB employs an innovative bottleneck-hypersurface model to quantify performance issues given a specific system state. Then, based on the model, APB identifies a list of apps that contribute most to the resource contention, which can well inform the end user to take action such as killing background apps to resolve the performance issue. We implement APB on commodity Android platforms and widely evaluate its effectiveness with real user studies. Results show that APB outperforms three baseline approaches and helps users to improve smartphone performance by 10 to 67 percent, with less than 1 percent runtime overhead.
Yi Gao 0001, Wei Dong 0001, Haocheng Huang, Jiajun Bu, Chun Chen 0001, Mingyuan Xia 0001, Xue (Steve) Liu
IEEE Trans. Mob. Comput.6
2015 A Tale of Two Erasure Codes in HDFS
Mingyuan Xia 0001, Mohit Saxena, Mario Blaum, David Pease
FAST1
2015 Distributed Optimal Datacenter Bandwidth Allocation for Dynamic Adaptive Video Streaming
abstract
Video streaming systems such as YouTube and Netflix are usually supported by the content delivery networks and datacenters that can consume many megawatts of power. Most existing works independently study the issues of improving quality of experience (QoE) for viewers and reducing the cost and emissions associated with the enormous energy usage of datacenters. By contrast, this paper addresses them both, and jointly optimizes the QoE, the energy cost and emissions by intelligently allocating datacenter bandwidth among different client groups. Specially, we propose a distributed algorithm for achieving the optimal bandwidth allocation. The algorithm novelly decomposes the optimization process into separate ones, which are solved iteratively across datacenters and clients. We demonstrate its convergence by both theoretical proof and experimental validation. The experimental results show that the proposed algorithm converges very fast and achieves much better QoE-cost balance than existing approaches.
Fanxin Kong, Xingjian Lu, Mingyuan Xia 0001, Xue (Steve) Liu, Haibing Guan
ACM Multimedia3
2015 Effective Real-Time Android Application Auditing
abstract
Mobile applications can access both sensitive personal data and the network, giving rise to threats of data leaks. App auditing is a fundamental program analysis task to reveal such leaks. Currently, static analysis is the de facto technique which exhaustively examines all data flows and pinpoints problematic ones. However, static analysis generates false alarms for being over-estimated and requires minutes or even hours to examine a real app. These shortcomings greatly limit the usability of automatic app auditing. To overcome these limitations, we design AppAudit that relies on the synergy of static and dynamic analysis to provide effective real-time app auditing. AppAudit embodies a novel dynamic analysis that can simulate the execution of part of the program and perform customized checks at each program state. AppAudit utilizes this to prune false positives of an efficient but over-estimating static analysis. Overall, AppAudit makes app auditing useful for app market operators, app developers and mobile end users, to reveal data leaks effectively and efficiently. We apply AppAudit to more than 1,000 known malware and 400 real apps from various markets. Overall, AppAudit reports comparative number of true data leaks and eliminates all false positives, while being 8.3x faster and using 90% less memory compared to existing approaches. AppAudit also uncovers 30 data leaks in real apps. Our further study reveals the common patterns behind these leaks: 1) most leaks are caused by 3rd-party advertising modules; 2) most data are leaked with simple unencrypted HTTP requests. We believe AppAudit serves as an effective tool to identify data-leaking apps and provides implications to design promising runtime techniques against data leaks.
Mingyuan Xia 0001, Lu Gong, Yuanhao Lyu, Zhengwei Qi, Xue (Steve) Liu
IEEE Symposium on Security and Privacy1
2014 Domo: Passive Per-Packet Delay Tomography in Wireless Ad-hoc Networks
abstract
In multi-hop wireless ad-hoc networks, packet delivery delay is one of the most important performance metrics. While a lot of research efforts have been spent on measuring and optimizing the end-to-end delay performance, there usually lack accurate and lightweight methods for decomposing the end-to-end delay into the per-hop delay for each packet. Knowledge on the per-hop per-packet delay can greatly improve the network visibility and facilitate network measurement and management. In this paper, we propose Domo, a passive, lightweight and accurate delay tomography approach to decomposing the packet end-to-end delay into each hop. The basic idea is to formulate the problem into a set of optimization problems by carefully considering the constraints among various timing quantities. At the network side, Domo attaches a small overhead to each packet for constructing constraints of the optimization problems. At the PC side, Domo employs semi-definite relaxation and several other methods to efficiently solve the optimization problems. We implement Domo and evaluate its performance extensively using large-scale simulations. Results show that Domo significantly outperforms two existing methods, nearly tripling the accuracy of the state-of-the-art.
Yi Gao 0001, Wei Dong 0001, Chun Chen 0001, Jiajun Bu, Mingyuan Xia 0001, Xue (Steve) Liu, Xianghua Xu
ICDCS6
2014 Taming IO Spikes in Enterprise and Campus VM Deployment
abstract
Enterprises and campuses have widely employed virtualization to improve resource utilization and save capital costs. The virtualized storage infrastructure usually comprises a pool of storage nodes (disks, RAID groups and storage servers), each consolidating a number of virtual machines (VMs). The IO workload on each storage node is highly bursty where a few overloaded periods, namely IO spikes, incur lags and greatly affect VM performance. In this paper, we design VMpart, a system that automatically reconfigures VM deployment to remove IO spikes from a deployed VM environment. Firstly, VMpart collects IO parameters and identifies IO spikes for deployed VMs during production periods. Secondly, during the maintenance phase, VMpart divides every VM based on its disk partitions and reconfigures the system with a fine-grained optimized partition deployment for better load balancing. We set up a representative VM environment with desktop and server VMs used in enterprises and campuses to evaluate VMpart. Our experiments show that the optimized deployment reduces the average booting time of desktop VMs during boot storm from 9% to 28%, improves server VM throughput by more than 60% and reduces VM storage migration time by about 57%.
Mingyuan Xia 0001, Pin Zhou, David Pease, Xue (Steve) Liu
SYSTOR1
2014 Data Loss and Reconstruction in Wireless Sensor Networks
abstract
Reconstructing the environment by sensory data is a fundamental operation for understanding the physical world in depth. A lot of basic scientific work (e.g., nature discovery, organic evolution) heavily relies on the accuracy of environment reconstruction. However, data loss in wireless sensor networks is common and has its special patterns due to noise, collision, unreliable link, and unexpected damage, which greatly reduces the reconstruction accuracy. Existing interpolation methods do not consider these patterns and thus fail to provide a satisfactory accuracy when the missing data rate becomes large. To address this problem, this paper proposes a novel approach based on compressive sensing to reconstruct the massive missing data. Firstly, we analyze the real sensory data from Intel Indoor, GreenOrbs, and Ocean Sense projects. They all exhibit the features of low-rank structure, spatial similarity, temporal stability and multi-attribute correlation. Motivated by these observations, we then develop an environmental space time improved compressive sensing (ESTI-CS) algorithm with a multi-attribute assistant (MAA) component for data reconstruction. Finally, extensive simulation results on real sensory datasets show that the proposed approach significantly outperforms existing solutions in terms of reconstruction accuracy.
Linghe Kong, Mingyuan Xia 0001, Xiao-Yang Liu, Guangshuo Chen, Yu Gu 0001, Min-You Wu, Xue (Steve) Liu
IEEE Trans. Parallel Distributed Syst.2
2013 Data loss and reconstruction in sensor networks
abstract
Reconstructing the environment in cyber space by sensory data is a fundamental operation for understanding the physical world in depth. A lot of basic scientific work (e.g., nature discovery, organic evolution) heavily relies on the accuracy of environment reconstruction. However, data loss in wireless sensor networks is common and has its special patterns due to noise, collision, unreliable link, and unexpected damage, which greatly reduces the accuracy of reconstruction. Existing interpolation methods do not consider these patterns and thus fail to provide a satisfactory accuracy when missing data become large. To address this problem, this paper proposes a novel approach based on compressive sensing to reconstruct the massive missing data. Firstly, we analyze the real sensory data from Intel Indoor, GreenOrbs, and Ocean Sense projects. They all exhibit the features of spatial correlation, temporal stability and low-rank structure. Motivated by these observations, we then develop an environmental space time improved compressive sensing (ESTICS) algorithm to optimize the missing data estimation. Finally, the extensive experiments with real-world sensory data shows that the proposed approach significantly outperforms existing solutions in terms of reconstruction accuracy. Typically, ESTICS can successfully reconstruct the environment with less than 20% error in face of 90% missing data.
Linghe Kong, Mingyuan Xia 0001, Xiao-Yang Liu, Min-You Wu, Xue (Steve) Liu
INFOCOM2
2010 Real-time Enhancement for Xen Hypervisor
abstract
System virtualization, which provides good isolation, is now widely used in server consolidation. Meanwhile, one of the hot topics in this field is to extend virtualization for embedded systems. However, current popular virtualization platforms do not support real-time operating systems such as embedded Linux well because the platform is not real-time ware, which will bring low-performance I/O and high scheduling latency. The goal of this paper is to optimize the Xen virtualization platform to be real-time operating system friendly. We improve two aspects of the Xen virtualization platform. First, we improve the xen scheduler to manage the scheduling latency and response time of the real-time operating system. Second, we import multiple real-time operating systems balancing method. Our experiment demonstrates that our enhancement to the Xen virtualization platform support real-time operating system well and the improvement to the real-time performance is about 20%.
Peijie Yu, Mingyuan Xia 0001, Qian Lin 0002, Shang Gao 0009, Zhengwei Qi, Kai Chen 0006, Haibing Guan
EUC2
2010 Enhanced Privilege Separation for Commodity Software on Virtualized Platform
abstract
Conventional privilege separation can effectively reduce the TCB size by granting privilege to only the privileged compartments. However, since they this approach relies on process isolation to ensure security assurance, malware exploiting against kernel components can easily compromise. Meanwhile, the frequent inter-process communications between separated processes inevitably incur notable overhead. To ameliorate these problems, we propose to perform privilege separation without partitioning application into two processes. Instead, we leverage virtualization to enforce the isolation of sensitive portions from other untrusted code. The virtual machine monitor intercepts all the code context switches transparently without requiring the application to explicitly use IPC as privilege context transition. We have implemented a prototype of our system, named Coir, based on commodity hypervisor Xen. Evaluation of our prototype includes a real-world remote control application, which is partitioned and protected in Coir-enabled hypervisor on unmodified Windows XP. We discuss the isolation strength as well as the performance penalty of our system based on the practical case.
Mingyuan Xia 0001, Qian Lin 0002, Zhengwei Qi, Haibing Guan
ICPADS1