Hua Zhong 0001

dblp:65/569-1 · DBLP profile ↗
← Back
50ranked-venue papers
4as first author
12since 2021 · last 2024
0000-0002-8535-8225ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 31 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5Systems, architecture and hardware · 2 · 1 since 2021Artificial intelligence and machine learning · 1Computer networks · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 Differential Optimization Testing of Gremlin-Based Graph Database Systems
abstract
Graph database systems (GDBs) allow efficiently creating, modifying, and retrieving graph data in a graph database. To accelerate graph queries, GDBs usually adopt various and complex optimization strategies. However, incorrect optimizations in GDBs can introduce optimization bugs, which cause a graph query to compute an incorrect query result, e.g., omitting a vertex in a graph database. In this paper, we propose Differential Optimization Testing (DOT), an effective and automated approach to detect optimization bugs in GDBs that adopt Gremlin as their query language. The main idea of DOT is that, given a Gremlin query$Q$, we execute it on the target GDB with two different optimization configurations and then verify whether they can compute the same query results for query$Q$. Any inconsistency between their query results indicates an optimization bug in the target GDB. To improve the efficiency of differential testing in DOT, we further propose an optimization-guided approach, aiming to explore more optimization strategies and more graph database features. We evaluate DOT on six popular and widely-used GDBs, i.e., Neo4j, OrientDB, JanusGraph, HugeGraph, TinkerGraph, and ArcadeDB. In total, we have found 28 unique optimization bugs, 16 of which have been confirmed as previously-unknown bugs.
Yingying Zheng, Wensheng Dou, Ziyu Cui, Jiansen Song, Ziyue Cheng, Wei Wang 0049, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
ICST9
2023 Detecting Isolation Bugs via Transaction Oracle Construction
abstract
Transactions are used to maintain the data integrity of databases, and have become an indispensable feature in modern Database Management Systems (DBMSs). Despite extensive efforts in testing DBMSs and verifying transaction processing mechanisms, isolation bugs still exist in widely-used DBMSs when these DBMSs violate their claimed transaction isolation levels. Isolation bugs can cause severe consequences, e.g., incorrect query results and database states. In this paper, we propose a novel transaction testing approach, Transaction oracle construction (Troc), to automatically detect isolation bugs in DBMSs. The core idea of Troc is to decouple a transaction into independent statements, and execute them on their own database views, which are constructed under the guidance of the claimed transaction isolation level. Any divergence between the actual transaction execution and the independent statement execution indicates an isolation bug. We implement and evaluate Troc on three widely-used DBMSs, i.e., MySQL, MariaDB, and TiDB. We have detected 5 previously-unknown isolation bugs in the latest versions of these DBMSs.
Wensheng Dou, Ziyu Cui, Qianwang Dai, Jiansen Song, Dong Wang 0048, Yu Gao 0002, Wei Wang 0049, Jun Wei 0001, Hanmo Wang, Hua Zhong 0001, Tao Huang 0001
ICSE11
2023 Coverage Guided Fault Injection for Cloud Systems
abstract
To support high reliability and availability, modern cloud systems are designed to be resilient to node crashes and reboots. That is, a cloud system should gracefully recover from node crashes/reboots and continue to function. However, node crashes/reboots that occur under special timing can trigger crash recovery bugs that lie in incorrect crash recovery protocols and their implementations. To ensure that a cloud system is free from crash recovery bugs, some fault injection approaches have been proposed to test whether a cloud system can correctly recover from various crash scenarios. These approaches are not effective in exploring the huge crash scenario space without developers' knowledge. In this paper, we propose Crash Fuzz, a fault injection testing approach that can effectively test crash recovery behaviors and reveal crash recovery bugs in cloud systems. CrashFuzz mutates the combinations of possible node crashes and reboots according to runtime feedbacks, and prioritizes the combinations that are prone to increase code coverage and trigger crash recovery bugs for smart exploration. We have implemented CrashFuzz and evaluated it on three popular open-source cloud systems, i.e., ZooKeeper, HDFS and HBase. CrashFuzz has detected 4 unknown bugs and 1 known bug. Compared with other fault injection approaches, CrashFuzz can detect more crash recovery bugs and achieve higher code coverage.
Yu Gao 0002, Wensheng Dou, Dong Wang 0048, Wenhan Feng, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
ICSE6
2023 Testing Database Systems via Differential Query Execution
abstract
Database Management Systems (DBMSs) provide efficient data retrieval and manipulation for many applications through Structured Query Language (SQL). Incorrect implementations of DBMSs can result in logic bugs, which cause SELECT queries to fetch incorrect results, or UPDATE and DELETE queries to generate incorrect database states. Existing approaches mainly focus on detecting logic bugs in SELECT queries. However, logic bugs in UPDATE and DELETE queries have not been tackled. In this paper, we propose a novel and general approach, which we have termed Differential Query Execution (DQE), to detect logic bugs in SELECT, UPDATE and DELETE queries of DBMSs. The core idea of DQE is that different SQL queries with the same predicate usually access the same rows in a database. For example, a row updated by an UPDATE query with a predicate φ should also be fetched by a SELECT query with the same predicate φ, If not, a logic bug is revealed in the target DBMS. To evaluate the effectiveness and generality of DQE, we apply DQE on five production-level DBMSs, i.e., MySQL, MariaDB, TiDB, CockroachDB and SQLite. In total, we have detected 50 unique bugs in these DBMSs, 41 of which have been confirmed, and 11 have been fixed. We expect that the simplicity and generality of DQE can greatly improve the reliability of DBMSs.
Jiansen Song, Wensheng Dou, Ziyu Cui, Qianwang Dai, Wei Wang 0049, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
ICSE7
2023 LPW: an efficient data-aware cache replacement strategy for Apache Spark
Shuping Ji, Hua Zhong 0001, Wei Wang 0049, Lijie Xu, Jun Wei 0001, Tao Huang 0001
Sci. China Inf. Sci.3
2023 Hydra: Deadline-Aware and Efficiency-Oriented Scheduling for Deep Learning Jobs on Heterogeneous GPUs
abstract
With the rapid proliferation of deep learning (DL) jobs running on heterogeneous GPUs, scheduling DL jobs to meet various scheduling requirements, such as meeting deadlines and reducing job completion time (JCT), is critical. Unfortunately, existing efficiency-oriented and deadline-aware efforts are still rudimentary. They lack the capability of scheduling jobs to meet deadline requirements while reducing total JCT, especially when the jobs have various execution times on heterogeneous GPUs. Therefore, we present Hydra, a novel quantitative cost comparison approach, to address this scheduling issue. Here, the cost represents the total JCT plus a dynamic penalty calculated from the total tardiness (i.e., the delay time of exceeding the deadline) of all jobs. Hydra adopts a sampling approach that exploits the inherent iterative periodicity of DL jobs to estimate job execution times accurately on heterogeneous GPUs. Then, Hydra considers various combinations of job sequences and GPUs to obtain the minimized cost by leveraging an efficient branch-and-bound algorithm. Finally, the results of evaluation experiments on Alibaba traces show that Hydra can reduce total tardiness by 85.8% while reducing total JCT as much as possible, compared with state-of-the-art efforts.
Heng Wu 0001, Yuanjia Xu, Yuewen Wu, Hua Zhong 0001, Wenbo Zhang 0006
IEEE Trans. Computers5
2022 Focusing Nonparallel-Track Bistatic SAR Data Using Modified Frequency Extended Nonlinear Chirp Scaling
abstract
Unsynchronization of the separate transmit–receive beams makes it a challenging task to obtain high-quality images for nonparallel-track bistatic synthetic aperture radar (NP-BiSAR). To accommodate this issue, we propose an imaging configuration where the receiver’s beam follows the transmitter’s one actively by adjusting the squint angle of receiver. And a frequency extended nonlinear chirp scaling (FENLCS) algorithm is modified to cope with new effects introduced by the innovative configuration, which is based on an improved quadratic ellipse model. Based on the new model, some innovative improvements on image formation are made, including a residual azimuth-dependent high-order range cell migration correction (ADH-RCMC) and a rederived FENLCS that takes highly varying Doppler centroid into consideration, which contribute to better imaging quality. Simulation results validate the effectiveness of the proposed configuration and algorithm.
Shiping Li, Hua Zhong 0001, Cunliang Yang, Huina Song, Ronghua Zhao
IEEE Geosci. Remote. Sens. Lett.2
2022 Detection of Lake Shoreline Based on Modified RSF Model Combined With Edge Energy and Global Energy for SAR Images
abstract
Lake shoreline detection plays an important role in hydrological structure analysis and urban ecology governance but is a challenging task in synthetic aperture radar (SAR) image interpretation. Due to the complex shoreline environment, the preservation of weak boundaries and fitting of global information in large-scale SAR images deserve further research. Thus, this letter proposes a novel coarse-to-fine lake shoreline detection approach for SAR images based on modified region-scalable fitting (RSF) model combined with edge energy and global energy. In this approach, SAR images are despeckled by the block-matching 3-D (BM3D) filter. Then a novel energy term based on Laplacian of Gaussian (LoG) operator and ratio of exponentially weighted averages (ROEWA) operator is constructed to accurately locate the boundary and reduce false boundary. Additionally, the global energy term is adopted to fit the global information well. The experimental results based on real data demonstrate that the proposed approach has a stronger ability to maintain weak edges compared with RSF model, which exhibits better effectiveness and reliability.
Huina Song, Junliang Xie, Yingcheng Ding, Hua Zhong 0001
IEEE Geosci. Remote. Sens. Lett.7
2022 A Fast Phase Optimization Approach of Distributed Scatterer for Multitemporal SAR Data Based on Gauss-Seidel Method
abstract
Distributed scatterer (DS) interferometric synthetic aperture can retrieve maximum available information by jointly processing persistent scatterers (PSs) and DSs. Unlike PSs, DSs are vulnerable to temporal, geometrical, and volumetric decorrelation. The phase optimization of DSs is essential for reliable parameter estimation. However, the preprocessing of DSs is very computationally expensive, and this drawback limits its engineering application to some degree. To improve computational efficiency, a fast scheme for reliable phase optimization of DSs is proposed based on the coherence-weighted model in this letter. The Gauss–Seidel iteration, having the advantages of fast convergence rate and small data memory, is introduced to solve the adopted phase optimization model. Experiments both on simulated data and real data are used to verify the reliability and efficiency of the presented method in this letter.
Huina Song, Hua Zhong 0001
IEEE Geosci. Remote. Sens. Lett.6
2021 A Curvature-Based Saliency Method for Ship Detection in SAR Images
abstract
According to the theory and interpretation method of information geometry, this work presents a novel curvature-based saliency method for ship detection in synthetic aperture radar (SAR) images. First, the nonlinear anisotropic diffusive process has been adopted to eliminate clutter while preserving the local ship target structure in SAR images. Then, a novel curvature-based saliency method for super-pixels in the filtered image is presented, which is used to exploit the microstructure feature of statistical manifold. Finally, a statistical classification method is used to realize the location of targets. The experimental results on real SAR images show that the proposed method can achieve good performance.
Meng Yang 0003, Chunsheng Guo, Hua Zhong 0001, Haibing Yin
IEEE Geosci. Remote. Sens. Lett.3
2021 An Improved Imaging Algorithm for High-Resolution and Highly Squinted One-Stationary Bistatic SAR Using Extended Nonlinear Chirp Scaling Based on Equi-Sum of Bistatic Ranges
abstract
The linear range walk correction (LRWC) and the inherent azimuth-variant geometric configuration in the case of one-stationary bistatic synthetic aperture radar (OS-BiSAR) imaging produce 2-D variant range cell migrations (RCMs) and azimuth-dependent Doppler parameters, which makes it more difficult to obtain high-quality image for high-resolution and highly squinted OS-BiSAR. To accommodate these issues, an improved extended nonlinear chirp scaling (ENLCS) algorithm is developed in this letter. An analytic model based on the equi-sum of bistatic ranges (ESBR) after the RCM correction is proposed to reveal the azimuth-variant property of the azimuth distributed targets. Based on this innovative model, analytic coefficients of the ENLCS are derived, and high-quality image formation for OS-BiSAR is accomplished. Simulations are conducted to demonstrate the validity of the proposed algorithm.
Hua Zhong 0001, Ronghua Zhao, Huina Song, Meng Yang 0003, Zongqi Ye
IEEE Geosci. Remote. Sens. Lett.1
2021 X-Check: Improving Effectiveness and Efficiency of Cross-Browser Issues Detection for JavaScript-Based Web Applications
abstract
Web 2.0 application based on JavaScript is a wide-spread application domain today as it delivers rich, interactive user experiences. However, with the increasing number of browsers and platforms on which the applications can be executed, cross-browser incompatibilities (XBIs) are becoming a serious problem for organizations to develop modern JavaScript-based Web applications. Although lots of XBIs detection techniques have been proposed, there are still some limitations: 1) existing techniques are prone to generating certain false positives/negatives that result from the fact that they ignore non-deterministic events (e.g., timer, asynchronous request/response) inside the browser; 2) detection process is inefficient, as the same elements located in different pages will be repeatedly checked even if they stay unchanged after an event is triggered. Leveraging existing record/replay technique, we proposed X-Check, a novel cross-browser testing technique, which supports automated XBIs detection effectively. To improve the efficiency of XBIs detection, this paper further designed an incremental detection algorithm by only checking DOM-mutated and layout-changed nodes. Our empirical evaluation shows that X-Check is effective and efficient. For the selected 21 real-world Web applications, it identifies XBIs with a fairly high precision (83 percent) and recall (93 percent), and improves the performance of XBIs detection about 5.79 times compared to its non-optimized version.
Guoquan Wu, Meimei He, Wei Chen 0018, Jun Wei 0001, Hua Zhong 0001
IEEE Trans. Serv. Comput.5
2020 Detecting cache-related bugs in Spark applications
abstract
Apache Spark has been widely used to build big data applications. Spark utilizes the abstraction of Resilient Distributed Dataset (RDD) to store and retrieve large-scale data. To reduce duplicate computation of an RDD, Spark can cache the RDD in memory and then reuse it later, thus improving performance. Spark relies on application developers to enforce caching decisions by using persist() and unpersist() APIs, e.g., which RDD is persisted and when the RDD is persisted / unpersisted. Incorrect RDD caching decisions can cause duplicate computations, or waste precious memory resource, thus introducing serious performance degradation in Spark applications. In this paper, we propose CacheCheck, to automatically detect cache-related bugs in Spark applications. We summarize six cache-related bug patterns in Spark applications, and then dynamically detect cache-related bugs by analyzing the execution traces of Spark applications. We evaluate CacheCheck on six real-world Spark applications. The experimental result shows that CacheCheck detects 72 previously unknown cache-related bugs, and 28 of them have been fixed by developers.
Dong Wang 0048, Yu Gao 0002, Wensheng Dou, Lijie Xu, Wei Wang 0049, Jun Wei 0001, Hua Zhong 0001
ISSTA9
2019 Focus Improvement for Highly Squinted One-Stationary BISAR Imaging Based On A Range Equivalent Model
abstract
Owing to the particular configuration of the bistatic synthetic aperture radar with stationary transmitter (ST-BISAR), image formation for such kind of radar is a great challenge for producing well-focused images. In this paper, an imaging algorithm based on the range equivalent ellipse model is proposed to deal with these issues. The construction of range equivalent model is the key for both the range processing and the equalization of azimuth-dependent DFMR. Analyzing the result of range processing combining linear range walk correction (LRWC), bulk range cell migration correction (RCMC) and secondary range compression (SRC), a range equivalent ellipse model is deduced to reveal the complicated relationship among azimuth variant echoes accurately, and helps derive the analytical expressions of the Doppler phases. Based on these results, an extended nonlinear chirp scaling (ENLCS) is utilized to equalize the space-variant ST-BISAR data, which contributes to a high-performance image formation. The effectiveness of the proposed approach is validated and demonstrated via simulations.
Hua Zhong 0001, Guangyong Zheng, Ronghua Zhao, Zongqi Ye, Guojin Chen, Aibo Yan
IGARSS1
2018 HW3C: A Heuristic based Workload Classification and Cloud Configuration Approach for Big Data Analytics
abstract
It is a big challenge to pick up the best cloud configuration for recurring big data analytics jobs running in clouds. Prior efforts may get in a sub-optimal configuration due to a broad spectrum of cloud configurations with a few test runs, such as CherryPick. We present HW3C which is a heuristic based workload classification and cloud configuration system for big data analytics jobs, our insight is classifying a job by comparing its resource preference and usage informantion with other jobs, and then using heuristic rules to distinguish bad samples from good ones in Bayesian Optimization algorithm. Our experiments on HiBench and SparkBench in Aliyun ECS show that the performance of job had been improved by 53% in average comparing with CherryPick, meanwhile the resource cost had been reduced by 40% in average.
Yuewen Wu, Heng Wu 0001, Wenbo Zhang 0006, Yuanjia Xu, Jun Wei 0001, Hua Zhong 0001
Internetware6
2018 Focus High-Resolution Highly Squint SAR Data Using Azimuth-Variant Residual RCMC and Extended Nonlinear Chirp Scaling Based on a New Circle Model
abstract
The combination of linear range walk correction and keystone transform is a good choice to focus high-resolution highly squint synthetic aperture radar (SAR) data because it is an effective way to remove linear range cell migration (RCM) completely and mitigate range-azimuth coupling. However, the results of this kind of imaging algorithm produce 2-D-variant residual RCM and variant-dependence Doppler phases. To obtain high-quality SAR image, an improved imaging algorithm using an azimuth-variant residual RCM correction (RCMC) and an extended nonlinear chirp scaling (ENLCS) is proposed in this letter. A new circle model is constructed to analyze the azimuth-variant properties of the residual high-order RCM and the Doppler phases. Based on this circle model, an azimuth-variant residual RCMC is implemented by multiplying a fourth-order phase function, and an improved ENLCS is derived to accomplish the azimuth equalization for azimuth compression. Simulation results validate the excellent performance of the proposed algorithm.
Hua Zhong 0001, Yuliang Chang, Erxiao Liu, Xianghong Tang, Jianwu Zhang
IEEE Geosci. Remote. Sens. Lett.1
2017 AppCheck: A Crowdsourced Testing Service for Android Applications
abstract
It is well known that the fragmentation of Android ecosystem has caused severe compatibility issues. Therefore, for Android apps, cross-platform testing (the apps must be tested on a multitude of devices and operating system versions) is particularly important to assure their quality. Although lots of cross-platform testing techniques have been proposed, there are still some limitations: 1) it is time-consuming and error-prone to encode platform-agnostic tests manually, 2) test scripts generated by existing record/replay techniques are brittle and will break when replayed on different platforms, 3) Developers, and even test vendors have not equipped some special Android devices. As a result, apps have not been tested sufficiently, leading to many compatibility issues after releasing. To address these limitations, this paper proposes AppCheck, a crowdsourced testing service for Android apps. To generate tests that will explore different behavior of the app automatically, AppCheck crowdsources event trace collection over the Internet, and various touch events will be captured when real users interact with the app. The collected event traces are then transformed into platform-agnostic test scripts, and directly replayed on the devices of real users. During the replay, various data (e.g., screenshots and layout information) will be extracted to identify compatibility issues. Our empirical evaluation shows that AppCheck is effective and improves the state of the art.
Guoquan Wu, Yuzhong Cao, Wei Chen 0018, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
ICWS5
2017 SpreadCluster: recovering versioned spreadsheets through similarity-based clustering
abstract
Version information plays an important role in spreadsheet understanding, maintaining and quality improving. However, end users rarely use version control tools to document spreadsheets' version information. Thus, the spreadsheets' version information is missing, and different versions of a spreadsheet coexist as individual and similar spreadsheets. Existing approaches try to recover spreadsheet version information through clustering these similar spreadsheets based on spreadsheet filenames or related email conversation. However, the applicability and accuracy of existing clustering approaches are limited due to the necessary information (e.g., filenames and email conversation) is usually missing. We inspected the versioned spreadsheets in VEnron, which is extracted from the Enron Corporation. In VEnron, the different versions of a spreadsheet are clustered into an evolution group. We observed that the versioned spreadsheets in each evolution group exhibit certain common features (e.g., similar table headers and worksheet names). Based on this observation, we proposed an automatic clustering algorithm, SpreadCluster. SpreadCluster learns the criteria of features from the versioned spreadsheets in VEnron, and then automatically clusters spreadsheets with the similar features into the same evolution group. We applied SpreadCluster on all spreadsheets in the Enron corpus. The evaluation result shows that SpreadCluster could cluster spreadsheets with higher precision (78.5% vs. 59.8%) and recall rate (70.7% vs. 48.7%) than the filename-based approach used by VEnron. Based on the clustering result by SpreadCluster, we further created a new versioned spreadsheet corpus VEnron2, which is much bigger than VEnron (12,254 vs. 7,294 spreadsheets). We also applied SpreadCluster on the other two spreadsheet corpora FUSE and EUSES. The results show that SpreadCluster can cluster the versioned spreadsheets in these two corpora with high precision (91.0% and 79.8%).
Wensheng Dou, Chushu Gao, Jie Wang 0035, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
MSR6
2017 ReSeer: Efficient search-based replay for multiprocessor virtual machines
Tao Wang 0030, Wenbo Zhang 0006, Jun Wei 0001, Hua Zhong 0001
J. Syst. Softw.6
2017 Focusing Nonparallel-Track Bistatic SAR Data Using Extended Nonlinear Chirp Scaling Algorithm Based on a Quadratic Ellipse Model
abstract
Image formation for bistatic synthetic aperture radar data with nonparallel tracks is a challenging task due to the existence of the spatial variance of range cell migration (RCM) and azimuth frequency modulation (FM) rate. In this letter, an extended nonlinear chirp scaling (ENLCS) algorithm based on a quadratic ellipse model with two motion platform parameters is proposed for this configuration. For range processing, a method combining linear range walk correction and keystone transform is adopted to remove linear RCM completely and mitigate range-azimuth cross coupling. For azimuth focusing, a quadratic ellipse model using the two motion platform parameters is established to depict the azimuth-variant properties of the azimuth FM rate, and its accurate high-order approximation is derived as well. Investigations of the quadratic phase error show that great improvement can be made by this model. Following which, coefficients of the ENLCS are derived to accomplish the azimuth equalization. Simulation results validate the effectiveness of the algorithm proposed by this letter.
Hua Zhong 0001, Minhong Sun
IEEE Geosci. Remote. Sens. Lett.1
2016 Determine Configuration Entry Correlations for Web Application Systems
abstract
Web application systems, comprising of heterogeneous and loosely coupled components, are usually highly-configurable due to the large number of configuration entries scattering in the components. The dependencies between components lead their entries correlate to one another, which makes the system deployment and migration daunting and error-prone. For two correlated entries, changing value of one entry requires the value change of the other. Otherwise, some implied constraints would be violated and the system failure will occur. Keeping track of entry correlations, which is essential to system reliabilities, is not a simple work as it often crosses products and requires in-depth domain knowledge. This paper proposes a method to automate the process of determining entry correlations. The method first narrows down the exploring scale to those frequently-set entries based on crawled sample data. Then, it generates a correlation score for each entry pair, which is calculated according to entry names, values and inferred types. Thirdly, a set of heuristics are provided to determine a candidate set of the likely correlations. Finally, a rank-ordered list of entry correlations is output so that system administrators can consult it to check system configuration systematically. Based on the method, we implement a tool, Correlation Explorer, and make experiments and evaluations with some real world systems. The result shows that Correlation Explorer is effective in finding a large portion of entry correlations.
Wei Chen 0018, Heng Wu 0001, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
COMPSAC4
2016 Crawling hidden objects with kNN queries
abstract
With rapidly growing popularity, Location Based Services (LBS), e.g., Google Maps, Yahoo Local, WeChat, FourSquare, etc., started offering web-based search features that resemble a kNN query interface. Specifically, for a user-specified query location q, these websites extract from the objects in their backend database the top-k nearest neighbors to q and return these k objects to the user through the web interface. Here k is often a small value like 50 or 100. For example, McDonald [1] returns the top 25 nearest restaurants for a user-specified location through its locations search webpage.
Zhiguo Gong, Nan Zhang 0004, Tao Huang 0001, Hua Zhong 0001, Jun Wei 0001
ICDE5
2016 X-Check: A Novel Cross-Browser Testing Service Based on Record/Replay
abstract
With the advent of Web 2.0 application, and the increasing number of browsers and platforms on which the applications can be executed, cross-browser incompatibilities (XBIs) are becoming a serious problem for organizations to develop web-based software. Although some techniques and tools have been proposed to identify XBIs, they cannot assure the same execution when the application runs across different browsers as only explicit user activity is considered, and thus prone to generating both false positives and false negatives. To address this limitation, this paper describes X-Check, a platform that enables cross-browser testing as a service by leveraging record/replay technique. Comparing to existing techniques and tools, X-Check supports to detect cross-browser issues with high accuracy. It also provides useful support to developers for diagnosis and (eventually) elimination of XBIs. Our empirical evaluation shows that X-Check is effective, improves the state of the art.
Meimei He, Guoquan Wu, Hongyin Tang, Wei Chen 0018, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
ICWS6
2016 Generating test cases to expose concurrency bugs in Android applications
abstract
Mobile systems usually support an event-based model of concurrent programming. This model, although advantageous to maintain responsive user interfaces, may lead to subtle concurrency errors due to unforeseen threads interleaving coupled with non-deterministic reordering of asynchronous events. These bugs are very difficult to reproduce even by the same user action sequences that trigger them, due to the undetermined schedules of underlying events and threads. In this paper, we proposed RacerDroid, a novel technique that aims to expose concurrency bugs in android applications by actively controlling event schedule and thread interleaving, given the test cases that have potential data races. By exploring the state model of the application constructed dynamically, our technique starts first to generate a test case that has potential data races based on the results obtained from existing static or dynamic race detection technique. Then it reschedules test cases execution by actively controlling event dispatching and thread interleaving to determine whether such potential races really lead to thrown exceptions or assertion violations. Our preliminary experiments show that RacerDroid is effective, and it confirms real data races, while at the same time eliminates false warnings for Android apps found in the wild.
Hongyin Tang, Guoquan Wu, Jun Wei 0001, Hua Zhong 0001
ASE4
2016 Crawling Hidden Objects with kNN Queries
abstract
Many websites offering Location Based Services (LBS) provide a$k$NN search interface that returns the top-$k$nearest-neighbor objects (e.g., nearest restaurants) for a given query location. This paper addresses the problem of crawling all objects efficiently from an LBS website, through the public$k$NN web search interface it provides. Specifically, we develop crawling algorithm for 2D and higher-dimensional spaces, respectively, and demonstrate through theoretical analysis that the overhead of our algorithms can be bounded by a function of the number of dimensions and the number of crawled objects, regardless of the underlying distributions of the objects. We also extend the algorithms to leverage scenarios where certain auxiliary information about the underlying data distribution, e.g., the population density of an area which is often positively correlated with the density of LBS objects, is available. Extensive experiments on real-world datasets demonstrate the superiority of our algorithms over the state-of-the-art competitors in the literature.
Zhiguo Gong, Nan Zhang 0004, Tao Huang 0001, Hua Zhong 0001, Jun Wei 0001
IEEE Trans. Knowl. Data Eng.5
2016 FD4C: Automatic Fault Diagnosis Framework for Web Applications in Cloud Computing
abstract
The large-scale dynamic cloud computing environment has raised great challenges for fault diagnosis in Web applications: First, fluctuating workloads cause traditional application models to change over time; second, modeling the behaviors of complex applications usually requires domain knowledge which is difficult to obtain; third, managing large-scale applications manually is impractical for operators. To address these issues, this paper proposes an automatic fault (F) diagnosis (D) framework for (4) Web applications in cloud (C) computing (FD4C). In this paper, we propose an online incremental clustering method to recognize access behavior patterns. We also use correlation analysis to model the correlations between the workloads and application performance/resource utilization metrics in a specific access behavior pattern. FD4C detects faults by discovering the abrupt changes of correlation coefficients with control charts. Then, FD4C identifies the fault-related metrics using a feature selection method. To evaluate our proposal, we inject typical faults into TPC-W benchmark and apply FD4C to diagnose the injected faults. The experimental results show that FD4C can effectively detect the typical faults and accurately locate the metrics related to the faults.
Tao Wang 0030, Wenbo Zhang 0006, Chunyang Ye, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
IEEE Trans. Syst. Man Cybern. Syst.5
2015 Fault detection for cloud computing systems with correlation analysis
abstract
The large-scale dynamic cloud computing environment has raised great challenges for fault diagnosis in Web applications. First, fluctuating workloads cause traditional application models to change over time. Moreover, modeling the behaviors of complex applications always requires domain knowledge which is difficult to obtain. Finally, managing large-scale applications manually is impractical for operators. This paper addresses these issues and proposes an automatic fault diagnosis method for Web applications in cloud computing. We propose an online incremental clustering method to recognize access behavior patterns, and uses CCA to model the correlation between workloads and the metrics of application performance/resource utilization in a specific access behavior pattern. Our method detects anomalies by discovering the abrupt change of correlation coefficients with a EWMA control chart, and then locates suspicious metrics using a feature selection method combining ReliefF and SVM-RFE. We validate our method by injecting typical faults in TPC-W an industry-standard benchmark, and the experimental results demonstrate that it can effectively detect typical faults.
Tao Wang 0030, Wenbo Zhang 0006, Jun Wei 0001, Hua Zhong 0001
IM4
2015 A Crowdsourcing framework for Detecting Cross-Browser Issues in Web Application
abstract
With the advent of Web 2.0 application, and the increasing number of browsers and platforms on which the applications can be executed, cross-browser incompatibilities (XBIs) are becoming a serious problem for organizations to develop web-based software with good user experience. Although some techniques and tools have been proposed to identify XBIs, some XBIs are still missed as only partial state space is explored (by the crawler) in the testing environment. To address this limitation, based on record/replay technique, this paper proposed a crowdsourcing framework to detect cross-browser issues for Web application deployed in the field. Our empirical evaluation shows that the proposed technique is effective and efficient, improves on the state of the art.
Meimei He, Hongyin Tang, Guoquan Wu, Jun Wei 0001, Hua Zhong 0001
Internetware5
2015 Aggregate Estimation in Hidden Databases with Checkbox Interfaces
abstract
A large number of web data repositories are hidden behind restrictive web interfaces, making it an important challenge to enable data analytics over these hidden web databases. Most existing techniques assume a form-like web interface which consists solely of categorical attributes (or numeric ones that can be discretized). Nonetheless, many real-world web interfaces (of hidden databases) also feature checkbox interfaces-e.g., the specification of a set of desired features, such as A/C, navigation, etc., for a car-search website like Yahoo! Autos. We find that, for the purpose of data analytics, such checkbox-represented attributes differ fundamentally from the categorical/numerical ones that were traditionally studied. In this paper, we address the problem of data analytics over hidden databases with checkbox interfaces. Extensive experiments on both synthetic and real datasets demonstrate the accuracy and efficiency of our proposed algorithms.
Zhiguo Gong, Nan Zhang 0004, Tao Huang 0001, Hua Zhong 0001, Jun Wei 0001
IEEE Trans. Knowl. Data Eng.5
2014 Inferring Data Contract for Web-Based API
abstract
Web-based API is a new trend for publishing services. To correctly use the API, developers should follow certain service specifications. Data contract is a service specification to express the constraints over the data model used in the APIs. Data contracts, however are not always readily available in a formalized format if not undocumented at all. In this paper, we present an approach to infer formal data contracts for Web-based API. The approach integrates information of the parameters, error messages and testing result of Web-based API. We demonstrate how this approach infers complicated data preconditions for Web-based API in the real-world Web API platforms.
Chushu Gao, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
ICWS3
2014 Runtime Enforcement of Data-centric Properties for Concurrent Service-Based Applications
abstract
For service-based applications which are composed of multiple independent third-parties, continuous monitoring is required to assure that runtime behavior of the systems complies with specified properties. However, most existing work only detects the violation while not consider how to enforce the properties so that the constraint can not be violated at runtime. To address this limitation, this paper presents EnforceBCL, a framework for enforcing data-centric properties for concurrent service-based applications. Users of EnforceBCL can specify the properties to be enforced using the expressive behavior constraint enforcement language. Data-centric property is enforced at runtime by blocking the process whose next action would violate it. The impacted processes can be unblocked and allowed to execute when the specified property eventually reaches a safe state. EnforceBCL also provides the mechanism to detect possible deadlock during the enforcement of the property, and executes corresponding handler to solve the deadlock. To evaluate the effectiveness and efficiency of the proposed approach, we conducted several experiments. Results show that EnforceBCL is able to effectively enforce data-centric properties for concurrent service-based applications and also incurs less performance overhead.
Guoquan Wu, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
ICWS3
2014 MC-Checker: Detecting Memory Consistency Errors in MPI One-Sided Applications
abstract
One-sided communication decouples data movement and synchronization by providing support for asynchronous reads and updates of distributed shared data. While such interfaces can be extremely efficient, they also impose challenges in properly performing asynchronous accesses to shared data. This paper presents MC-Checker, a new tool that detects memory consistency errors in MPI one-sided applications. MCChecker first performs online instrumentation and captures relevant dynamic events, such as one-sided communications and load/store operations. MC-Checker then performs analysis to detect memory consistency errors. When found, errors are reported along with useful diagnostic information. Experiments indicate that MC-Checker is effective at detecting and diagnosing memory consistency bugs in MPI one-sided applications, with low overhead, ranging from 24.6% to 71.1%, with an average of 45.2%.
Zhezhe Chen, James Dinan, Pavan Balaji, Hua Zhong 0001, Jun Wei 0001, Tao Huang 0001
SC5
2014 Workload-aware anomaly detection for Web applications
Tao Wang 0030, Jun Wei 0001, Wenbo Zhang 0006, Hua Zhong 0001, Tao Huang 0001
J. Syst. Softw.4
2013 Detecting performance anomaly with correlation analysis for Internetware
Tao Wang 0030, Jun Wei 0001, Wenbo Zhang 0006, Hua Zhong 0001, Tao Huang 0001
Sci. China Inf. Sci.5
2012 Application-Level CPU Consumption Estimation: Towards Performance Isolation of Multi-tenancy Web Applications
abstract
Performance isolation is a key requirement for application-level multi-tenant sharing hosting environments. It requires knowledge of the resource consumption of the various tenants. It is of great importance not only to be aware of the resource consumption of a tenant's given kind of transaction mix, but also to be able to be aware of the resource consumption of a given transaction type. However, direct measurement of CPU resource consumption requires instrumentation and incurs overhead. Recently, regression analysis has been applied to indirectly approximate resource consumption, but challenges still remain for cases with non-determinism and multicollinearity. In this work, we adapts Kalman filter to estimate CPU consumptions from easily observed data. We also propose techniques to deal with the non-determinism and the multicollinearity issues. Experimental results show that estimation results are in agreement with the corresponding measurements with acceptable estimation errors, especially with appropriately tuned filter settings taken into account. Experiments also demonstrate the utility of the approach in avoiding performance interference and CPU overloading.
Wei Wang 0049, Xiang Huang 0005, Xiulei Qin, Wenbo Zhang 0006, Jun Wei 0001, Hua Zhong 0001
IEEE CLOUD6
2012 Workload-Aware Online Anomaly Detection in Enterprise Applications with Local Outlier Factor
abstract
Detecting anomalies are essential for improving the reliability of enterprise applications. Current approaches set thresholds for metrics or model correlations between metrics, and anomalies are detected when the thresholds are violated or the correlations are broken. However, we have found that the dynamic workload fluctuating over multiple time scales causes system metrics and their correlations to change. Moreover, it is difficult to model various metric correlations in complex applications. This paper addresses these problems and proposes an online anomaly detection approach for enterprise applications. A method is presented for recognizing workload patterns with an incremental clustering algorithm. The Local Outlier Factor (LOF) based on the specific workload pattern is adopted for detecting anomalies. Our approach is evaluated on a testbed running the TPC-W benchmark. The experimental results show that our approach can capture workload fluctuations accurately and detect the typical faults effectively.
Tao Wang 0030, Wenbo Zhang 0006, Jun Wei 0001, Hua Zhong 0001
COMPSAC4
2012 Online Anomaly Detection for Components in OSGi-based Software
Tao Wang 0030, Wenbo Zhang 0006, Jun Wei 0001, Hua Zhong 0001
SEKE5
2012 Specification and monitoring of data-centric temporal properties for service-based systems
Guoquan Wu, Jun Wei 0001, Chunyang Ye, Hua Zhong 0001, Tao Huang 0001, Hong He 0004
J. Syst. Softw.4
2011 On-line Cache Strategy Reconfiguration for Elastic Caching Platform: A Machine Learning Approach
abstract
Cloud computing provide scalability and high availability for web applications using such techniques as distributed caching and clustering. As one database offloading strategy, elastic caching platforms (ECPs) are introduced to speed up the performance or handle application state management with fault tolerance. Several cache strategies for ECPs have been proposed, say replicated strategy, partitioned strategy and near strategy. We first evaluate the impact of the three cache strategies using the TPC-W benchmark and find that there is no single cache strategy suitable for all conditions, the selection of the best strategy is related with workload patterns, cluster size and the number of concurrent users. This raises the question of when and how the cache strategy should be reconfigured as the condition varies which has received comparatively less attention. In this paper, we present a machine learning based approach to solving this problem. The key features of the approach are off-line training coupled with on-line system monitoring and robust synchronization process after triggering a reconfiguration, at the same time the performance model is periodically updated. More explicitly, first a rule set used to identify which cache strategy is optimal under the current condition are trained with the system statistics and performance results. We then introduce a framework to switch the cache strategy on-line as the workload varies and keep its overhead to acceptable levels. Finally, we illustrate the advantages of this approach by carrying out a set of experiments.
Xiulei Qin, Wenbo Zhang 0006, Wei Wang 0049, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
COMPSAC5
2011 A Statistical Approach for Estimating CPU Consumption in Shared Java Middleware Server
abstract
Middleware sharing is one of the important resource sharing approaches which enables sharing of costs across a large pool of users. However, the shared Java middleware server easily causes interference on performance between concurrent user requests. A key requirement to an effective performance isolation is the knowledge of the resource consumption of the various kinds of use requests classified according to different application context information. Direct measurement of resource consumption requires instrumentation which is impractical. In this paper, we demonstrate that CPU consumptions of various kinds of user requests on a given hardware can be approximated by a proposed Kalman filter based approach. Experimental results derived from testing the approach by using the TPC-W e-commerce suite deployed on a widely-used Java middleware server (Tomcat) illustrate the potential of this approach.
Wei Wang 0049, Xiang Huang 0005, Yunkui Song, Wenbo Zhang 0006, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
COMPSAC6
2011 Runtime Monitoring of Data-centric Temporal Properties for Web Services
abstract
Runtime monitoring of Web service compositions has been widely acknowledged as a significant approach to understand and guarantee the quality of services. However, existing runtime monitoring solutions consider only the constraints on the sequence of messages exchanged between partner services and ignore the actual data contents inside the messages. As a result, it is difficult to monitor some dynamic properties such as how message data of interest is processed between different participants. To address this issue, we propose an efficient, non-intrusive online monitoring approach to dynamically analyze data-centric properties for service-oriented applications involving multiple participants. By introducing Par-BCL - a Parametric Behavior Constraint Language for web services - to define monitoring parameters, various data-centric temporal behavior properties for Web services can be specified and monitored. This approach broadens the monitored patterns to include not only message exchange orders, but also the data contents bound to the parameters. To reduce runtime overhead, we statically analyze the monitored properties to generate parameter state machine from the event pattern automata to optimize monitoring. The experiments show that our solution is efficient and promising.
Guoquan Wu, Jun Wei 0001, Chunyang Ye, Xiaozhe Shao, Hua Zhong 0001, Tao Huang 0001
ICWS5
2011 Runtime Verification of Data-Centric Properties in Service Based Systems
Guoquan Wu, Jun Wei 0001, Chunyang Ye, Xiaozhe Shao, Hua Zhong 0001, Tao Huang 0001
RV5
2010 A Two-Phase Approach to Subscription Subsumption Checking for Content-Based Publish/Subscribe Systems
abstract
The efficiency of subscription subsumption checking remains a key issue for content-based publish/subscribe systems. In this paper, we propose an efficient data structure called subscription subsumption graph (SSG). This data structure could differentiate the two types of subsumption relationships and help speed up the process of subsumption checking and subscription cancellation. We then present a two-phase approach to subscription subsumption checking. Phase one is mainly about checking of non-numeric constraints by using an index structure which could help filter out most of irrelevant subscriptions while phase two is about checking of remaining numeric constraints where SSG is employed. Finally, we introduce an efficient SSG-based unsubscription algorithm that could find out which subscriptions need to be forwarded without any redundant computing. We illustrate the advantages of this approach by carrying out extensive experiments.
Xiulei Qin, Jun Wei 0001, Wenbo Zhang 0006, Hua Zhong 0001, Tao Huang 0001
AINA4
2010 Detecting Data Inconsistency Failure of Composite Web Services Through Parametric Stateful Aspect
abstract
Runtime monitoring of Web service compositions with WS-BPEL has been widely acknowledged as a significant approach to understand and guarantee the quality of services. However, most existing monitoring technologies only track patterns related to the execution of an individual process. As a result, the possible inconsistency failure caused by implicit interactions among concurrent process instances cannot be detected. To address this issue, this paper proposes an approach to specify the behavior properties related to shared resources for web service compositions and verify their consistency with the aid of a parametric stateful aspect extension to WS-BPEL. Parameters are introduced in pattern specification, which allows monitoring not only events but also their values bound to the parameters at runtime to keep track of data flow among concurrent process instances. An efficient implementation is also provided to reduce the runtime overhead of monitoring and event observation. Our experiments show that the proposed approach is promising.
Guoquan Wu, Jun Wei 0001, Chunyang Ye, Hua Zhong 0001, Tao Huang 0001
ICWS4
2010 Middleware support for internetware: a service perspective
abstract
The advent of Internet technology introduces a revolution to software application and development paradigms. Traditional software development and application patterns have been shifted to Internet-based service sharing and collaboration among partners all over the Internet. This imposes new challenges and complexity in the lifecycle of software development, deployment and maintenance. Middleware, an intermediate layer to abstract the homogeneity and hide the difference of underlying systems, can be used to reduce the complexity for Internet application development. In this paper, we exploit the needs of middleware support for Internet-based applications from a service perspective. We investigate the potential requirements and features of Internetware, and the state-of-the-art solutions. We also analyze the remaining issues, the challenges and potential future research directions.
Chunyang Ye, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
Internetware3
2009 A study on the replaceability of context-aware middleware
abstract
In context-ware computing paradigm, context-aware middleware plays a key role. The middleware collects and manipulates contexts from environments, providing context-aware applications well-defined interfaces to adapt their behaviors when environments change. However, some minor difference in the implementation of context-aware middleware may cause the same context-aware application behave differently. Such behavior deviation may lead to serious problems or even disasters for a context-aware application. It is thus desirable to check whether a mobile context-aware application behaves consistently before moving it from one middleware to another, or whether a context-aware application still works correctly when upgrading the underlying middleware? Existing approaches for context-aware applications are not adequate for detecting such behavior deviation because these approaches do not consider the impacts of the difference in the middleware implementation. In this paper, we study the strategies in the implementation of context-aware middleware and their impacts on the behavior of context-aware applications. By exploring the implied scenarios where a context-aware application may behave differently, new testing approach is proposed to detect the behavior deviation of a context-aware application running on different middleware by generating test cases to cover these implied scenarios.
Chunyang Ye, Shing-Chi Cheung, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
Internetware4
2007 An Interaction Instance Oriented Approach for Web Application Integration in Portals
abstract
Currently, the significance of a portal stems not only from being a handy way to access data but also from being the means of facilitating the integration with Web applications. This paper proposes an interaction instance oriented approach for integrating web applications in portals. The key aspect is to enable a user to have the same interaction experiences from a portal that he/she accesses the web application directly. The approach is thus focused on the description of the presentation layer of the interaction instances of a web application and a portlet, which defines an interaction instance as consecutive web pages or fragments. Web application integration is then transformed to the problem that how to translate all web pages of an interaction instance of a Web application to certain view equivalence regions, which form the interaction instance of a portlet. Experiments show that the approach is effective and efficient.
Jingyu Song, Jun Wei 0001, Shuchao Wan, Hua Zhong 0001
COMPSAC (1)4
2007 A Satisfaction Driven Approach for the Composition of Interactive Web Services
abstract
The paradigmatic shift from function-oriented Web services to interactive Web services(IWSs) addresses lots of issues such as repetitious development for user interface and tiresome understanding of underlying service operations. However, the composition of existing IWSs to create value-added ones is still a prominent problem. Considering the special interactive characteristics of IWSs, current approaches to compose function-oriented services are inappropriate. This paper proposes a novel user satisfaction model to evaluate the interactive quality of a composite IWS completely. Based on the satisfaction model, an effective satisfaction-driven approach is also developed for service selection in IWSs composition, which can meet diverse interactive requirements of users.
Shuchao Wan, Jun Wei 0001, Jingyu Song, Hua Zhong 0001
COMPSAC (1)4
2007 A New Approach for Overload Management in Content-based Publish/Subscribe
abstract
Overload management is of vital importance in wide-area publish/subscribe systems, yet current solutions are best-effort. In this paper, we present an admission control scheme for overload management in large-scale and scalable content-based publish/subscribe systems. We analyze the stumbling block for implementing admission control in publish/subscribe systems, and point out how it differs from admission control schemes in other research areas. We propose a cover relation based algorithm to compute subscription resource requirements and an admission control algorithm based on subscription routing. The scheme ensures time, space and flows decoupling without sacrificing scalability of publish/subscribe systems. Finally, we conduct experiments to verify the effectiveness of the scheme.
Xiangfeng Guo, Hua Zhong 0001, Jun Wei 0001, Dongli Han
ICSEA2
2007 Sequential Pattern-Based Cache Replacement in Servlet Container
Lin Zuo, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
ICWE4