VLDB 2026 Research / reviewers in the wild / expert
Tao Huang 0001
dblp:34/808-1
· DBLP profile ↗
90ranked-venue papers
3as first author
25since 2021 · last 2025
0009-0006-4875-732XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 58 · 17 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 2 since 2021Systems, architecture and hardware · 9 · 4 since 2021Artificial intelligence and machine learning · 2Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Proving Cypher Query EquivalenceabstractGraph database systems store graph data as nodes and relationships, and utilize graph query languages (e.g., Cypher) for efficiently querying graph data. Proving the equivalence of graph queries is an important foundation for optimizing graph query performance, ensuring graph query reliability, etc. Although researchers have proposed many SQL query equivalence provers for relational database systems, these provers cannot be directly applied to prove the equivalence of graph queries. The difficulty lies in the fact that graph query languages (e.g., Cypher) adopt significantly different data models (property graph model vs. relational model) and query patterns (graph pattern matching vs. tabular tuple calculus) from SQL. In this paper, we propose GraphQE, an automated prover to determine whether two Cypher queries are semantically equivalent. We design a U-semiring based Cypher algebraic representation to model the semantics of Cypher queries. Our Cypher algebraic representation is built on the algebraic structure of unbounded semirings, and can sufficiently express nodes and relationships in property graphs and complex Cypher queries. Then, determining the equivalence of two Cypher queries is transformed into determining the equivalence of the corresponding Cypher algebraic representations, which can be verified by SMT solvers. To evaluate the effectiveness of GraphQE, we construct a dataset consisting of 148 pairs of equivalent Cypher queries. Among them, we have successfully proven 138 pairs of equivalent Cypher queries, demonstrating the effectiveness of GraphQE. Wensheng Dou, Yingying Zheng, Lijie Xu, Wei Wang 0049, Jun Wei 0001, Tao Huang 0001 |
ICDE | 7 |
| 2025 | Evaluating Garbage Collection Performance Across Managed Language RuntimesabstractModern managed language runtimes (e.g., Java, Go and C#) rely on garbage collection (GC) mechanisms to automatically allocate and reclaim in-memory objects. The efficiency of GC implementations can greatly impact the overall performance of runtime-based applications. To improve GC performance, the academic and industrial communities have proposed several approaches to evaluate the GC implementations in an individual runtime. However, these approaches target a specific managed language (e.g., Java), and cannot be used to compare the GC implementations in different runtimes. In this paper, we propose GEAR, an automated approach to construct consistent GC workloads for different managed language runtimes, which can further be used to evaluate GC implementations across different runtimes. Specifically, we design a group of runtime-agnostic Memory Operation Primitives (MOP), which can portray the memory usage information that influences GC. GEAR can further automatically convert a MOP program into runtime-specific programs for the target runtimes, which serve as a consistent GC workload for different runtimes. To build MOP programs with real-world GC workloads, we instrument the commonly-used runtime Java Virtual Machine (JVM) to collect the memory operation trace during a Java application's execution, and then transform the memory operation trace into a MOP program. The experimental result on three widely-used runtimes (i.e., Java, Go and C#) shows that GEAR can generate consistent GC workloads for different runtimes. We further conduct a comprehensive study on these three runtimes, and reveal some interesting findings about their GC performance, providing useful guidance for improving their GC implementations. Wensheng Dou, Yi Wang 0069, Wei Wang 0049, Jun Wei 0001, Tao Huang 0001 |
ICSE | 7 |
| 2025 | SCodeGen: A Real-Time Trustworthy Constrained Decoding Framework for Secure Code Generation with LLMsabstractLarge language models (LLMs) are increasingly integrated into software development workflows to accelerate code generation, but often produce insecure and uncontrollable code due to vulnerable training data and unconstrained decoding strategies. This poses severe risks in security-critical systems, where post-generation vulnerability detection and manual remediation incur significant overhead. While constrained decoding offers a practical mitigation strategy, existing methods suffer from degraded trustworthiness, constraint conflicts, and high latency—especially when enforcing multiple concurrent security constraints.We propose SCodeGen, a real-time constrained decoding framework designed to enforce fine-grained security controls during LLM code generation. To improve trustworthiness and controllability, SCodeGen introduces (1) a matching-length-aware logit modulation strategy that enhances trustworthiness and controllability without semantic disruption, and (2) a two-stage low-latency decoding architecture, which compiles constraint phrases into a runtime-enforceable constraint automaton (RCA) with precomputed logit bias vectors for efficient online decoding. Extensive evaluations on CodeGuard+ show that SCodeGen significantly improves secure pass rates under both single and multi-constraint settings, while maintaining latency comparable to unconstrained decoding. This work demonstrates a practical and scalable solution toward trustworthy LLM-assisted software development under security constraints. Muzi Qu, Jie Liu 0008, Liangyi Kang, Shuyi Ling, Dan Ye 0004, Tao Huang 0001 |
TrustCom | 7 |
| 2025 | BridgeGC: An Efficient Cross-Level Garbage Collector for Big Data FrameworksabstractPopular big data frameworks commonly run atop Java Virtual Machine (JVM) and rely on garbage collection (GC) mechanism to automatically allocate/reclaim in-memory objects. Existing garbage collectors are designed based on the hypothesis that most objects are short lived. However, big data frameworks usually generate many long-lived data objects, which can cause heavy GC overhead. Recent approaches have reduced GC overhead in big data frameworks but still suffer from heavy human efforts, additional runtime overhead, or suboptimal GC efficiency. This article describes the design of BridgeGC , a big-data-friendly garbage collector that significantly reduces GC overhead introduced by long-lived data objects. BridgeGC follows a cross-level co-design. At the big data framework level, BridgeGC provides two annotations for framework developers to denote the creation and release of data objects. Based on the annotations, BridgeGC tracks the lifecycles of annotated data objects and optimizes their allocation/reclamation at the GC level. At the GC level, we design a label-based allocator that stores data objects separately from other objects and balances their memory usage in the same JVM, leading to fewer GC cycles. We further design an efficient collector to eliminate unnecessary marking and copying of data objects during GC cycles, lowering the GC time. We have integrated BridgeGC into OpenJDK ZGC. The extensive evaluation, using two popular big data frameworks (Flink and Spark) and a key–value database (Cassandra), shows that BridgeGC achieves 31–82% GC time reduction compared to the baseline ZGC. BridgeGC also outperforms other traditional and academic garbage collectors in end-to-end performance. Lijie Xu, Tian Guo 0001, Wensheng Dou, Hongbin Zeng, Wei Wang 0049, Jun Wei 0001, Tao Huang 0001 |
ACM Trans. Archit. Code Optim. | 8 |
| 2024 | How to Fit the SCC Algorithm Efficiently into Distributed Graph Iterative ComputationabstractThis paper reviews the sequential, parallel and distributed implementations of strongly connected component algorithms, and analyzes the challenges of each implementation in the graph iteration paradigm of distributed processing. We also review the graph data layout and communication mode of each distributed graph processing system, and analyze the defect of high memory usage in the implementation of the strongly connected component algorithm of the existing distributed graph processing system. Therefore, we propose a strongly connected component algorithm that performs pull mode communication on CSR instead of traversing the transposed graph. Experiments show that the memory consumption of our method is greatly reduced, and the running time of the algorithm is much lower than that of the most advanced implementation. On four publicly accessible data sets, our approach reduces computation time and storage space by an average of 26% and 39%, respectively. In addition, we optimize the general pull communication mode, resulting in 15.2°/0 computation time reduction on the four publicly available data sets. Xiaochen Sun, Wei Wang 0049, Tao Huang 0001 |
COMPSAC | 3 |
| 2024 | Efficient Multi-network Community Search Method for Distributed Graph Iterative ComputationabstractGraph is often used for data analysis. Distributed graph processing is gaining traction as it becomes more difficult for a single machine to store and process the complete graph due to the growing volume of data. We investigated 26 popular distributed graph processing systems and the graph algorithms and datasets provided by these systems. The computational logic of these graph algorithms does not distinguish between the types of vertices and edges, so distributed graph processing systems treat all vertices and edges in an undifferentiated way. However, using the hidden data connections of different types of vertices in multi-networks can greatly improve the accuracy of the community search algorithm. So we describe the challenges for the existing distributed graph processing systems to deal with different types of vertices and edges in multi-networks, and propose an index-based multi-network storage abstraction to store various vertices and edges, and a heuristic greedy algorithm to complete the partition job for different vertices and edges. Base on the two jobs, we finish the research of efficient community search for distributed graph processing in multi-networks, making it possible for future research of more algorithms for distributed graph processing in multi-networks. Xiaochen Sun, Wei Wang 0049, Tao Huang 0001 |
COMPSAC | 3 |
| 2024 | Understanding Transaction Bugs in Database SystemsabstractTransactions are used to guarantee data consistency and integrity in Database Management Systems (DBMSs), and have become an indispensable component in DBMSs. However, faulty designs and implementations of DBMSs' transaction processing mechanisms can introduce transaction bugs, and lead to severe consequences, e.g., incorrect database states and DBMS crashes. An in-depth understanding of real-world transaction bugs can significantly promote effective techniques in combating transaction bugs in DBMSs. Ziyu Cui, Wensheng Dou, Yu Gao 0002, Dong Wang 0048, Jiansen Song, Yingying Zheng, Tao Wang 0030, Rui Yang 0039, Jun Wei 0001, Tao Huang 0001 |
ICSE | 12 |
| 2024 | Differential Optimization Testing of Gremlin-Based Graph Database SystemsabstractGraph database systems (GDBs) allow efficiently creating, modifying, and retrieving graph data in a graph database. To accelerate graph queries, GDBs usually adopt various and complex optimization strategies. However, incorrect optimizations in GDBs can introduce optimization bugs, which cause a graph query to compute an incorrect query result, e.g., omitting a vertex in a graph database. In this paper, we propose Differential Optimization Testing (DOT), an effective and automated approach to detect optimization bugs in GDBs that adopt Gremlin as their query language. The main idea of DOT is that, given a Gremlin query$Q$, we execute it on the target GDB with two different optimization configurations and then verify whether they can compute the same query results for query$Q$. Any inconsistency between their query results indicates an optimization bug in the target GDB. To improve the efficiency of differential testing in DOT, we further propose an optimization-guided approach, aiming to explore more optimization strategies and more graph database features. We evaluate DOT on six popular and widely-used GDBs, i.e., Neo4j, OrientDB, JanusGraph, HugeGraph, TinkerGraph, and ArcadeDB. In total, we have found 28 unique optimization bugs, 16 of which have been confirmed as previously-unknown bugs. Yingying Zheng, Wensheng Dou, Ziyu Cui, Jiansen Song, Ziyue Cheng, Wei Wang 0049, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001 |
ICST | 10 |
| 2024 | Testing Gremlin-Based Graph Database Systems via Query DisassemblingabstractGraph Database Systems (GDBs) support efficiently storing and retrieving graph data, and have become a critical component in many important applications. Many widely-used GDBs utilize the Gremlin query language to create, modify, and retrieve data in graph databases, in which developers can assemble a sequence of Gremlin APIs to perform a complex query. However, incorrect implementations and optimizations of GDBs can introduce logic bugs, which can cause Gremlin queries to return incorrect query results, e.g., omitting vertices in a graph database. In this paper, we propose Query Di sassembling (QuDi), an effective testing technique to automatically detect logic bugs in Gremlin-based GDBs. Given a Gremlin query Q, QuDi disassembles Q into a sequence of atomic graph traversals TList, which shares the equivalent execution semantics with Q. If the execution results of Q and TList are different, a logic bug is revealed in the target GDB. We evaluate QuDi on six popular GDBs, and have found 25 logic bugs in these GDBs, 10 of which have been confirmed as previously-unknown bugs by GDB developers. Yingying Zheng, Wensheng Dou, Ziyu Cui, Yu Gao 0002, Jiansen Song, Wei Wang 0049, Jun Wei 0001, Hua Zhong 0007, Tao Huang 0001 |
ISSTA | 12 |
| 2024 | Dynamic Scoring Code Token Tree: A Novel Decoding Strategy for Generating High-Performance CodeabstractWithin the realms of scientific computing, large-scale data processing, and artificial intelligence-powered computation, disparities in performance, which originate from differing code implementations, directly influence the practicality of the code. Although existing works tried to utilize code knowledge to enhance the execution performance of codes generated by large language models, they neglect code evaluation outcomes which directly refer to the code execution details, resulting in inefficient computation. To address this issue, we propose DSCT-Decode, an innovative adaptive decoding strategy for large language models, that employs a data structure named 'Code Token Tree' (CTT), which guides token selection based on code evaluation outcomes. DSCT-Decode assesses generated code across three dimensions---correctness, performance, and similarity---and utilizes a dynamic penalty-based boundary intersection method to compute multi-objective scores, which are then used to adjust the scores of nodes in the CTT during backpropagation. By maintaining a balance between exploration, through token selection probabilities, and exploitation, through multi-objective scoring, DSCT-Decode effectively navigates the code space to swiftly identify high-performance code solutions. To substantiate our framework, we developed a new benchmark, big-DS-1000, which is an extension of DS-1000. This benchmark is the first of its kind to specifically evaluate code generation methods based on execution performance. Comparative evaluations with leading large language models, such as CodeLlama and GPT-4, show that our framework achieves an average performance enhancement of nearly 30%. Furthermore, 30% of the codes exhibited a performance improvement of more than 20%, underscoring the effectiveness and potential of our framework for practical applications. Muzi Qu, Jie Liu 0008, Liangyi Kang, Dan Ye 0004, Tao Huang 0001 |
ASE | 6 |
| 2024 | Match Word with Deed: Maintaining Consistency for IoT Systems with Behavior ModelsabstractEnsuring the reliability and consistency of Internet of Things (IoT) systems is critical. Traditional approaches to maintaining consistency often rely on retry and rollback mechanisms, which can be inadequate and lead to further complications. These methods struggle with the complexity and heterogeneity of IoT systems, failing to provide robust and general solutions for real-time consistency assurance. Tao Wang 0030, Wei Chen 0018, Guoquan Wu, Jun Wei 0001, Tao Huang 0001 |
ASE | 6 |
| 2024 | Detecting Metadata-Related Logic Bugs in Database Systems via Raw Database ConstructionabstractDatabase Management Systems (DBMSs) are widely used to efficiently store and retrieve data. DBMSs usually support various metadata, e.g., integrity constraints for ensuring data integrity and indexes for locating data. DBMSs can further utilize these metadata to optimize query evaluation. However, incorrect metadata-related optimizations can introduce metadata-related logic bugs, which can cause a DBMS to return an incorrect query result for a given query. In this paper, we propose a general and effective testing approach, Raw database construction (Radar), to detect metadata-related logic bugs in DBMSs. Given a database db containing some metadata, Radar first constructs a raw database rawDb , which wipes out the metadata in db and contains the same data as db. Since db and rawDb have the same data, they should return the same query result for a given query. Any inconsistency in their returned query results indicates a metadata-related logic bug. To effectively detect metadata-related logic bugs, we further propose a metadata-oriented testing optimization strategy to focus on testing previously unseen metadata, thus detecting more metadata-related logic bugs quickly. We implement and evaluate Radar on five widely-used DBMSs, and have detected 42 bugs, of which 38 have been confirmed as new bugs and 16 have been fixed by DBMS developers. Jiansen Song, Wensheng Dou, Yu Gao 0002, Ziyu Cui, Yingying Zheng, Dong Wang 0048, Wei Wang 0049, Jun Wei 0001, Tao Huang 0001 |
Proc. VLDB Endow. | 9 |
| 2023 | A Reinforcement Learning Approach to Generating Test Cases for Web ApplicationsabstractWeb applications play an important role in modern society. Quality assurance of web applications requires lots of manual efforts. In this paper, we propose WebQT, an automatic test case generator for web applications based on reinforcement learning. Specifically, to increase testing efficiency, we design a new reward model, which encourages the agent to mimic human testers to interact with the web applications. To alleviate the problem of state redundancy, we further propose a novel state abstraction technique, which can identify different web pages with the same functionality as the same state, and yields a simplified state space. We evaluate WebQT on seven open-source web applications. The experimental results show that WebQT achieves 45.4% more code coverage along with higher efficiency than the state-of-the-art technique. In addition, WebQT also reveals 69 exceptions in 11 real-world web applications. Xiaoning Chang, Zheheng Liang, Zhenyue Long, Guoquan Wu, Yu Gao 0002, Wei Chen 0018, Jun Wei 0001, Tao Huang 0001 |
AST | 10 |
| 2023 | Model Checking Guided Testing for Distributed SystemsabstractDistributed systems have become the backbone of cloud computing. Incorrect system designs and implementations can greatly impair the reliability of distributed systems. Although a distributed system design modelled in the formal specification can be verified by formal model checking, it is still challenging to figure out whether its corresponding implementation conforms to the verified specification. An incorrect system implementation can violate its verified specification, and causes intricate bugs. Dong Wang 0048, Wensheng Dou, Yu Gao 0002, Chenao Wu, Jun Wei 0001, Tao Huang 0001 |
EuroSys | 6 |
| 2023 | Detecting Isolation Bugs via Transaction Oracle ConstructionabstractTransactions are used to maintain the data integrity of databases, and have become an indispensable feature in modern Database Management Systems (DBMSs). Despite extensive efforts in testing DBMSs and verifying transaction processing mechanisms, isolation bugs still exist in widely-used DBMSs when these DBMSs violate their claimed transaction isolation levels. Isolation bugs can cause severe consequences, e.g., incorrect query results and database states. In this paper, we propose a novel transaction testing approach, Transaction oracle construction (Troc), to automatically detect isolation bugs in DBMSs. The core idea of Troc is to decouple a transaction into independent statements, and execute them on their own database views, which are constructed under the guidance of the claimed transaction isolation level. Any divergence between the actual transaction execution and the independent statement execution indicates an isolation bug. We implement and evaluate Troc on three widely-used DBMSs, i.e., MySQL, MariaDB, and TiDB. We have detected 5 previously-unknown isolation bugs in the latest versions of these DBMSs. Wensheng Dou, Ziyu Cui, Qianwang Dai, Jiansen Song, Dong Wang 0048, Yu Gao 0002, Wei Wang 0049, Jun Wei 0001, Hanmo Wang, Hua Zhong 0001, Tao Huang 0001 |
ICSE | 12 |
| 2023 | Coverage Guided Fault Injection for Cloud SystemsabstractTo support high reliability and availability, modern cloud systems are designed to be resilient to node crashes and reboots. That is, a cloud system should gracefully recover from node crashes/reboots and continue to function. However, node crashes/reboots that occur under special timing can trigger crash recovery bugs that lie in incorrect crash recovery protocols and their implementations. To ensure that a cloud system is free from crash recovery bugs, some fault injection approaches have been proposed to test whether a cloud system can correctly recover from various crash scenarios. These approaches are not effective in exploring the huge crash scenario space without developers' knowledge. In this paper, we propose Crash Fuzz, a fault injection testing approach that can effectively test crash recovery behaviors and reveal crash recovery bugs in cloud systems. CrashFuzz mutates the combinations of possible node crashes and reboots according to runtime feedbacks, and prioritizes the combinations that are prone to increase code coverage and trigger crash recovery bugs for smart exploration. We have implemented CrashFuzz and evaluated it on three popular open-source cloud systems, i.e., ZooKeeper, HDFS and HBase. CrashFuzz has detected 4 unknown bugs and 1 known bug. Compared with other fault injection approaches, CrashFuzz can detect more crash recovery bugs and achieve higher code coverage. Yu Gao 0002, Wensheng Dou, Dong Wang 0048, Wenhan Feng, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001 |
ICSE | 7 |
| 2023 | Testing Database Systems via Differential Query ExecutionabstractDatabase Management Systems (DBMSs) provide efficient data retrieval and manipulation for many applications through Structured Query Language (SQL). Incorrect implementations of DBMSs can result in logic bugs, which cause SELECT queries to fetch incorrect results, or UPDATE and DELETE queries to generate incorrect database states. Existing approaches mainly focus on detecting logic bugs in SELECT queries. However, logic bugs in UPDATE and DELETE queries have not been tackled. In this paper, we propose a novel and general approach, which we have termed Differential Query Execution (DQE), to detect logic bugs in SELECT, UPDATE and DELETE queries of DBMSs. The core idea of DQE is that different SQL queries with the same predicate usually access the same rows in a database. For example, a row updated by an UPDATE query with a predicate φ should also be fetched by a SELECT query with the same predicate φ, If not, a logic bug is revealed in the target DBMS. To evaluate the effectiveness and generality of DQE, we apply DQE on five production-level DBMSs, i.e., MySQL, MariaDB, TiDB, CockroachDB and SQLite. In total, we have detected 50 unique bugs in these DBMSs, 41 of which have been confirmed, and 11 have been fixed. We expect that the simplicity and generality of DQE can greatly improve the reliability of DBMSs. Jiansen Song, Wensheng Dou, Ziyu Cui, Qianwang Dai, Wei Wang 0049, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001 |
ICSE | 8 |
| 2023 | Characterizing Flaky Tests in Node.js ApplicationsabstractRegression testing is an important means of assessing the quality of Node.js applications. However, non-deterministic executions inside Node.js framework could make test cases intermittently pass or fail on the same version of code, which are called flaky tests. Flaky tests can cause unreliable test results, and make developers waste a significant amount of time debugging the bugs that do not belong to the target application. In this paper, we conduct an empirical study on 87 flaky tests from 7 popular Node.js applications, and analyze the non-determinism that causes these flaky tests. Through this study, there is a wide range of non-determinism to cause flaky tests, including non-deterministic event triggering order, non-deterministic function calls, non-deterministic process/thread scheduling order, non-deterministic execution of asynchronous tasks and non-deterministic event triggering data. The result reveals that, existing approaches on event race detection are not sufficient for flaky test detection. In future, researchers can design flaky test detection approaches targeted at different categories of non-determinism. Xiaoning Chang, Zheheng Liang, Guoquan Wu, Yu Gao 0002, Wei Chen 0018, Jun Wei 0001, Zhenyue Long, Tao Huang 0001 |
ASE | 9 |
| 2023 | Detecting Smart Home Automation Application Interferences with Domain KnowledgeabstractTrigger-action programming (TAP) is a widely used development paradigm that simplifies the Internet of Things (loT) automation. However, the exceptional interactions between automation applications may result in interferences, such as conflicts and infinite loops, which cause undesirable consequences and even security and safety risks. While several techniques have been proposed to address this problem, they are often restricted in handling explicit and simple conflicts without considering contextual influences. In addition, they suffer from performance issues when applying to large-scale applications. To address these challenges, we design an effective and practical tool KnowDetector with comprehensive domain knowledge to detect application interferences. To detect application interferences, KnowDetector constructs an automation graph with 1) events, conditions, and actions from automation applications, 2) vertices representing physical environment channels, and 3) edges derived from potential semantic relations between the vertices. In order to make the graph extensively capture the interactions between automation applications, we propose a knowledge model named KnowloT that accurately characterizes loT devices with command-level loT services and the intricate relations between these services and the contextual environment. We abstract the interference detection into a graph pattern-matching problem and summarize ten application interference patterns of four types. Finally, KnowDetector can efficiently detect application interferences by searching for sub-graphs matching the patterns within the automation graph. We evaluated KnowDetector on three real-world datasets. The results demonstrated that it outperformed the other state-of-the-art tools with the highest precision, recall, and F-measure. In addition, KnowDetector is scalable to detect application interferences within a large number of applications with a minimal time overhead. Tao Wang 0030, Wei Chen 0018, Guoquan Wu, Jun Wei 0001, Tao Huang 0001 |
ASE | 6 |
| 2023 | LPW: an efficient data-aware cache replacement strategy for Apache Spark
Shuping Ji, Hua Zhong 0001, Wei Wang 0049, Lijie Xu, Jun Wei 0001, Tao Huang 0001 |
Sci. China Inf. Sci. | 8 |
| 2022 | Characterizing and Detecting Bugs in WeChat Mini-ProgramsabstractBuilt on the WeChat social platform, WeChat Mini-Programs are widely used by more than 400 million users every day. Consequently, the reliability of Mini-Programs is particularly crucial. However, WeChat Mini-Programs suffer from various bugs related to execution environment, lifecycle management, asynchronous mechanism, etc. These bugs have seriously affected users' experience and caused serious impacts. Tao Wang 0030, Qingxin Xu, Xiaoning Chang, Wensheng Dou, Jinhui Xie, Yuetang Deng, Jianbo Yang, Jiaheng Yang, Jun Wei 0001, Tao Huang 0001 |
ICSE | 11 |
| 2022 | Understanding device integration bugs in smart home systemabstractSmart devices have been widely adopted in our daily life. A smart home system, e.g., Home Assistant and openHAB, can be equipped with hundreds and even thousands of smart devices. A smart home system communicates with smart devices through various device integrations, each of which is responsible for a specific kind of devices. Developing high-quality device integrations is a challenging task, in which developers have to properly handle the heterogeneity of different devices, unexpected exceptions, etc. We find that device integration bugs, i.e., iBugs, are prevalent and have caused various consequences, e.g., causing devices unavailable, unexpected device behaviors. Tao Wang 0030, Kangkang Zhang, Wei Chen 0018, Wensheng Dou, Jun Wei 0001, Tao Huang 0001 |
ISSTA | 7 |
| 2021 | Evaluating the Parallel Execution Schemes of Smart Contract Transactions in Different Blockchains: An Empirical Study
Chengzhi Li, Heng Wu 0001, Heran Gao, Songchang Jin, Tao Huang 0001, Wenbo Zhang 0006 |
ICA3PP (3) | 6 |
| 2021 | Best VM Selection for Big Data Applications across Multiple Frameworks by Transfer LearningabstractCloud providers are presented with a bewildering choice of VM types for a range of contemporary data processing frameworks today. However, existing performance modeling and machine learning efforts cannot pick optimal VM types for multiple frameworks simultaneously, since they are difficult to balance model accuracy and model training cost. Yuewen Wu, Heng Wu 0001, Yuanjia Xu, Wenbo Zhang 0006, Hua Zhong 0007, Tao Huang 0001 |
ICPP | 7 |
| 2021 | Race Detection for Event-Driven Node.js ApplicationsabstractNode.js has become a widely-used event-driven architecture for server-side and desktop applications. Node.js provides an effective asynchronous event-driven programming model, and supports asynchronous tasks and multi-priority event queues. Unexpected races among events and asynchronous tasks can cause severe consequences. Existing race detection approaches in Node.js applications mainly adopt random fuzzing technique, and can miss races due to large schedule space.In this paper, we propose a dynamic race detection approach NRace for Node.js applications. In NRace, we build precise happens-before relations among events and asynchronous tasks in Node.js applications, which also take multi-priority event queues into consideration. We further develop a predictive race detection technique based on these relations. We evaluate NRace on 10 realworld Node.js applications. The experimental result shows that NRace can precisely detect 6 races, and 5 of them have been confirmed by developers. Xiaoning Chang, Wensheng Dou, Jun Wei 0001, Tao Huang 0001, Jinhui Xie, Yuetang Deng, Jianbo Yang, Jiaheng Yang |
ASE | 4 |
| 2020 | Hermes: Efficient Cache Management for Container-based Serverless ComputingabstractServerless computing systems are shifting towards shorter function durations and larger degrees of parallelism to eliminate intolerable latency. For container-based serverless computing, the state-of-the-art efforts fail to ensure low latency because on-demand container images reloading from remote storage can increase the data transmission rate and downgrades system performance. Heran Gao, Heng Wu 0001, Wenbo Zhang 0006, Tao Huang 0001 |
Internetware | 6 |
| 2019 | Detecting atomicity violations for event-driven Node.js applicationsabstractNode.js has been widely-used as an event-driven server-side architecture. To improve performance, a task in a Node.js application is usually divided into a group of events, which are non-deterministically scheduled by Node.js. Developers may assume that the group of events (named atomic event group) should be atomically processed, without interruption. However, the atomicity of an atomic event group is not guaranteed by Node.js, and thus other events may interrupt the execution of the atomic event group, break down the atomicity and cause unexpected results. Existing approaches mainly focus on event race among two events, and cannot detect high-level atomicity violations among a group of events. In this paper, we propose NodeAV, which can predictively detect atomicity violations in Node.js applications based on an execution trace. Based on happens-before relations among events in an execution trace, we automatically identify a pair of events that should be atomically processed, and use predefined atomicity violation patterns to detect atomicity violations. We have evaluated NodeAV on real-world Node.js applications. The experimental results show that NodeAV can effectively detect atomicity violations in these Node.js applications. Xiaoning Chang, Wensheng Dou, Yu Gao 0002, Jie Wang 0035, Jun Wei 0001, Tao Huang 0001 |
ICSE | 6 |
| 2019 | Aladdin: Optimized Maximum Flow Management for Shared Production ClustersabstractThe rise in popularity of long-lived applications (LLAs), such as deep learning and latency-sensitive online Web services, has brought new challenges for cluster schedulers in shared production environments. Scheduling LLAs needs to support complex placement constraints (e.g., to run multiple containers of an application on different machines) and larger degrees of parallelism to provide global optimization. But existing schedulers usually suffer severe constraint violations, high latency and low resource efficiency. This paper describes Aladdin, a novel cluster scheduler that can maximize resource efficiency while avoiding constraint violations: (i) it proposes a multidimensional and nonlinear capacity function to support constraint expressions; (ii) it applies an optimized maximum flow algorithm to improve resource efficiency. Experiments with an Alibaba workload trace from a 10,000-machine cluster show that Aladdin can reduce violated constraints by as mush as 20%. Meanwhile, it improves resource efficiency by 50% compared with state-of-the-art schedulers. Heng Wu 0001, Wenbo Zhang 0006, Yuanjia Xu, Tao Huang 0001, Haiyang Ding |
IPDPS | 5 |
| 2018 | Migrating Web Applications from Monolithic Structure to Microservices ArchitectureabstractIn the traditional software development and deployment, the centralized monolithic is always adopted, as the modules are tightly coupled, which caused many inconvenience in software DevOps. The modules with bottlenecks in monolithic application cannot be extend separately as the application is an integral part, and different module cannot use different technology stack. To prolong the lifecycle of the monolithic applications, its need to migrated it to microservice architecture. Due to the complex logic and large number of third party framework libraries depended, get an accurate comprehensive of the application characteristics is challenging. The existing research mostly based on the static characteristics, lack of consideration of the runtime dynamic characteristics, and the completeness and accuracy of the static analysis is inadequate. To resolve above problems, we combined static and dynamic analysis to get static structure and runtime behavior characteristics of monolithic application. We employed the coupling among functions to evaluate the degree of dependence, and through function clustering to achieve the migration of legacy monolithic applications and its data to microservices architecture. Through the empirical study of migrate the typical legacy project to microservices, it is proved that we proposed method can offer precise guidance and assistance in the migration procedure. Experiments show that the method has high accuracy and low performance cost. Zhongshan Ren, Wei Wang 0049, Guoquan Wu, Chushu Gao, Wei Chen 0018, Jun Wei 0001, Tao Huang 0001 |
Internetware | 7 |
| 2018 | How are spreadsheet templates used in practice: a case study on EnronabstractTo reduce the effort of creating similar spreadsheets, end users may create expected spreadsheets from some predesigned templates, which contain necessary table layouts (e.g., headers and styles) and formulas, other than from scratch. When there are no explicitly predesigned spreadsheet templates, end users often take an existing spreadsheet as the instance template to create a new spreadsheet. However, improper template design and usage can introduce various issues. For example, a formula error in the template can be easily propagated to all its instances without users’ noticing. Since template design and usage are rarely documented in literature and practice, practitioners and researchers lack understanding of them to achieve effective improvement. In this paper, we conduct the first empirical study on the design and the usage of spreadsheet templates based on 47 predesigned templates (490 instances in total), and 21 instance template groups (168 template and instance pairs in total), extracted from the Enron corpus. Our study reveals a number of spreadsheet template design and usage issues in practice, and also sheds lights on several interesting research directions. Wensheng Dou, Chushu Gao, Jun Wei 0001, Tao Huang 0001 |
ESEC/SIGSOFT FSE | 6 |
| 2018 | Detecting faulty empty cells in spreadsheetsabstractSpreadsheets play an important role in various business tasks, such as financial reports and data analysis. In spreadsheets, empty cells are widely used for different purposes, e.g., separating different tables, or default value "0". However, a user may delete a formula unintentionally, and leave a cell empty. Such ad-hoc modification may introduce a faulty empty cell that should have a formula. We observe that the context of an empty cell can help determine whether the empty cell is faulty. For example, is the empty cell next to a cell array in which all cells share the same semantics? Does the empty cell have headers similar to other non-empty cells'? In this paper, we propose EmptyCheck, to detect faulty empty cells in spreadsheets. By analyzing the context of an empty cell, EmptyCheck validates whether the cell belong to a cell array. If yes, the empty cell is faulty since it does not contain a formula. We evaluate EmptyCheck on 100 randomly sampled EUSES spreadsheets. The experimental result shows that EmptyCheck can detect faulty empty cells with high precision (75.00%) and recall (87.04%). Existing techniques can detect only 4.26% of the true faulty empty cells that EmptyCheck detects. Wensheng Dou, Chushu Gao, Jun Wei 0001, Tao Huang 0001 |
SANER | 7 |
| 2018 | IO dependent SSD cache allocation for elastic Hadoop applications
Wei Wang 0049, Yu Huang 0002, Heng Wu 0001, Jun Wei 0001, Tao Huang 0001 |
Sci. China Inf. Sci. | 7 |
| 2017 | Mining API Type Specifications for JavaScriptabstractAPI specifications play an important role in software development. However, API specifications are often not well documented, especially for JavaScript. Many JavaScript API specifications lack of precise type information for API parameters and return values. In this paper, we propose a static approach for mining JavaScript type specifications automatically. We gather the usage information of return values and parameters statically, and infer types of return values based their usages, by identifying a known type which they are used most likely to be, and infer parameters by identifying the most used parameters. We evaluate the approach on the homepages of Alexa top 1000 websites, the experimental results show that our approach can gain high precision. Our case study on jQuery shows that our approach gains high precision and reasonable recall on jQuery, and we can use our inferred API type specifications to detect 2 jQuery misusage errors in real-world web sites, and 1 missing type error in jQuery documentations. Wensheng Dou, Chushu Gao, Jun Wei 0001, Tao Huang 0001 |
APSEC | 5 |
| 2017 | AppCheck: A Crowdsourced Testing Service for Android ApplicationsabstractIt is well known that the fragmentation of Android ecosystem has caused severe compatibility issues. Therefore, for Android apps, cross-platform testing (the apps must be tested on a multitude of devices and operating system versions) is particularly important to assure their quality. Although lots of cross-platform testing techniques have been proposed, there are still some limitations: 1) it is time-consuming and error-prone to encode platform-agnostic tests manually, 2) test scripts generated by existing record/replay techniques are brittle and will break when replayed on different platforms, 3) Developers, and even test vendors have not equipped some special Android devices. As a result, apps have not been tested sufficiently, leading to many compatibility issues after releasing. To address these limitations, this paper proposes AppCheck, a crowdsourced testing service for Android apps. To generate tests that will explore different behavior of the app automatically, AppCheck crowdsources event trace collection over the Internet, and various touch events will be captured when real users interact with the app. The collected event traces are then transformed into platform-agnostic test scripts, and directly replayed on the devices of real users. During the replay, various data (e.g., screenshots and layout information) will be extracted to identify compatibility issues. Our empirical evaluation shows that AppCheck is effective and improves the state of the art. Guoquan Wu, Yuzhong Cao, Wei Chen 0018, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001 |
ICWS | 6 |
| 2017 | Application-centric SSD Cache Allocation for Hadoop ApplicationsabstractFlash-based Solid State Drive (SSD) is widely used in the virtualization environment, usually as the cache of the hard disk drive-based Virtual Machine (VM) storage, to improve the IO performance. Existing SSD caching schemes are mainly driven by VM-centric metrics. They treat the VMs as independent units and focus on critical low-level performance metrics of individual VMs, such as the working set, the IO latency, or the throughput. However, for elastic Hadoop applications consisting of multiple VMs, the workload is rapidly changing, and the importance of differnet VMs may be different even if they have the same low-level IO pattern. In this situation, the VM-centric SSD caching schemes may not lead to the best performance, i.e., the shortest job completion time. Considering the importance of VMs and relationships among VMs inside the application may potentially better improve the performance, which we regard as the application-centric metrics. We propose the Application-Centric SSD caching for Hadoop applications (ACSSD), which reduces the job completion time from the application level. AC-SSD uses the genetic algorithm based approach to calculate the nearly optimal weights of virtual machines for allocating SSD cache space and controlling the I/O Operations Per Second (IOPS) based on the importance of the VMs. Moreover, AC-SSD introduces the closed-loop adaptation to face the rapidly changing workload. The evaluation shows that AC-SSD reduces the job completion time by up to 39% for IO sensitive workloads, and up to 29% for rapidly changing workloads. Wei Wang 0049, Yu Huang 0002, Heng Wu 0001, Jun Wei 0001, Tao Huang 0001 |
Internetware | 6 |
| 2017 | SpreadCluster: recovering versioned spreadsheets through similarity-based clusteringabstractVersion information plays an important role in spreadsheet understanding, maintaining and quality improving. However, end users rarely use version control tools to document spreadsheets' version information. Thus, the spreadsheets' version information is missing, and different versions of a spreadsheet coexist as individual and similar spreadsheets. Existing approaches try to recover spreadsheet version information through clustering these similar spreadsheets based on spreadsheet filenames or related email conversation. However, the applicability and accuracy of existing clustering approaches are limited due to the necessary information (e.g., filenames and email conversation) is usually missing. We inspected the versioned spreadsheets in VEnron, which is extracted from the Enron Corporation. In VEnron, the different versions of a spreadsheet are clustered into an evolution group. We observed that the versioned spreadsheets in each evolution group exhibit certain common features (e.g., similar table headers and worksheet names). Based on this observation, we proposed an automatic clustering algorithm, SpreadCluster. SpreadCluster learns the criteria of features from the versioned spreadsheets in VEnron, and then automatically clusters spreadsheets with the similar features into the same evolution group. We applied SpreadCluster on all spreadsheets in the Enron corpus. The evaluation result shows that SpreadCluster could cluster spreadsheets with higher precision (78.5% vs. 59.8%) and recall rate (70.7% vs. 48.7%) than the filename-based approach used by VEnron. Based on the clustering result by SpreadCluster, we further created a new versioned spreadsheet corpus VEnron2, which is much bigger than VEnron (12,254 vs. 7,294 spreadsheets). We also applied SpreadCluster on the other two spreadsheet corpora FUSE and EUSES. The results show that SpreadCluster can cluster the versioned spreadsheets in these two corpora with high precision (91.0% and 79.8%). Wensheng Dou, Chushu Gao, Jie Wang 0035, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001 |
MSR | 7 |
| 2016 | Determine Configuration Entry Correlations for Web Application SystemsabstractWeb application systems, comprising of heterogeneous and loosely coupled components, are usually highly-configurable due to the large number of configuration entries scattering in the components. The dependencies between components lead their entries correlate to one another, which makes the system deployment and migration daunting and error-prone. For two correlated entries, changing value of one entry requires the value change of the other. Otherwise, some implied constraints would be violated and the system failure will occur. Keeping track of entry correlations, which is essential to system reliabilities, is not a simple work as it often crosses products and requires in-depth domain knowledge. This paper proposes a method to automate the process of determining entry correlations. The method first narrows down the exploring scale to those frequently-set entries based on crawled sample data. Then, it generates a correlation score for each entry pair, which is calculated according to entry names, values and inferred types. Thirdly, a set of heuristics are provided to determine a candidate set of the likely correlations. Finally, a rank-ordered list of entry correlations is output so that system administrators can consult it to check system configuration systematically. Based on the method, we implement a tool, Correlation Explorer, and make experiments and evaluations with some real world systems. The result shows that Correlation Explorer is effective in finding a large portion of entry correlations. Wei Chen 0018, Heng Wu 0001, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001 |
COMPSAC | 5 |
| 2016 | Crawling hidden objects with kNN queriesabstractWith rapidly growing popularity, Location Based Services (LBS), e.g., Google Maps, Yahoo Local, WeChat, FourSquare, etc., started offering web-based search features that resemble a kNN query interface. Specifically, for a user-specified query location q, these websites extract from the objects in their backend database the top-k nearest neighbors to q and return these k objects to the user through the web interface. Here k is often a small value like 50 or 100. For example, McDonald [1] returns the top 25 nearest restaurants for a user-specified location through its locations search webpage. Zhiguo Gong, Nan Zhang 0004, Tao Huang 0001, Hua Zhong 0001, Jun Wei 0001 |
ICDE | 4 |
| 2016 | Hug the Elephant: Migrating a Legacy Data Analytics Application to Hadoop EcosystemabstractBig data applications that rely on relational databases gradually expose limitations on scalability and performance. In recent years, Hadoop ecosystem has been widely adopted as an evolving solution. This paper presents the migration of a legacy data analytics application in a provincial data center. The target platform follows "no one size fits all" method. Considering different workloads, data storage is hybrid with distributed file system (HDFS) and distributed NoSQL database. Beyond the architecture re-design, we focus on the problem of data model transformation from relational database to NoSQL database. We propose a query-aware approach to free developers from tedious manual work. The approach generates query-specific views (NoView) for NoSQL and re-structures the views to align with NoSQL's data model. Our results show that the migrated application achieves high scalability and high performance. We believe that our practice provides valuable insights (such as NoSQL data modeling methodology), and the techniques can be easily applied to other similar migrations. Jie Liu 0008, Sa Wang, Lijie Xu, Jixin Ren, Dan Ye 0004, Jun Wei 0001, Tao Huang 0001 |
ICSME | 9 |
| 2016 | X-Check: A Novel Cross-Browser Testing Service Based on Record/ReplayabstractWith the advent of Web 2.0 application, and the increasing number of browsers and platforms on which the applications can be executed, cross-browser incompatibilities (XBIs) are becoming a serious problem for organizations to develop web-based software. Although some techniques and tools have been proposed to identify XBIs, they cannot assure the same execution when the application runs across different browsers as only explicit user activity is considered, and thus prone to generating both false positives and false negatives. To address this limitation, this paper describes X-Check, a platform that enables cross-browser testing as a service by leveraging record/replay technique. Comparing to existing techniques and tools, X-Check supports to detect cross-browser issues with high accuracy. It also provides useful support to developers for diagnosis and (eventually) elimination of XBIs. Our empirical evaluation shows that X-Check is effective, improves the state of the art. Meimei He, Guoquan Wu, Hongyin Tang, Wei Chen 0018, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001 |
ICWS | 7 |
| 2016 | Clustering-based acceleration for virtual machine image deduplication in the cloud environment
Wenbo Zhang 0006, Tao Wang 0030, Tao Huang 0001 |
J. Syst. Softw. | 5 |
| 2016 | Crawling Hidden Objects with kNN QueriesabstractMany websites offering Location Based Services (LBS) provide a$k$NN search interface that returns the top-$k$nearest-neighbor objects (e.g., nearest restaurants) for a given query location. This paper addresses the problem of crawling all objects efficiently from an LBS website, through the public$k$NN web search interface it provides. Specifically, we develop crawling algorithm for 2D and higher-dimensional spaces, respectively, and demonstrate through theoretical analysis that the overhead of our algorithms can be bounded by a function of the number of dimensions and the number of crawled objects, regardless of the underlying distributions of the objects. We also extend the algorithms to leverage scenarios where certain auxiliary information about the underlying data distribution, e.g., the population density of an area which is often positively correlated with the density of LBS objects, is available. Extensive experiments on real-world datasets demonstrate the superiority of our algorithms over the state-of-the-art competitors in the literature. Zhiguo Gong, Nan Zhang 0004, Tao Huang 0001, Hua Zhong 0001, Jun Wei 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2016 | FD4C: Automatic Fault Diagnosis Framework for Web Applications in Cloud ComputingabstractThe large-scale dynamic cloud computing environment has raised great challenges for fault diagnosis in Web applications: First, fluctuating workloads cause traditional application models to change over time; second, modeling the behaviors of complex applications usually requires domain knowledge which is difficult to obtain; third, managing large-scale applications manually is impractical for operators. To address these issues, this paper proposes an automatic fault (F) diagnosis (D) framework for (4) Web applications in cloud (C) computing (FD4C). In this paper, we propose an online incremental clustering method to recognize access behavior patterns. We also use correlation analysis to model the correlations between the workloads and application performance/resource utilization metrics in a specific access behavior pattern. FD4C detects faults by discovering the abrupt changes of correlation coefficients with control charts. Then, FD4C identifies the fault-related metrics using a feature selection method. To evaluate our proposal, we inject typical faults into TPC-W benchmark and apply FD4C to diagnose the injected faults. The experimental results show that FD4C can effectively detect the typical faults and accurately locate the metrics related to the faults. Tao Wang 0030, Wenbo Zhang 0006, Chunyang Ye, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2015 | Discovering User-Defined Event Handlers in Presence of JavaScript LibrariesabstractJavaScript libraries, such as JQuery, are widely used in web applications. In these libraries' event delegation models, a DOM element's event handler is usually bound to its parent nodes. This makes it difficult for developers to figure out the user-defined event handlers of a specified DOM element. In this paper, we propose an approach that identifies the user-defined event handlers of DOM elements in a web page. We dynamically collect the execution trace for each triggered event in a web page, and analyze how each function is used in the execution trace to discover the event handlers for each event. We evaluate our approach on seven real-world web applications. The result shows that our approach is effective, with an overall precision of 100% and recall of 99.8%. Wensheng Dou, Chushu Gao, Jun Wei 0001, Tao Huang 0001 |
APSEC | 5 |
| 2015 | A Lightweight Evaluation Framework for Table Layouts in MapReduce Based Query Systems
Jie Liu 0008, Lijie Xu, Dan Ye 0004, Jun Wei 0001, Tao Huang 0001 |
APWeb | 6 |
| 2015 | VMon: Monitoring and Quantifying Virtual Machine Interference via Hardware Performance CounterabstractVirtualization greatly improves resource utilization in IaaS platforms, but it also introduces potential interference between virtual machines (VMs). For example, VMs may suffer from performance degradation, when they are located in one host and compete for sharing physical resources. Thus, how to efficiently monitor and quantify the VMs interference becomes a key challenge for IaaS providers. In this paper, we present Vmon, a system to transparently monitor and quantify the interference between VMs with the hardware performance counters (HPCs). By collecting the HPCs of different VMs and exploring the LLC miss rates within HPCs, Vmon analyzes the relationship between the LLC miss rates and VM performance degradation to predict the interference between different resource-intensive VMs, and mitigate the VMs interference. The experimental results show that Vmon predicts the performance degradation in the accuracy of more than 90% with less than 10% performance overhead. Sa Wang, Wenbo Zhang 0006, Tao Wang 0030, Chunyang Ye, Tao Huang 0001 |
COMPSAC | 5 |
| 2015 | Towards Web Application Mobilization via Efficient Web Control ExtractionabstractTraditional web applications are not suitable for mobile devices, because mobile devices are usually equipped with small screens and use slow and expensive mobile network. In order to adapt web applications to mobile devices, existing approaches reconstruct particular web applications, or adapt only partial views of web pages. They require a lot of additional reconstructing work or network bandwidth. In this paper we propose an approach that can extract a part of a web page as an executable web control efficiently. Our approach monitors the execution of user code, builds a dependency graph of executed user code, and performs slicing based on the dependency graph. The evaluation on two real-world web applications shows that our approach is able to extract executable web controls efficiently, and for the two web applications, visiting extracted web controls instead of the original web pages can save 98% and 23% of bandwidth respectively. Wensheng Dou, Guoquan Wu, Jie Wang 0035, Chushu Gao, Jun Wei 0001, Tao Huang 0001 |
Internetware | 7 |
| 2015 | Aggregate Estimation in Hidden Databases with Checkbox InterfacesabstractA large number of web data repositories are hidden behind restrictive web interfaces, making it an important challenge to enable data analytics over these hidden web databases. Most existing techniques assume a form-like web interface which consists solely of categorical attributes (or numeric ones that can be discretized). Nonetheless, many real-world web interfaces (of hidden databases) also feature checkbox interfaces-e.g., the specification of a set of desired features, such as A/C, navigation, etc., for a car-search website like Yahoo! Autos. We find that, for the purpose of data analytics, such checkbox-represented attributes differ fundamentally from the categorical/numerical ones that were traditionally studied. In this paper, we address the problem of data analytics over hidden databases with checkbox interfaces. Extensive experiments on both synthetic and real datasets demonstrate the accuracy and efficiency of our proposed algorithms. Zhiguo Gong, Nan Zhang 0004, Tao Huang 0001, Hua Zhong 0001, Jun Wei 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2014 | A Lightweight Virtual Machine Image Deduplication Backup Approach in Cloud EnvironmentabstractAs most clouds are based on virtualization technology, more and more virtual machine images are created within data centers. Depending on the need of disaster recovery, the storage space used for backup would easily sprawl to a TB or PB level with the growth of images. Unfortunately, different images have a large amount of same data segments. Those duplicated data segments will lead to serious waste of storage resource. Although there is a lot of work focus on deduplication storage and could achieve a good result in removing duplicate copies, they are not very suitable for virtual machine image deduplication in a cloud environment. Because huge resource usage of deduplication operations could lead to serious performance interference to the hosting virtual machines. This paper propose a local deduplication method which can speed up the operation progress of virtual machine image deduplication and reduce the operation time. The method is based on an improved k-means clustering algorithm, which could classify the metadata of backup image to reduce the search space of index lookup and improve the index lookup performance. Experiments show that our approach is robust and effective. It can significantly reduce the performance interference to hosting virtual machine with an acceptable increase in disk space usage. Wenbo Zhang 0006, Shiyang Ye, Jun Wei 0001, Tao Huang 0001 |
COMPSAC | 5 |
| 2014 | Inferring Data Contract for Web-Based APIabstractWeb-based API is a new trend for publishing services. To correctly use the API, developers should follow certain service specifications. Data contract is a service specification to express the constraints over the data model used in the APIs. Data contracts, however are not always readily available in a formalized format if not undocumented at all. In this paper, we present an approach to infer formal data contracts for Web-based API. The approach integrates information of the parameters, error messages and testing result of Web-based API. We demonstrate how this approach infers complicated data preconditions for Web-based API in the real-world Web API platforms. Chushu Gao, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001 |
ICWS | 4 |
| 2014 | Runtime Enforcement of Data-centric Properties for Concurrent Service-Based ApplicationsabstractFor service-based applications which are composed of multiple independent third-parties, continuous monitoring is required to assure that runtime behavior of the systems complies with specified properties. However, most existing work only detects the violation while not consider how to enforce the properties so that the constraint can not be violated at runtime. To address this limitation, this paper presents EnforceBCL, a framework for enforcing data-centric properties for concurrent service-based applications. Users of EnforceBCL can specify the properties to be enforced using the expressive behavior constraint enforcement language. Data-centric property is enforced at runtime by blocking the process whose next action would violate it. The impacted processes can be unblocked and allowed to execute when the specified property eventually reaches a safe state. EnforceBCL also provides the mechanism to detect possible deadlock during the enforcement of the property, and executes corresponding handler to solve the deadlock. To evaluate the effectiveness and efficiency of the proposed approach, we conducted several experiments. Results show that EnforceBCL is able to effectively enforce data-centric properties for concurrent service-based applications and also incurs less performance overhead. Guoquan Wu, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001 |
ICWS | 4 |
| 2014 | MC-Checker: Detecting Memory Consistency Errors in MPI One-Sided ApplicationsabstractOne-sided communication decouples data movement and synchronization by providing support for asynchronous reads and updates of distributed shared data. While such interfaces can be extremely efficient, they also impose challenges in properly performing asynchronous accesses to shared data. This paper presents MC-Checker, a new tool that detects memory consistency errors in MPI one-sided applications. MCChecker first performs online instrumentation and captures relevant dynamic events, such as one-sided communications and load/store operations. MC-Checker then performs analysis to detect memory consistency errors. When found, errors are reported along with useful diagnostic information. Experiments indicate that MC-Checker is effective at detecting and diagnosing memory consistency bugs in MPI one-sided applications, with low overhead, ranging from 24.6% to 71.1%, with an average of 45.2%. Zhezhe Chen, James Dinan, Pavan Balaji, Hua Zhong 0001, Jun Wei 0001, Tao Huang 0001 |
SC | 7 |
| 2014 | Workload-aware anomaly detection for Web applications
Tao Wang 0030, Jun Wei 0001, Wenbo Zhang 0006, Hua Zhong 0001, Tao Huang 0001 |
J. Syst. Softw. | 5 |
| 2013 | VM image update notification mechanism based on pub/sub paradigm in cloudabstractVirtual machine image encapsulates the whole software stack including operating system, middleware, user application and other software products. Failure occurred in any layer of the software stack will be treated as image failure. However, virtual machine image with potential failures can be convert to template and spread to a wide range by means of template replication. And this paper refer to this phenomenon as "image failure propagation". Usually, patching is a widely adopted solution to resolve software failures. Nevertheless, virtual machine image patches are difficult to deliver to the final users in cloud computing environment for its openness and multi-tenancy features. This paper described image failure propagation model for the first time and proposed a promoting mechanism based on pub/sub computing paradigm to combat with the patching delivery problem. Shiyang Ye, Wenbo Zhang 0006, Jun Wei 0001, Tao Huang 0001 |
Internetware | 5 |
| 2013 | Detecting performance anomaly with correlation analysis for Internetware
Tao Wang 0030, Jun Wei 0001, Wenbo Zhang 0006, Hua Zhong 0001, Tao Huang 0001 |
Sci. China Inf. Sci. | 6 |
| 2013 | A benefit-aware on-demand provisioning approach for multi-tier applications in cloud computing
Heng Wu 0001, Wenbo Zhang 0006, Jun Wei 0001, Tao Huang 0001 |
Frontiers Comput. Sci. | 5 |
| 2012 | A Profit-Aware Virtual Machine Deployment Optimization Framework for Cloud Platform ProvidersabstractAs a rising application paradigm, cloud computing enables the resources to be virtualized and shared among applications. In a typical cloud computing scenario, customers, Service Providers (SP), and Platform Providers (PP) are independent participants, and they have their own objectives with different revenues and costs. From PPs' viewpoints, much research work reduced the costs by optimizing VM placement and deciding when and how to perform the VM migrations. However, some work ignored the fact that the balanced use of the multi-dimensional resources can affect overall resource utilization significantly. Furthermore, some work focuses on the selection of the VMs and the target servers without considering how to perform the reconfigurations. In this paper, with a comprehensive consideration of PPs' interests, we propose a framework to improve their profits by maximizing the resource utilization and reducing the reconfiguration costs. Firstly, we use the vector arithmetic to model the objective of balancing the multi-dimensional resources use and propose a VM deployment optimization method to maximize the resource utilization. Then a two-level runtime reconfiguration strategy, including local adjustment and VM parallel migration, is presented to reduce the VM migration and shorten the total migration time. Finally, we conduct some preliminary experiments, and the results show that our framework is effective in maximizing the resource utilization and reducing the costs of the runtime reconfiguration. Wei Chen 0018, Xiaoqiang Qiao, Jun Wei 0001, Tao Huang 0001 |
IEEE CLOUD | 4 |
| 2012 | Optimizing data migration for cloud-based key-value storesabstractAs one database offloading strategy, elastic key-value stores are often introduced to speed up the application performance with dynamic scalability. Since the workload is varied, efficient data migration with minimal impact in service is critical for the issue of elasticity and scalability. However, due to the new virtualization technology, real-time and low-latency requirements, data migration within cloud-based key-value stores has to face new challenges: effects of VM interference, and the need to trade off between the two ingredients of migration cost, namely migration time and performance impact. To fulfill these challenges, in this paper we explore a new approach to optimize the data migration. Explicitly, we build two interference-aware models to predict the migration time and performance impact for each migration action using statistical machine learning, and then create a cost model to strike a balance between the two ingredients. Using the load rebalancing scenario as a case study, we have designed one cost-aware migration algorithm that utilizes the cost model to guide the choice of possible migration actions. Finally, we demonstrate the effectiveness of the approach using Yahoo! Cloud Serving Benchmark (YCSB). Xiulei Qin, Wenbo Zhang 0006, Wei Wang 0049, Jun Wei 0001, Tao Huang 0001 |
CIKM | 6 |
| 2012 | Towards a Cost-Aware Data Migration Approach for Key-Value StoresabstractLive data migration is an important technique for key-value stores. However, due to the stateful feature, new virtualization technology, stringent low latency requirements and unexpected workload changes, key-value stores deployed in cloud environment have to face new challenges for data migration: effects of VM interference, and the need to trade off between the two ingredients of migration cost, say migration time and performance impact. To address these challenges, we focus on the data migration problem in a load rebalancing scenario and build a new framework that aims to rebalance load while minimizing migration costs. We build two interference-aware prediction models to predict the migration time and performance impact for each action using statistical machine learning and then create a cost model to strike a right balance between the two ingredients of cost. A cost-aware migration algorithm is designed to utilize the cost model and balance rate to guide the choice of possible migration actions. We demonstrate the effectiveness of the data migration approach as well as the cost model and two prediction models using YCSB. Xiulei Qin, Wenbo Zhang 0006, Wei Wang 0049, Jun Wei 0001, Tao Huang 0001 |
CLUSTER | 6 |
| 2012 | Elasticat: A load rebalancing framework for cloud-based key-value storesabstractThe problem of load rebalancing is an important issue for cloud-based key-value stores. However, the new virtualization environment and the store's stateful feature make this classical issue more challenging. In this paper, we build a new load rebalancing framework for cloud-based key-value stores, namely ElastiCat. It can be used for auto reconfiguring the store system with minimal costs and no disruption to the availability of the service. To evaluate and minimize the rebalancing costs, we firstly build two interference-aware prediction models to predict the data migration time and performance impact for each action using statistical machine learning and then create a cost model to strike a right balance between them. A cost-aware rebalancing algorithm is designed to utilize the cost model and balance rate to create a rebalancing plan and guide the choice of possible rebalancing actions. To maintain the availability of storage service, we propose a lightweight piggy-back based data access protocol. Finally, we demonstrate the effectiveness of the framework as well as the cost model using YCSB. Xiulei Qin, Wei Wang 0049, Wenbo Zhang 0006, Jun Wei 0001, Tao Huang 0001 |
HiPC | 6 |
| 2012 | Specification and monitoring of data-centric temporal properties for service-based systems
Guoquan Wu, Jun Wei 0001, Chunyang Ye, Hua Zhong 0001, Tao Huang 0001, Hong He 0004 |
J. Syst. Softw. | 5 |
| 2011 | An Adaptive Performance Modeling Approach to Performance Profiling of Multi-service Web ApplicationsabstractThe performance of multi-service applications are known to be determined mainly by the interactions between workload and behaviors of the application. The change of workload can lead to dynamic service demands on system resources, and even cause dynamic bottleneck switches between services inside the application. In this paper, to profiling large-applications' behaviors, and help to locate the bottleneck and optimize their capacities, we focus on modeling their behavior according to the workload. Although this topic has been well studied at testing stage, building such a model under live workload remains a challenge, because the workload and application behaviors are time-varying. To tackle this problem, we propose an adaptive approach to build and rebuild performance model according to log files. Both the user behaviors and their corresponding internal service relations are modeled, and the CPU time consumed by each service is also obtained through Kalman filter, which can "absorb" some level of noise in real-world data. Our model can explain the behaviors of both the whole application and the individual services, and provide valuable information for capacity planning and bottleneck detection. At last, our work is evaluated with TPC-W bench mark, whose results can demonstrate the effectiveness of our approach. Xiang Huang 0005, Wei Wang 0049, Wenbo Zhang 0006, Jun Wei 0001, Tao Huang 0001 |
COMPSAC | 5 |
| 2011 | On-line Cache Strategy Reconfiguration for Elastic Caching Platform: A Machine Learning ApproachabstractCloud computing provide scalability and high availability for web applications using such techniques as distributed caching and clustering. As one database offloading strategy, elastic caching platforms (ECPs) are introduced to speed up the performance or handle application state management with fault tolerance. Several cache strategies for ECPs have been proposed, say replicated strategy, partitioned strategy and near strategy. We first evaluate the impact of the three cache strategies using the TPC-W benchmark and find that there is no single cache strategy suitable for all conditions, the selection of the best strategy is related with workload patterns, cluster size and the number of concurrent users. This raises the question of when and how the cache strategy should be reconfigured as the condition varies which has received comparatively less attention. In this paper, we present a machine learning based approach to solving this problem. The key features of the approach are off-line training coupled with on-line system monitoring and robust synchronization process after triggering a reconfiguration, at the same time the performance model is periodically updated. More explicitly, first a rule set used to identify which cache strategy is optimal under the current condition are trained with the system statistics and performance results. We then introduce a framework to switch the cache strategy on-line as the workload varies and keep its overhead to acceptable levels. Finally, we illustrate the advantages of this approach by carrying out a set of experiments. Xiulei Qin, Wenbo Zhang 0006, Wei Wang 0049, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001 |
COMPSAC | 6 |
| 2011 | A Statistical Approach for Estimating CPU Consumption in Shared Java Middleware ServerabstractMiddleware sharing is one of the important resource sharing approaches which enables sharing of costs across a large pool of users. However, the shared Java middleware server easily causes interference on performance between concurrent user requests. A key requirement to an effective performance isolation is the knowledge of the resource consumption of the various kinds of use requests classified according to different application context information. Direct measurement of resource consumption requires instrumentation which is impractical. In this paper, we demonstrate that CPU consumptions of various kinds of user requests on a given hardware can be approximated by a proposed Kalman filter based approach. Experimental results derived from testing the approach by using the TPC-W e-commerce suite deployed on a widely-used Java middleware server (Tomcat) illustrate the potential of this approach. Wei Wang 0049, Xiang Huang 0005, Yunkui Song, Wenbo Zhang 0006, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001 |
COMPSAC | 7 |
| 2011 | Runtime Monitoring of Data-centric Temporal Properties for Web ServicesabstractRuntime monitoring of Web service compositions has been widely acknowledged as a significant approach to understand and guarantee the quality of services. However, existing runtime monitoring solutions consider only the constraints on the sequence of messages exchanged between partner services and ignore the actual data contents inside the messages. As a result, it is difficult to monitor some dynamic properties such as how message data of interest is processed between different participants. To address this issue, we propose an efficient, non-intrusive online monitoring approach to dynamically analyze data-centric properties for service-oriented applications involving multiple participants. By introducing Par-BCL - a Parametric Behavior Constraint Language for web services - to define monitoring parameters, various data-centric temporal behavior properties for Web services can be specified and monitored. This approach broadens the monitored patterns to include not only message exchange orders, but also the data contents bound to the parameters. To reduce runtime overhead, we statically analyze the monitored properties to generate parameter state machine from the event pattern automata to optimize monitoring. The experiments show that our solution is efficient and promising. Guoquan Wu, Jun Wei 0001, Chunyang Ye, Xiaozhe Shao, Hua Zhong 0001, Tao Huang 0001 |
ICWS | 6 |
| 2011 | Runtime Verification of Data-Centric Properties in Service Based Systems
Guoquan Wu, Jun Wei 0001, Chunyang Ye, Xiaozhe Shao, Hua Zhong 0001, Tao Huang 0001 |
RV | 6 |
| 2011 | Stable cohesion metrics for evolving ontologiesabstractAbstract With the drastic development of semantic‐driven applications, assessing the quality of ontologies has received more attention. Measuring and assessing the quality of ontologies can help ontology engineers to control project management and reduce the risk of project failures. However, most of the existing ontology metrics for measuring and assessing the quality of ontologies are defined based on ontology structure, and neglect the stability of ontology measurement. In this paper, we concentrate on stable ontology measurement by using semantically derived ontology metrics. We propose four ontology cohesion metrics, which fully consider the implicitly expressed semantic information and are defined based on ontological semantics rather than ontology structure. Before measuring and assessing an ontology, we materialize a pre‐processing for stable ontology measurement by treating the ontology. The proposed ontology cohesion metrics are theoretically validated by the validation criteria of object‐oriented software. The experimental results show that we can successfully collect more semantic knowledge from the testing ontologies for stable ontology measurement by using the proposed ontology cohesion metrics. The ontology cohesion metrics proposed in this paper can be reasonably used as a cogent complementarity of existing ontology metrics. Copyright © 2010 John Wiley & Sons, Ltd. Yinglong Ma 0001, Haijiang Wu, Beihong Jin, Tao Huang 0001, Jun Wei 0001 |
J. Softw. Maintenance Res. Pract. | 5 |
| 2010 | A Two-Phase Approach to Subscription Subsumption Checking for Content-Based Publish/Subscribe SystemsabstractThe efficiency of subscription subsumption checking remains a key issue for content-based publish/subscribe systems. In this paper, we propose an efficient data structure called subscription subsumption graph (SSG). This data structure could differentiate the two types of subsumption relationships and help speed up the process of subsumption checking and subscription cancellation. We then present a two-phase approach to subscription subsumption checking. Phase one is mainly about checking of non-numeric constraints by using an index structure which could help filter out most of irrelevant subscriptions while phase two is about checking of remaining numeric constraints where SSG is employed. Finally, we introduce an efficient SSG-based unsubscription algorithm that could find out which subscriptions need to be forwarded without any redundant computing. We illustrate the advantages of this approach by carrying out extensive experiments. Xiulei Qin, Jun Wei 0001, Wenbo Zhang 0006, Hua Zhong 0001, Tao Huang 0001 |
AINA | 5 |
| 2010 | Detecting Data Inconsistency Failure of Composite Web Services Through Parametric Stateful AspectabstractRuntime monitoring of Web service compositions with WS-BPEL has been widely acknowledged as a significant approach to understand and guarantee the quality of services. However, most existing monitoring technologies only track patterns related to the execution of an individual process. As a result, the possible inconsistency failure caused by implicit interactions among concurrent process instances cannot be detected. To address this issue, this paper proposes an approach to specify the behavior properties related to shared resources for web service compositions and verify their consistency with the aid of a parametric stateful aspect extension to WS-BPEL. Parameters are introduced in pattern specification, which allows monitoring not only events but also their values bound to the parameters at runtime to keep track of data flow among concurrent process instances. An efficient implementation is also provided to reduce the runtime overhead of monitoring and event observation. Our experiments show that the proposed approach is promising. Guoquan Wu, Jun Wei 0001, Chunyang Ye, Hua Zhong 0001, Tao Huang 0001 |
ICWS | 5 |
| 2010 | An adaptive fine-grained performance modeling approach for internetwareabstractWith the great success of internet technology, internetware has become one of the most important software paradigms. But the open, dynamic and uncertain network makes it difficult to guarantee the performance of internetwares. Feed forward control method has been proved to be an effective mechanism for performance guarantee in advance, but it is difficult to work well in such a dynamic environment, in which performance aspects are highly changeable because for the load fluctuation and software updates. In this paper, we proposed an adaptive performance modeling approach to adapt the environment and provide fine-grained performance guarantee. In our approach, the service invocation sequences corresponding to the load of internetware are constructed adaptively. And the service time of each service, which is the most performance parameter of our performance tool, is accurately acquired through Kalman filter. Xiang Huang 0005, Wei Wang 0049, Wenbo Zhang 0006, Jun Wei 0001, Tao Huang 0001 |
Internetware | 5 |
| 2010 | A new approach to performance optimization of mashups via data flow refactoringabstractMashup tools allow end users graphically build complex mashups using pipes to connect web data sources into a data flow. Because end users are of poor technical expertise, the designed data flows may be inefficient. This paper targets on enhancing the performance of mashups via automatically refactoring the structure of its data flows. First a set of operational semantics features are selected for annotating the operators in data flows and refactoring rules are defined to generate all candidate semantics equivalent data flows. Then a heuristic algorithm is described for accurately searching the data flow of minimal execution time by constructing a partially ordered set of data flows based on their cost estimation. This approach is applicable to general mashup data flows without knowing complete operational semantics of their operators and the efficiency improvement is demonstrated by experiments. Jie Liu 0008, Jun Wei 0001, Dan Ye 0004, Tao Huang 0001 |
Internetware | 4 |
| 2010 | Middleware support for internetware: a service perspectiveabstractThe advent of Internet technology introduces a revolution to software application and development paradigms. Traditional software development and application patterns have been shifted to Internet-based service sharing and collaboration among partners all over the Internet. This imposes new challenges and complexity in the lifecycle of software development, deployment and maintenance. Middleware, an intermediate layer to abstract the homogeneity and hide the difference of underlying systems, can be used to reduce the complexity for Internet application development. In this paper, we exploit the needs of middleware support for Internet-based applications from a service perspective. We investigate the potential requirements and features of Internetware, and the state-of-the-art solutions. We also analyze the remaining issues, the challenges and potential future research directions. Chunyang Ye, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001 |
Internetware | 4 |
| 2009 | ETL Workflow Analysis and Verification Using Backwards Constraint Propagation
Jie Liu 0008, Senlin Liang, Dan Ye 0004, Jun Wei 0001, Tao Huang 0001 |
CAiSE | 5 |
| 2009 | Towards self-healing web services compositionabstractTo achieve self-healing web services composition, much work has been studied in the area of web services composition recently. However, most work addresses the problem of runtime monitoring, diagnosis and recovery in isolation. What is missing, however, is a unified solution that can be used to tackle this challenge in a principled manner. This paper presents a fresh view on self-healing web services composition. In particular, rather than building baseline system model a priori, we advocate using statistical learning theory(SLT) technique to extract it by observing the behavior of web services composition and locate the potential anomaly. Guoquan Wu, Jun Wei 0001, Tao Huang 0001 |
Internetware | 3 |
| 2009 | A study on the replaceability of context-aware middlewareabstractIn context-ware computing paradigm, context-aware middleware plays a key role. The middleware collects and manipulates contexts from environments, providing context-aware applications well-defined interfaces to adapt their behaviors when environments change. However, some minor difference in the implementation of context-aware middleware may cause the same context-aware application behave differently. Such behavior deviation may lead to serious problems or even disasters for a context-aware application. It is thus desirable to check whether a mobile context-aware application behaves consistently before moving it from one middleware to another, or whether a context-aware application still works correctly when upgrading the underlying middleware? Existing approaches for context-aware applications are not adequate for detecting such behavior deviation because these approaches do not consider the impacts of the difference in the middleware implementation. In this paper, we study the strategies in the implementation of context-aware middleware and their impacts on the behavior of context-aware applications. By exploring the implied scenarios where a context-aware application may behave differently, new testing approach is proposed to detect the behavior deviation of a context-aware application running on different middleware by generating test cases to cover these implied scenarios. Chunyang Ye, Shing-Chi Cheung, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001 |
Internetware | 5 |
| 2009 | Runtime Monitoring CompositeWeb Services Through Stateful Aspect Extension
Tao Huang 0001, Guoquan Wu, Jun Wei 0001 |
J. Comput. Sci. Technol. | 1 |
| 2008 | Flexible Pattern Monitoring for WS-BPEL through Stateful Aspect ExtensionabstractThe execution of composite web services with WS-BPEL relies on externally autonomous Web services. This implies the need to constantly monitor the running behavior of the involved parties. Moreover, monitoring the execution of such processes is critical to enforce business policies and meet reliability goals. This paper proposes a stateful aspect extension to WS-BPEL, as a solution to support flexible behavior pattern monitoring for composite Web services. Specifically, in the stateful aspect, history-based pointcut specifies the pattern of interest within a range, while advice describes the associated action to manage the process if the specified pattern occurs. We also present its implementation based on finite state automata through runtime weaving mechanism. Our experiments indicate the proposed monitoring approach incurs minimal overhead. Guoquan Wu, Jun Wei 0001, Tao Huang 0001 |
ICWS | 3 |
| 2008 | Efficient Approach for Web Services Selection with Multi-QOS ConstraintsabstractWith the increasing number of Web Services with similar or identical functionality, the non-functional properties of a Web Service will become more and more important. Hence, a choice needs to be made to determine which services are to participate in a given composite service. In general, multi-QoS constrained Web Services composition, with or without optimization, is a NP-complete problem on computational complexity that cannot be exactly solved in polynomial time. A lot of heuristics and approximation algorithms with polynomial- and pseudo-polynomial-time complexities have been designed to deal with this problem. However, these approaches suffer from excessive computational complexities that cannot be used for service composition in runtime. In this paper, we propose a efficient approach for multi-QoS constrained Web Services selection. Firstly, a user preference model was proposed to collect the user's preference. And then, a correlation model of candidate services are established in order to reduce the search space. Based on these two model, a heuristic algorithm is then proposed to find a feasible solution for multi-QoS constrained Web Services selection with high performance and high precision. The experimental results show that the proposed approach can achieve the expecting goal. Tao Huang 0001, Jun Wei 0001 |
Int. J. Cooperative Inf. Syst. | 1 |
| 2007 | High Performance Approach for Multi-QoS Constrained Web Services Selection
Jun Wei 0001, Tao Huang 0001 |
ICSOC | 3 |
| 2007 | Sequential Pattern-Based Cache Replacement in Servlet Container
Lin Zuo, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001 |
ICWE | 5 |
| 2007 | An Ontology-Based Approach for Semantic Conflict Resolution in Database Integration
Tao Huang 0001, Shaohua Liu 0002, Hua Zhong 0007 |
J. Comput. Sci. Technol. | 2 |
| 2006 | User-Defined Atomicity Constraint: A More Flexible Transaction Model for Reliable Service Composition
Xiaoning Ding, Jun Wei 0001, Tao Huang 0001 |
ICFEM | 3 |
| 2006 | Declarative Performance Modeling for Component-Based System using UML Profile for Schedulability, Performance and TimeabstractIn this paper we propose a declarative method for modeling the performance impact of container middleware to component-based system. The method abstracts the major performance impact factors, and these factors are modeled as sub-models by using UML activity diagram with UML profile for schedulability, performance and time annotations. Through a model assembly descriptor file associated with specific component application, the model assembling algorithm proposed in this paper can automatically composite the performance impacting factors into component application UML model. The declarative method helps to realize the separation of concerns between component application performance modeling and middleware impact. Using a case study, we validate proposed method Tao Huang 0001, Jun Wei 0001 |
SEFM | 2 |
| 2006 | Performance Evaluation of Component System based on Container style Middleware
Ningjiang Chen, Jun Wei 0001, Tao Huang 0001 |
SEKE | 4 |
| 2006 | An application-semantics-based relaxed transaction model for internetware
Tao Huang 0001, Xiaoning Ding, Jun Wei 0001 |
Sci. China Ser. F Inf. Sci. | 1 |
| 2006 | Extending Interactive Web Services for Improving Presentation Level Integration in Web Portals
Jingyu Song, Jun Wei 0001, Shuchao Wan, Tao Huang 0001 |
J. Comput. Sci. Technol. | 4 |
| 2005 | An Extended Event Matching Approach in Content-based Pub/Sub Systems for EAIabstractContent-based publish/subscribe offers a convenient abstraction for information producers and consumers, supporting a large-scale system design and evolution by integrating several distributed independent application systems. Unlike in the traditional address-based unicast or multicast, its core problem is how to match events by predicates on the content of events. In existing matching approaches, matching predicates are composed by the conjunction and disjunction of non-semantic constraints. But, in context of enterprise application integration, although they can match events by their contents, this traditional matching predicates are not expressive enough in manipulating the complex event matching, such as the "one-to-many" and "many-to-one" matching. Therefore, traditional matching approaches should be extended to solve the complex matching problems. After analyzing information matching patterns in enterprise application integration, we propose three matching models, extend this simple matching approach to the multi-semantic matching approach and further introduce the temporal constraint variable. The multi-semantic matching approach allows using different operations in accordance with different semantics; the temporal constraint variable supports processing several discrete events in temporal sequences. Then, we extend OBDD graphs into hierarchy coloured OBDD graphs and prove the equivalence of the transformation. Based on the extended OBDD graphs, the composite matching algorithm is presented and analysed. By experiments, we show the proposed algorithm is efficient. Tao Huang 0001 |
EDOC | 3 |
| 2005 | A QoS-enable failure detection framework for J2EE application serverabstractAs a basic reliability guarantee technology in distributed systems, failure detection provides the ability of timely detecting the liveliness of runtime systems. Effective failure detection is very important to J2EE application server (JAS), the leading middleware in Web computing environment, and it also needs to meet the requirements of reconfiguration, flexibility and adaptability. Based on the QoS (quality of service) specification of failure detector, this paper presents a QoS-enable failure detection framework for JAS, which satisfies the requirements of dynamically adjusting qualities and flexible integration of failure detectors. The work has been implemented in OnceAS application server that is developed by Institute of Software, Chinese Academy of Sciences. The experiments show that the framework can provide good QoS of failure detection in JAS. Ningjiang Chen, Jun Wei 0001, Tao Huang 0001 |
ISADS | 5 |
| 2005 | Extending OBDD Graphs for Composite Event Matching in Content-based Pub/Sub SystemsabstractContent-based publish/subscribe offers a convenient abstraction for the information producers and consumers, supporting a large-scale system design and evolution by integrating several distributed independent application systems. Unlike in the traditional address-based unicast or multicast, its core problem is how to match events by predicates on the content of events. In existing matching approaches, matching predicates are composed by the conjunction and disjunction of non-semantic constraints. But, in context of enterprise application integration, although they can match events by their contents, this traditional matching predicates are not enough expressive in manipulating the complex event matching, such as the "one-to-many" and "many-to-one" matching. Therefore, traditional matching approaches should be extended to solve the complex matching problems. After analyzing information matching patterns in enterprise application integration, we propose three matching models, extend the simple matching to the multi-semantic matching and introduce the temporal constraint variable. The multi-semantic matching allows using different operations in accordance with different semantics; the temporal constraint variable supports processing the discrete events in the temporal sequence. Then, we extend OBDD graphs into hierarchy coloured OBDD graphs and prove the equivalence of the transformation. Based on the extended OBDD graphs, the composite matching algorithm is presented and analysed. By experiments, we show the proposed algorithm is efficient Tao Huang 0001 |
ISPDC | 3 |
| 2004 | Performance Tuning for Application Server OnceAS
Wenbo Zhang 0006, Beihong Jin, Ningjiang Chen, Tao Huang 0001 |
ISPA | 5 |