VLDB 2026 Research / reviewers in the wild / expert
Wei Song 0003
dblp:62/1539-3
· DBLP profile ↗
65ranked-venue papers
22as first author
29since 2021 · last 2026
0000-0002-4324-3382ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 49 · 14 first-author · 25 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Computer networks · 2Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Consumption-Generation Driven Deep Embedding Framework for Next Activity PredictionabstractNext activity prediction, a critical subtask of Predictive Process Monitoring (PPM), aims to forecast the upcoming activity of an ongoing process instance based on event logs generated by process-aware information systems, thereby enabling businesses to proactively mitigate potential risks. While recent deep sequence models have shown strong predictive capabilities, they often fail to explicitly capture the intrinsic dependencies that govern the conditions and impacts of event occurrences in business processes. To address this limitation, we propose a novel Place-Transition Embedding (PTE) framework that captures these dependencies through Petri-net–inspired token consumption and generation semantics. PTE employs a deep embedding network to simulate the evolution of markings, i.e., how each observed event drives token dynamics over places over time under our resource interpretation. In addition, we introduce a temporal scoring mechanism that ranks candidate activities by jointly considering resource feasibility and temporal urgency, improving both interpretability and predictive accuracy. Extensive experiments on six real-world datasets show that PTE achieves state-of-the-art performance for next-activity prediction. Moreover, PTE generalizes robustly under sparse training conditions, highlighting its practical applicability in real-world scenarios with limited business data. Jian Wang 0018, Maodong Li 0006, Bing Li 0010, Wei Song 0003 |
IEEE Trans. Serv. Comput. | 5 |
| 2025 | Finding Compiler Bugs through Cross-Language Code Generator and Differential TestingabstractCompilers play a central role in translating high-level code into executable programs, making their correctness essential for ensuring code safety and reliability. While extensive research has focused on verifying the correctness of compilers for single-language compilation, the correctness of cross-language compilation — which involves the interaction between two languages and their respective compilers — remains largely unexplored. To fill this research gap, we propose CrossLangFuzzer , a novel framework that introduces a universal intermediate representation (IR) for JVM-based languages and automatically generates cross-language test programs with diverse type parameters and complex inheritance structures. After generating the initial IR, CrossLangFuzzer applies three mutation techniques — LangShuffler, FunctionRemoval , and TypeChanger — to enhance program diversity. By evaluating both the original and mutated programs across multiple compiler versions, CrossLangFuzzer successfully uncovered 10 confirmed bugs in the Kotlin compiler, 4 confirmed bugs in the Groovy compiler, 7 confirmed bugs in the Scala 3 compiler, 2 confirmed bugs in the Scala 2 compiler, and 1 confirmed bug in the Java compiler. Among all mutators, TypeChanger is the most effective, detecting 11 of the 24 compiler bugs. Furthermore, we analyze the symptoms and root causes of cross-compilation bugs, examining the respective responsibilities of language compilers when incorrect behavior occurs during cross-language compilation. To the best of our knowledge, this is the first work specifically focused on identifying and diagnosing compiler bugs in cross-language compilation scenarios. Our research helps to understand these challenges and contributes to improving compiler correctness in multi-language environments. Qiong Feng, Ziyuan Feng, Marat Kh. Akhin, Wei Song 0003, Peng Liang 0001 |
Proc. ACM Program. Lang. | 5 |
| 2025 | DepFuzz: Efficient Smart Contract Fuzzing with Function Dependence GuidanceabstractFuzzing is an effective technique to detect vulnerabilities in smart contracts. The challenge of smart contract fuzzing lies in the statefulness of contracts, which indicates that certain vulnerabilities can only be manifested in specific contract states. State-of-the-art fuzzers may generate and execute a plethora of meaningless or redundant transaction sequences during fuzzing, incurring a penalty in efficiency. To this end, we present DepFuzz , a hybrid fuzzer for efficient smart contract fuzzing, which introduces a symbolic execution module into the feedback-based fuzzer. Guided by the distance-based function dependencies between functions, DepFuzz can efficiently yield meaningful transaction sequences that contribute to vulnerability exposure or code coverage. The experiments on 286 benchmark smart contracts and 500 large real-world smart contracts corroborate that, compared to state-of-the-art approaches, DepFuzz achieves higher instruction coverage rate and uncovers many more vulnerabilities with less time. Wei Song 0003, Jeff Huang 0001 |
Proc. ACM Program. Lang. | 2 |
| 2025 | COTE: Predicting Code-to-Test Co-Evolution by Integrating Link Analysis and Pre-Trained Language Model TechniquesabstractTests, as an essential artifact, should co-evolve with the production code to ensure that the associated production code satisfies specification. However, developers often postpone or even forget to update tests, making the tests outdated and lag behind the code. To predict which tests need to be updated when production code is changed, it is challenging to identify all related tests and determine their change probabilities due to complex change scenarios. This paper fills the gap and proposes a hybrid approach named COTE to predict code-to-test co-evolution. We first compute the linked test candidates based on different code-to-test dependencies. After that, we identify common co-change patterns by building a method-level dependence graph. For the remaining ambiguous patterns, we leverage a pre-trained language model which captures the semantic features of code and the change reasons contained in commit messages to judge one test’s likelihood of being updated. Experiments on our datasets consisting of 6,314 samples extracted from 5,000 Java projects show that COTE outperforms state-of-the-art approaches, achieving a precision of 89.0% and a recall of 71.6%. This work can help practitioners reduce test maintenance costs and improve software quality. Yuyong Liu, Zhifei Chen, Lin Chen 0015, Yanhui Li 0001, Xuansong Li, Wei Song 0003 |
IEEE Trans. Software Eng. | 6 |
| 2024 | FunRedisp: Reordering Function Dispatch in Smart Contract to Reduce Invocation Gas FeesabstractSmart contracts mostly written in Solidity are Turing-complete programs executed on the blockchain platforms such as Ethereum. To prevent resource abuse, a gas fee is required when users deploy or invoke smart contracts. Although saving gas consumption has received much attention, no work investigates the effect of function dispatch on the invocation gas consumption. In this paper, after demystifying how the function dispatch affects the invocation gas consumption, we present FunRedisp, a bytecode refactoring method and an open-source tool, to reduce the overall invocation gas consumption of smart contracts. At the source code level, FunRedisp initially identifies hot functions in a smart contract that have a big chance to be invoked, and then move them to the front of the function dispatch at the bytecode level. We implement FunRedisp and evaluate it on 50 real-world smart contracts randomly selected from Ethereum. The experimental results demonstrate that FunRedisp can save approximately 125.17 units of gas per transaction with the compilation overhead increased by only 0.37 seconds. Wei Song 0003 |
ISSTA | 2 |
| 2024 | FunRedisp: A Function Redispatch Tool to Reduce Invocation Gas Fees in Solidity Smart ContractsabstractFunRedisp is a function dispatch refactoring tool to reduce the overall invocation gas consumption of Solidity smart contracts. It initially extracts all external functions in a contract at the source code level. After identifying the functions which are highly likely to be invoked among these functions by a pre-trained TextCNN model, FunRedisp moves them to the front of the function dispatch at the bytecode level. FunRedisp saves about 125.17 units of gas per contract transaction on the 50 real-world smart contracts, which are randomly selected from Ethereum. The time overhead for refactoring each contract is only 0.37 seconds on average, demonstrating its effectiveness and efficiency. Wei Song 0003 |
ISSTA | 2 |
| 2024 | Depends-Kotlin: A Cross-Language Kotlin Dependency ExtractorabstractSince Google introduced Kotlin as an official programming language for developing Android apps in 2017, Kotlin has gained widespread adoption in Android development. However, compared to Java, there is limited support for Kotlin code dependency analysis, which is the foundation to software analysis. To bridge this gap, we develop Depends-Kotlin to extract entities and their dependencies in Kotlin source code. Not only does Depends-Kotlin support extracting entities' dependencies in Kotlin code, but it can also extract dependency relations between Kotlin and Java. Using three open-source Kotlin-Java mixing projects as our subjects, Depends-Kotlin demonstrates high accuracy and performance in resolving Kotlin-Kotlin and Kotlin-Java dependencies relations. The source code of Depends-Kotlin and the dataset used have been made available at https://github.com/XYZboom/depends-kotlin. We also provide a screen-cast presenting Depends-Kotlin at https://youtu.be/ZPq8SRhgXzM. Qiong Feng, Huan Ji, Wei Song 0003, Peng Liang 0001 |
ASE | 4 |
| 2024 | VarLifter: Recovering Variables and Types from Bytecode of Solidity Smart ContractsabstractSince funds or tokens in smart contracts are maintained through specific state variables, contract audit, an effective means for security assurance, particularly focuses on these variables and their related operations. However, the absence of publicly accessible source code for numerous contracts, with only bytecode exposed, hinders audit efforts. Recovering variables and their types from Solidity bytecode is thus a critical task in smart contract analysis and audit, yet this is a challenging task because the bytecode loses variable and type information, only with low-level data operated by stack manipulations and untyped memory/storage accesses. The state-of-the-art smart contract decompilers miss identifying many variables and incorrectly infer the types for many identified variables. To this end, we propose VarLifter , a lifter dedicated to the precise and efficient recovery of typed variables. VarLifter interprets every read or written field of a data region as at least one potential variable, and after discarding falsely identified variables, it progressively refines the variable types based on the variable behaviors in the form of operation sequences. We evaluate VarLifter on 34,832 real-world Solidity smart contracts. VarLifter attains a precision of 97.48% and a recall of 91.84% for typed variable recovery. Moreover, VarLifter finishes analyzing 77% of smart contracts in around 10 seconds per contract. If VarLifter is used to replace the variable recovery modules of the two state-of-the-art Solidity bytecode decompilers, 52.4%, and 74.6% more typed variables will be correctly recovered, respectively. The applications of VarLifter to contract decompilation, contract audit, and contract bytecode fuzzing illustrate that the recovered variable information improves many contract analysis tasks. Yichuan Li 0005, Wei Song 0003, Jeff Huang 0001 |
Proc. ACM Program. Lang. | 2 |
| 2024 | Risky Dynamic Typing-related Practices in Python: An Empirical StudyabstractPython’s dynamic typing nature provides developers with powerful programming abstractions. However, many type-related bugs are accumulated in code bases of Python due to the misuse of dynamic typing. The goal of this article is to aid in the understanding of developers’ high-risk practices toward dynamic typing and the early detection of type-related bugs. We first formulate the rules of six types of risky dynamic typing-related practices (type smells for short) in Python. We then develop a rule-based tool named RUPOR, which builds an accurate type base to detect type smells. Our evaluation shows that RUPOR outperforms the existing type smell detection techniques (including the Large Language Models–based approaches, Mypy, and PYDYPE) on a benchmark of 900 Python methods. Based on RUPOR, we conduct an empirical study on 25 real-world projects. We find that type smells are significantly related to the occurrence of post-release faults. The fault-proneness prediction model built with type smell features slightly outperforms the model built without them. We also summarize the common patterns, including inserting type check to fix type smell bugs. These findings provide valuable insights for preventing and fixing type-related bugs in the programs written in dynamic-typed languages. Zhifei Chen, Lin Chen 0015, Yibiao Yang, Qiong Feng, Xuansong Li, Wei Song 0003 |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2023 | Discovering Structural Errors From Business Process Event Logs (Extended Abstract)abstractWhile process mining has gained much attention in the past decade, surprisingly, discovering structural errors (i.e., deadlock and lack of synchronization) from event logs has seldom been studied. Since event logs may involve erroneous event occurrences caused by unsynchronized activities, discovering deadlocks and lack of synchronization errors may influence each other. To this end, we first extract from the original event log two independent event logs which are employed to discover deadlocks and lack of synchronization errors, respectively. We then discard the erroneous event occurrences in the two event logs, from which our event relation based mining rules can discover the corresponding structural errors. We have implemented our approach, and the experimental results corroborate that our approach can effectively and efficiently discover process structural errors from the event logs involving sufficient event sequences. Wei Song 0003, Zhen Chang, Hans-Arno Jacobsen, Pengcheng Zhang 0001 |
ICDE | 1 |
| 2023 | PTPDroid: Detecting Violated User Privacy Disclosures to Third-Parties of Android AppsabstractAndroid apps frequently access personal information to provide customized services. Since such information is sensitive in general, regulators require Android app vendors to publish privacy policies that describe what information is collected and why it is collected. Existing work mainly focuses on the types of the collected data but seldom considers the entities that collect user privacy, which could falsely classify problematic declarations about user privacy collected by third-parties into clear disclosures. To address this problem, we propose PTPDroid, a flow-to-policy consistency checking approach and an automated tool, to comprehensively uncover from the privacy policy the violated disclosures to third-parties. Our experiments on real-world apps demonstrate the effectiveness and superiority of PTPDroid, and our empirical study on 1,000 popular real-world apps reveals that violated user privacy disclosures to third-parties are prevalent in practice. Zeya Tan, Wei Song 0003 |
ICSE | 2 |
| 2023 | DDLDroid: Efficiently Detecting Data Loss Issues in Android AppsabstractData loss issues in Android apps triggered by activity restart or app relaunch significantly reduce the user experience and undermine the app quality. While data loss detection has received much attention, the state-of-the-art techniques still miss many data loss issues due to the inaccuracy of the static analysis or the low coverage of the dynamic exploration. To this end, we present DDLDroid, a static analysis approach and an open-source tool, to systematically and efficiently detect data loss issues based on the data flow analysis. DDLDroid is bootstrapped by a saving-restoring bipartite graph which correlates variables that need saving to the corresponding variables that need restoring according to their carrier widgets. The missed or broken saving or restoring data flows lead to data loss issues. The experimental evaluation on 66 Android apps demonstrates the effectiveness and efficiency of our approach: DDLDroid successfully detects 302 true data loss issues in 73 minutes, 180 of which are previously unknown. Wei Song 0003 |
ISSTA | 2 |
| 2023 | DDLDroid: A Static Analyzer for Automatically Detecting Data Loss Issues in Android ApplicationsabstractDDLDroid is a static analyzer for detecting data loss issues in Android apps during activity restart or app relaunch. It is bootstrapped by a saving-restoring bipartite graph which correlates variables that need saving to those that need restoring according to their carrier widgets, and is based on the analysis of saving and restoring data flows. It reports data loss issues once missed or broken data flows are identified. DDLDroid detects 302 true data loss issues out of 66 Android apps in 73 minutes, including 180 previously-unknown issues, demonstrating its effectiveness and efficiency. Wei Song 0003 |
ISSTA | 2 |
| 2023 | TransRacer: Function Dependence-Guided Transaction Race Detection for Smart ContractsabstractSmart contracts are programs that define rules for transactions running on blockchains. Since any qualified transaction sequence within the same block can be orchestrated by the blockchain miner, unexpected results may occur due to data races between transactions (called transaction races). Surprisingly, transaction races in smart contracts have not been fully investigated. To address this, we propose TransRacer, an automated approach and open-source tool that employs symbolic execution to detect transaction races in smart contracts. TransRacer analyzes function dependencies to identify transaction races hidden in specific contract states. It also generates witness transactions that can trigger such races. The experimental results on 50 real-world smart contracts show the effectiveness and efficiency of TransRacer: it detects 426 races in 255.9 minutes, including 149 race bugs leading to inconsistent states. Wei Song 0003, Jeff Huang 0001 |
ESEC/SIGSOFT FSE | 2 |
| 2022 | PermDroid: automatically testing permission-related behaviour of Android applicationsabstractThe Android runtime permission model allows users to grant and revoke permissions at runtime. To verify the robustness of apps, developers have to test the apps repeatedly under a wide range of permission combinations, which is time-consuming and unsuited for regression testing. Existing app testing techniques are of limited help in this context, as they seldom consider different permission combinations explicitly. To address this issue, we present PermDroid to automatically test the permission-related behaviour of apps with permissions granted/revoked dynamically. PermDroid first statically constructs a state transition graph (STG) for the app; it then utilizes the STG for the permission-directed exploration to test permission-related behaviour only under the combinations of the relevant permissions. The experimental results on 50 real-world Android apps demonstrate the effectiveness and efficiency of PermDroid: the average permission-related API invocation coverage achieves 72.38% in 10 minutes, and seven permission-related bugs are uncovered, six of which are not detected by the competitors. Shuaihao Yang, Zigang Zeng, Wei Song 0003 |
ISSTA | 3 |
| 2022 | Collaboration in software ecosystems: A study of work groups in open environment
Zhifei Chen, Wanwangying Ma, Lin Chen 0015, Wei Song 0003 |
Inf. Softw. Technol. | 4 |
| 2022 | Discovering Structural Errors From Business Process Event LogsabstractProcess mining aims at discovering behavioral knowledge of business processes from their event logs, which has received an increasing attention in the era of cloud computing and big data. Surprisingly, to date, discovering structural errors (e.g., deadlocks and lack of synchronization) from event logs has not been considered in state-of-the-art process mining techniques. Moreover, existing process discovery approaches cannot be directly applied to event logs of processes with structural errors due to erroneous event occurrences caused by unsynchronized activities. To address this problem, we first preprocess the event log to obtain two separate event logs that are used to discover deadlocks and lack of synchronization, respectively. Erroneous event occurrences caused by unsynchronized activities are discarded in the two processed event logs, from which our error mining algorithms can discover all process fragments involving structural errors, without the need to obtain the overall process first. We implement our approach in a ProM plugin and evaluate it on event logs of real-life business processes, the results of which demonstrate that our approach can effectively and efficiently discover deadlocks and lack of synchronization if event logs contain sufficient event sequences. Wei Song 0003, Zhen Chang, Hans-Arno Jacobsen, Pengcheng Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Development of Collaborative Business Processes: A Correctness Enforcement ApproachabstractCollaborative business processes gather a set of business processes with complementary competencies and knowledge to cooperate to achieve more business successes. To ensure their successful implementation, correctness is a key issue that needs to be addressed during their development. To this end, a novel correctness enforcement approach is proposed to support the development of collaborative business processes. In this approach, we first give an algorithm to check the correctness of an original process specified by Petri nets. Then, we prune its reachability graph to obtain its core in case of partially correct, which is a reduced reachability graph that doesn't cover invalid states. Finally, we generate a set of controllers from the core using coordination mapping (i.e., inserting some coordination activities into controllers), and then an enforced process is built by the composition of the original process and the controllers. Our approach is implemented as an analysis module called cetool in the PIPE (Platform Independent Petri Net Editor) and it is validated on a set of real-world cases. The results show the effectiveness and efficiency of the proposed approach. Wei Song 0003, Fei Dai 0002, Leilei Lin, Tong Li 0004 |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | Identifying a Minimum Sequence of High-Level Changes Between WorkflowsabstractAdaptive workflow management systems allow workflows to be changed in both the modeling and runtime stages, resulting in many workflow variants. Identifying a minimum sequence of high-level changes between two workflows represents a fundamental yet critical issue. The state-of-the-art approach utilizes digital logic to seek the optimal solution; however, this approach may face difficulties when advanced workflow patterns (e.g., loops) are involved, and it does not scale well. To address this problem, we first propose a naive approach that applies all valid changes to one workflow until the other workflow is found. Then, the approach is optimized from two aspects. First, we present advanced heuristics that significantly reduce the search space without pruning the optimal solution. Second, we employ the A$^\ast$search algorithm to direct the search procedure. Because the heuristic function used in the A$^\ast$algorithm is problem specific, we devise a consistent heuristic function to approximate the edit distance between two workflows, thereby accelerating the search. We implement our approach in a prototype tool and conduct extensive experiments on two data sets to evaluate its effectiveness and efficiency. The experimental results demonstrate that our approach outperforms the state of the art in terms of both application scope and scalability. Wei Song 0003, Fangfei Chen, Hans-Arno Jacobsen, Chengzhen Zhang |
IEEE Trans. Serv. Comput. | 1 |
| 2022 | M-BSRM: Multivariate BayeSian Runtime QoS Monitoring Using Point Mutual InformationabstractQuality of Service (QoS) is well acknowledged as a decisive means for ascertaining the performance of third-party Web services. QoS has high uncertainty in complex and dynamic network environments. QoS monitoring is considered as one of the most effective techniques to detect QoS violations at runtime. However, existing QoS monitoring approaches only consider single QoS attribute and do not provide a promising solution for comprehensively monitoring multivariate QoS attributes. To overcome this problem, a novel QoS monitoring approach, named M-BSRM (MultivariateBayeSianRuntimeMonitoring), is proposed. First, M-BSRM adopts the point mutual information theory to initialize the weights of different environmental impact factors and solves the problem of uneven distribution between classes brought by traditional algorithms. Second, each single QoS attribute is integrated with user preference using the information fusion theory. Finally, a Bayesian classifier is used to comprehensively evaluate multivariate QoS attributes at runtime. The experimental results on both the real-world and simulated data sets show that M-BSRM is more effective, practical, and efficient than the other approaches. Pengcheng Zhang 0001, Huiying Jin, Hai Dong 0001, Wei Song 0003 |
IEEE Trans. Serv. Comput. | 4 |
| 2022 | Privacy-Preserving QoS Forecasting in Mobile Edge EnvironmentsabstractMobile Edge Computing is an emerging technology offering low latency responses by deploying edge servers near mobile devices. We propose a novel privacy-preserving QoS forecasting approach – Edge-Laplace QoS (QoS forecasting with Laplace noise in mobile Edge environments) to address the challenges of user mobility and information leakage encountered by QoS forecasting in mobile edge environments. Edge-Laplace QoS is able to accurately and efficiently forecast Quality of Service (QoS) of various Web Services, while effectively protecting user privacy in mobile edge environments. We employ an improved differential privacy method to add dynamic disguises to the original QoS data in the edge environment to protect user data privacy. A collaborative filtering method is adopted to retrieve similar users’ accessing records based on geographic locations of their accessed servers for QoS forecasting. We conduct a set of experiments using several public network data sets. The results show that the efficiency of Edge-Laplace QoS is superior to traditional forecasting approaches. Edge-Laplace QoS is also validated to be more suitable for edge environments than traditional privacy-preserving approaches. Pengcheng Zhang 0001, Huiying Jin, Hai Dong 0001, Wei Song 0003, Athman Bouguettaya |
IEEE Trans. Serv. Comput. | 4 |
| 2022 | Automated Use-After-Free Detection and Exploit Mitigation: How Far Have We Gone?abstractC/C++ programs frequently encounter memory errors, such as Use-After-Free (UAF), buffer overflow, and integer overflow. Among these memory errors, UAF vulnerabilities are increasingly being exploited by attackers to disrupt critical software systems, leading to serious consequences, such as remote code execution and data breaches. Researchers have proposed dozens of approaches to detect UAFs in testing environments and to mitigate UAF exploit in production environments. However, to the best of our knowledge, no comprehensive studies have evaluated and compared these approaches. In this paper, we shed light on the current UAF detection and exploit mitigation approaches and provide a systematic overview, comprehensive comparison, and evaluation. Specifically, we evaluate the effectiveness and efficiency of publicly available UAF detection and exploit mitigation tools. The experimental results show that static UAF detectors are suitable for detecting intra-procedural UAFs but are not sufficient to detect inter-procedural UAFs in real-world programs. Dynamic UAF detectors are still the first choice for detecting inter-procedural UAFs. Our evaluation also demonstrates that the runtime overhead of existing UAF exploit mitigation tools is relatively stable whereas the memory overhead may vary dramatically with respect to different programs. Finally, we envision potential valuable future research directions. Binfa Gui, Wei Song 0003, Hailong Xiong, Jeff Huang 0001 |
IEEE Trans. Software Eng. | 2 |
| 2021 | IMGDroid: Detecting Image Loading Defects in Android ApplicationsabstractImages are essential for many Android applications or apps. Although images play a critical role in app functionalities and user experience, inefficient or improper image loading and displaying operations may severely impact the app performance and quality. Additionally, since these image loading defects may not be manifested by immediate failures, e.g., app crashes, existing GUI testing approaches cannot detect them effectively. In this paper, we identify five anti-patterns of such image loading defects, including image passing by intent, image decoding without resizing, local image loading without permission, repeated decoding without caching, and image decoding in UI thread. Based on these anti-patterns, we propose a static analysis technique, IMGDroid, to automatically and effectively detect such defects. We have applied IMGDroid to a benchmark of 21 open-source Android apps, and found that it not only successfully detects the 45 previously-known image loading defects but also finds 15 new such defects. Our empirical study on 1,000 commercial Android apps demonstrates that the image loading defects are prevalent. Wei Song 0003, Mengqi Han, Jeff Huang 0001 |
ICSE | 1 |
| 2021 | UAFSan: an object-identifier-based dynamic approach for detecting use-after-free vulnerabilitiesabstractUse-After-Free (UAF) vulnerabilities constitute severe threats to software security. In contrast to other memory errors, UAFs are more difficult to detect through manual or static analysis due to pointer aliases and complicated relationships between pointers and objects. Existing evidence-based dynamic detection approaches only track either pointers or objects to record the availability of objects, which become invalid when the memory that stored the freed object is reallocated. To this end, we propose an approach UAFSan dedicated to comprehensively detecting UAFs at runtime. Specifically, we assign a unique identifier to each newly-allocated object and its pointers; when a pointer dereferences a memory object, we determine whether a UAF occurs by checking the consistency of their identifiers. We implement UAFSan in an open-source tool and evaluate it on a large collection of popular benchmarks and real-world programs. The experiment results demonstrate that UAFSan successfully detect all UAFs with reasonable overhead, whereas existing publicly-available dynamic detectors all miss certain UAFs. Binfa Gui, Wei Song 0003, Jeff Huang 0001 |
ISSTA | 2 |
| 2021 | An Empirical Study on Data Flow Bugs in Business ProcessesabstractAn increasing number of service-based business processes are being developed with the booming of BPaaS (Business Process as a Service) in cloud computing. The profits and performance of enterprises strongly depend on the soundness of their processes being bereft of control flow and data flow bugs. Although some work has focused on the detection of control flow bugs, few studies have comprehensively and empirically investigated data flow bugs in business processes. To this end, we report an empirical study on data flow bugs in business (BPEL) processes. Our analysis of 178 real-world BPEL processes reveals that data flow bugs are surprisingly common: 94 BPEL processes involve data flow bugs, among which redundant output is predominant. The distribution and common scenarios of data flow bugs provide a reference for BPEL process designers. We also investigate the correlation between process complexity metrics and data flow bugs. Based on the statistics of the process complexity metrics and data flow bugs in our empirical study, we present a method to select appropriate metrics as features of BPEL processes and utilize state-of-the-art supervised learning algorithms to predict data flow bugs in an unseen BPEL process. The prediction accuracies of the different classification algorithms exceed 90 percent on average when using our selected metrics. Wei Song 0003, Chengzhen Zhang, Hans-Arno Jacobsen |
IEEE Trans. Cloud Comput. | 1 |
| 2021 | Self-Healing Event LogsabstractEvent logs of process-aware information systems play an increasingly critical role in today's enterprises because they are the basis for a number of business intelligence applications such as complex event processing, provenance analysis, performance analysis, and process mining. However, due to incorrect manual recording, system errors, and resource constraints, event logs inevitably contain noise in the form of deviating event sequences with redundant, missing, or dislocated events. To repair event logs, existing approaches rely on predefined process models to obtain a minimum recovery for each deviating event sequence. However, process models are typically unavailable in practice, rendering existing approaches inapplicable. In this scenario, can event logs be self-healing? To address this problem, we propose an approach that leverages compliant event sequences to repair deviating sequences. Our approach is effective if the compliant event sequences contain sufficient knowledge for repair. We implement our approach in a prototype and employ the tool to conduct experiments. The experimental results demonstrate that our approach can achieve efficient repairs without the help of process models. Wei Song 0003, Hans-Arno Jacobsen, Pengcheng Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | Workflow Refactoring for Maximizing Concurrency and Block-StructurednessabstractIn the era of Internet and big data, contemporary workflows become increasingly large in scale and complex in structure, introducing greater challenges for workflow modeling. Workflows are not with maximized concurrency and block-structuredness in terms of control flow, though languages supporting block-structuredness (e.g., BPEL) are employed. Existing workflow refactoring approaches mostly focus on maximizing concurrency according to dependences between activities, but do not consider the block-structuredness of the refactored workflow. It is easier to comprehend and analyze a workflow that is block-structured and to transform it into BPEL-like processes. In this paper, we aim at maximizing both concurrency and block-structuredness. Nevertheless, not all workflows can be refactored with a block-structured representation, and it is intractable to make sure that the refactored workflows are as block-structured as possible. We first define a well-formed dependence pattern of activities. The control flow among the activities in this pattern can be represented in block-structured forms with maximized concurrency. Then, we propose a greedy heuristics-based graph reduction approach to recursively find such patterns. In this way, the resulting workflow is with maximized concurrency and its block-structuredness approximates optimality. We show the effectiveness and efficiency of our approach with real-world scientific workflows. Wei Song 0003, Hans-Arno Jacobsen, Shing-Chi Cheung, Xiaoxing Ma |
IEEE Trans. Serv. Comput. | 1 |
| 2021 | Dependence-Based Data-Aware Process Conformance CheckingabstractData-aware executable processes are an effective and efficient means to build service-oriented applications. However, since the services involved are loosely-coupled and self-managed, the process is flexible by nature and it executions may deviate from their specifications. In contrast to existing approaches that focus on control flow deviations, we leverage activity dependences for data-aware process conformance checking. To analyze the conformance of a process instance to its process definition, we seek a process reference trace “best-fitting” the instance trace such that the conformance degree of the input trace to the process equals the consistency degree of both traces. We measure the consistency between two traces based on their activity dependences. Since finding the reference trace is NP-hard, we resort to heuristics based on process decomposition and trace replaying to determine the trace. Our approach can identify conformance decrease caused by activity dependence deviations, thus, complementing existing approaches. We implement our approach as a ProM plugin. Experimental results on 102 real-world WS-BPEL processes and 26,880 synthetic input traces confirm the effectiveness and efficiency of our approach. Wei Song 0003, Hans-Arno Jacobsen, Chengzhen Zhang, Xiaoxing Ma |
IEEE Trans. Serv. Comput. | 1 |
| 2021 | LA-LMRBF: Online and Long-Term Web Service QoS ForecastingabstractWe propose aLong-term Quality of Service (QoS) forecasting approach usingAdvertisement andLevenberg-Marquardt improvedRadialBasisFunction (LA-LMRBF)—a novel online QoS forecasting approach. LA-LMRBF aims to accurately predict QoS attributes of Web services in the form of multivariate time series via three stages. First, the phase space reconstruction theory is employed to restore multi-dimensional and nonlinear relations among the multivariate QoS attributes. Second, short-term QoS advertisement data is incorporated to enable long-term QoS forecasting. Finally, an optimized Radial Basis Function (RBF) neural network is constructed to forecast long-term multivariate QoS values, where the Affinity Propagation clustering algorithm is used to determine the number of hidden nodes and the Levenberg-Marquardt (LM) algorithm is utilized to dynamically update some parameters of the RBF neural network. A series of experiments are performed on a mixture of public and self-collected data sets. The results show that LA-LMRBF is superior to the other approaches and more suitable for long-term QoS forecasting. Pengcheng Zhang 0001, Huiying Jin, Hai Dong 0001, Wei Song 0003 |
IEEE Trans. Serv. Comput. | 4 |
| 2020 | Towards Programming and Verification for Activity-Oriented Smart Home SystemsabstractSmart home systems are becoming increasingly popular. Software engineering of such systems hence becomes a prominent challenge. In this engineering paradigm, users are often interested in considering sensor states while they are performing various activities. Existing works have proposed initial efforts on incremental development method for activity-oriented requirements. However, there is no systematic way of ensuring reliability and security of such systems which may be developed by various developers and may execute in a complex environment. Some properties, especially those including timing constraints, need to be satisfied. In this paper, we introduce Actom, a framework for identification of activity-oriented requirements and runtime verification. Actom supports the development of the mapping between activities and required sensor readings (activity-sensor mapping). At runtime, Actom receives results of activity recognition and is able to trigger actuators to provide the required physical conditions for the activities, as determined by the activity-sensor mapping. Moreover, Actom continuously monitors whether activity-sensor mapping holds over a time period during the activity. We also discuss the evaluation plan to demonstrate the effectiveness and efficiency of Actom. The end product will be a systematic framework to facilitate the development of activity-oriented requirements and monitor properties with timing constraints to improve reliability and security. Xuansong Li, Wei Song 0003, Xiangyu Zhang 0001 |
ASE | 2 |
| 2020 | Efficient Service Entity Chain Placement in Mobile Edge ComputingabstractEdge service entity placement is a fundamental issue in mobile edge computing, which tries to place service entities on edge servers to achieve better economic benefits and quality of service for users. Most existing studies towards this issue usually deploy application services separately; however, we observe that many application services can be broken down into smaller service components/entities, which may enable us to share these smaller entities between application services. Therefore, in this paper, we propose the concept of service entity chain, which is a chain of ordered service entities that represent an application service. We study the problem of placing service entities in the form of chains on edge servers within a given cost budget, so as to minimize the total latency experienced by users. We provide a formal problem formulation and design an efficient algorithm for it. Extensive simulations are conducted to demonstrate the advantages of the proposed algorithm compared with two state-of-the-art algorithms. Yu Liang 0001, Jidong Ge, Sheng Zhang 0001, Changan Niu, Wei Song 0003, Bin Luo 0003 |
MSN | 5 |
| 2020 | Short-Term Rainfall Forecasting Using Multi-Layer PerceptronabstractRainfall forecasting is crucial in the field of meteorology and hydrology. However, existing solutions always achieve low prediction accuracy for short-term rainfall forecasting. Atmospheric forecasting models perform worse in many conditions. Machine learning approaches neglect the influences of physical factors in upstream or downstream regions, which make forecasting accuracy fluctuate in different areas. To improve the overall forecasting accuracy for short-term rainfall, this paper proposes a novel solution called Dynamic Regional Combined short-term rainfall Forecasting approach (DRCF) using Multi-layer Perceptron (MLP). First, Principal Component Analysis (PCA) is used to reduce the dimension of thirteen physical factors, which serves as the input of MLP. Second, a greedy algorithm is applied to determine the structure of MLP. The surrounding sites are perceived based on the forecasting site. Finally, to solve the clutter interference which is caused by the extension of the perception range, DRCF is enhanced with several dynamic strategies. Experiments are conducted on data from 56 real-world meteorology sites in China, and we compare DRCF with atmospheric models and other machine learning approaches. The experimental results show that DRCF outperforms existing approaches in both threat score (TS) and root mean square error (RMSE). Pengcheng Zhang 0001, Yangyang Jia, Jerry Zeyu Gao, Wei Song 0003, Hareton K. N. Leung |
IEEE Trans. Big Data | 4 |
| 2020 | Understanding JavaScript Vulnerabilities in Large Real-World Android ApplicationsabstractJavaScript-related vulnerabilities are becoming a major security threat to hybrid mobile applications. In this article, we present a systematic study to understand how JavaScript is used in real-world Android apps and how it may lead to security vulnerabilities. We begin by conducting an empirical study on the top-100 most popular Android apps to investigate JavaScript usage and its related security vulnerabilities. Our study identifies four categories of JavaScript usage and finds that three of these categories, if inappropriately used, can respectively lead to three types of vulnerabilities. We also design and implement an automatic tool named JSDroid to detect JavaScript-related vulnerabilities. We have applied JSDroid to 1,000 large real-world Android apps and found that over 70 percent of these apps have potential JavaScript-related vulnerabilities and 20 percent of them can be successfully exploited. Moreover, based on the vulnerabilities identified by JSDroid, we have successfully launched real attacks on 30 real-world apps. Wei Song 0003, Jeff Huang 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2020 | Scientific Workflow Protocol Discovery from Public Event Logs in CloudsabstractWith the advancement of cloud computing, many challenging scientific problems can be solved using scientific workflow technology which integrates geo-distributed instruments, applications, and big data effectively and efficiently. For workflow collaboration, the workflow protocols of all participants are needed. However, workflow protocols are not always available and are often outdated as the workflow evolve frequently. To address this problem, we propose a novel workflow discovery approach which can extract up-to-date scientific workflow protocols from public event logs in clouds, without the need to access the full-fledged event logs involving private events. Our approach leverages transitive precedence relations between events to achieve this. We implement our approach as a ProM plug-in, and evaluate it through extensive experiments on event logs of real-world scientific workflows. The experimental results demonstrate that our approach requires a weaker completeness notion of event logs than the state-of-the-art do, and our approach derives the same workflow protocol from the public event log as that discovered from the original event log, and thus the private events can be protected. Wei Song 0003, Hans-Arno Jacobsen, Fangfei Chen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2019 | ServDroid: detecting service usage inefficiencies in Android applicationsabstractServices in Android applications are frequently-used components for performing time-consuming operations in the background. While services play a crucial role in the app performance, our study shows that service uses in practice are not as efficient as expected, e.g., they tend to cause unnecessary resource occupation and/or energy consumption. Moreover, as service usage inefficiencies do not manifest with immediate failures, e.g., app crashes, existing testing-based approaches fall short in finding them. In this paper, we identify four anti-patterns of such service usage inefficiency bugs, including premature create, late destroy, premature destroy, and service leak, and present a static analysis technique, ServDroid, to automatically and effectively detect them based on the anti-patterns. We have applied ServDroid to a large collection of popular real-world Android apps. Our results show that, surprisingly, service usage inefficiencies are prevalent and can severely impact the app performance. Wei Song 0003, Jeff Huang 0001 |
ESEC/SIGSOFT FSE | 1 |
| 2019 | From event streams to process models and back: Challenges and opportunities
Pnina Soffer, Annika Hinze, Agnes Koschmider, Holger Ziekow, Claudio Di Ciccio, Boris Koldehofe, Oliver Kopp, Hans-Arno Jacobsen, Jan Sürmeli, Wei Song 0003 |
Inf. Syst. | 10 |
| 2019 | Measuring Business Process Consistency Across Different Abstraction LevelsabstractBusiness process modeling can take place at three abstraction levels, namely the conceptual level for system requirements, the logical level for system specification, and the physical level for software development. The consistency between these process models is crucial in process mapping, process integration, and difference detection. Existing work either only provides a simple “yes” or “no” answer as the consistency result, or simply checks the consistency from the control flow perspective. This paper presents a systematic approach to the quantitative measurement of the consistency between business processes across different abstraction levels. We use the essential event constraints to quantify the consistency from the perspectives of control flow and data flow, where the importance of different essential event constraints can also be distinguished. Our approach is implemented in a prototype tool. We evaluate our approach using synthetic datasets, the results of which demonstrate the effectiveness and efficiency of our approach. Wei Song 0003, Jiacun Wang 0001, Jianchun Xing, Qizhen Zhou |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2019 | Instance Migration Validity for Dynamic Evolution of Data-Aware ProcessesabstractLikely more than many other software artifacts, business processes constantly evolve to adapt to ever changing application requirements. To enable dynamic process evolution, where changes are applied to in-flight processes, running process instances have to be migrated. On the one hand, as many instances as possible should be migrated to the changed process. On the other hand, the validity to migrate an instance should be guaranteed to avoid introducing dynamic change bugs after migration. As our theoretical results show, when the state of variables is taken into account, migration validity of data-aware process instances is undecidable. Based on the trace of an instance, existing approaches leverage trace replaying to check migration validity. However, they err on the side of caution, not identifying many instances as potentially safe to migrate. We present a more relaxed migration validity checking approach based on the dependence graph of a trace. We evaluate effectiveness and efficiency of our approach experimentally showing that it allows for more instances to safely migrate than for existing approaches and that it scales in the number of instances checked. Wei Song 0003, Xiaoxing Ma, Hans-Arno Jacobsen |
IEEE Trans. Software Eng. | 1 |
| 2018 | Response Time Aware Operator Placement for Complex Event Processing in Edge Computing
Xinchen Cai, Hongyu Kuang, Hao Hu 0001, Wei Song 0003, Jian Lu 0001 |
ICSOC | 4 |
| 2018 | Weighted Bayesian Runtime Monitor: A Novel QoS Monitoring Approach Sensitive to Environmental FactorsabstractHow to assure Quality of Service (QoS) of the third-party services is very important for the SOA. Effective monitoring technique towards QoS, which is an important measurement for third-party service quality, is necessary to ensure quality of Web service. Current monitoring approaches do not consider the influences of environment factors such as the position of server, user usage, and the load at runtime. Ignoring these influences, which do exist among the monitoring process, may cause existing monitoring approaches producing unpredictable monitoring results. In order to overcome this limitation, this paper proposes a novel Web Service QoS (WS-Qos) monitoring approach sensitive to environmental factors called weighted Bayesian Runtime Monitor (wBSRM) based on weighted naïve Bayesian classifiers and Term Frequency-Inverse Document Frequency (TF-IDF) algorithm. wBSRM constructs weighted naïve Bayesian classifier by learning a part of samples to classify the monitoring results. The results meeting QoS standard are classified as [Formula: see text] and the one that does not meet is classified as [Formula: see text]. Classifier can also output ratio between posterior probability of [Formula: see text] and [Formula: see text], and consequently the analysis can lead to three monitoring results including [Formula: see text], [Formula: see text] or inconclusive. A set of dedicated experiments are conducted to validate wBSRM. The experiments are based on a public dataset and a simulated dataset under the given standard. The experimental results demonstrate that wBSRM is better than previous approaches. Pengcheng Zhang 0001, Huiying Jin, Hareton K. N. Leung, Wei Song 0003, Yu Zhou 0010 |
Int. J. Softw. Eng. Knowl. Eng. | 5 |
| 2018 | IgS-wBSRM: A time-aware Web Service QoS monitoring approach in dynamic environments
Pengcheng Zhang 0001, Huiying Jin, Zhipeng He 0004, Hareton K. N. Leung, Wei Song 0003 |
Inf. Softw. Technol. | 5 |
| 2018 | AocML: A Domain-Specific Language for Model-Driven Development of Activity-Oriented Context-Aware Applications
Xuansong Li, XianPing Tao, Wei Song 0003, Kai Dong 0001 |
J. Comput. Sci. Technol. | 3 |
| 2018 | Cost and Energy Aware Scheduling Algorithm for Scientific Workflows with Deadline Constraint in CloudsabstractCloud computing is a suitable platform to execute the deadline-constrained scientific workflows which are typical big data applications and often require many hours to finish. Moreover, the problem of energy consumption has become one of the major concerns in clouds. In this paper, we present a cost and energy aware scheduling (CEAS) algorithm for cloud scheduler to minimize the execution cost of workflow and reduce the energy consumption while meeting the deadline constraint. The CEAS algorithm consists of five sub-algorithms. First, we use the VM selection algorithm which applies the concept of cost utility to map tasks to their optimal virtual machine (VM) types by the sub-makespan constraint. Then, two tasks merging methods are employed to reduce execution cost and energy consumption of workflow. Further, In order to reuse the idle VM instances which have been leased, the VM reuse policy is also proposed. Finally, the scheme of slack time reclamation is utilized to save energy of leased VM instances. According to the time complexity analysis, we conclude that the time complexity of each sub-algorithm is polynomial. The CEAS algorithm is evaluated using Cloudsim and four real-world scientific workflow applications, which demonstrates that it outperforms the related well-known approaches. Zhongjin Li, Jidong Ge, Wei Song 0003, Hao Hu 0001, Bin Luo 0003 |
IEEE Trans. Serv. Comput. | 4 |
| 2018 | Static and Dynamic Process ChangeabstractApproaches for modifying processes both at build time and at run time are commonly referred to as process change, which play an increasingly important role in the enterprise today, where more than ever before, changing requirements must be rapidly accommodated. Over the years, approaches for supporting process change have received much attention from the research community. In spite of that, no comprehensive survey of this important subject exits. To draw a clear picture that analyzes the status of research in this area, in this paper, we conduct a systematic literature review on process change. The resulting survey sheds light on how to classify approaches for process change, determines what the principal research questions and challenges are, and identifies several research directions for further study. Wei Song 0003, Hans-Arno Jacobsen |
IEEE Trans. Serv. Comput. | 1 |
| 2017 | Parallelized Mobility-Aware Complex Event ProcessingabstractThe concept of complex event processing (CEP) and complex-event-aware service have been extensively studied to retrieve relevant information from massive amount of realtime streaming events. In mobile environment, the Mobilityaware CEP (MCEP) system was proposed to address the issue of synchronization problem between different query ranges and MCEP operators. We noticed that MCEP systems lack the ability to process event in parallel and scale out when system load is high. In this paper, we proposed a parallel architecture for MCEP. The architecture can handle the synchronization problem and guarantee the correctness of event processing result. We also proposed a scaling strategy that can automatically scale out operators while ensures semantic transparency. An empirical evaluation based on up to 10 ViMs demonstrated that our approach is able to achieve higher throughput while keeping the MCEP synchronization mechanism valid. Yuhao Gong, Hongyu Kuang, Xinchen Cai, Hao Hu 0001, Wei Song 0003, Jian Lu 0001 |
ICWS | 5 |
| 2017 | A Web Service QoS Forecasting Approach Based on Multivariate Time SeriesabstractIn order to accurately forecast Quality of Service (QoS) of different Web Services, this paper proposes a novel QoS forecasting approach called MulA-LMRBF (Multi-step fore-casting with Advertisement and Levenberg-Marquardt improved Radial Basis Function) based on multivariate time series. Considering the correlation among different QoS attributes, we use phase-space reconstruction to map historical multivariate QoS data into a dynamic system, use Average Dimension (AD) to estimate the embedding dimension and delay time of reconstructed phase space. We also add the short-term QoS advertisement data of service provider to form a more comprehensive data set. Then, RBF (Radial Basis Function) neural network improved by the Levenberg-Marquardt (LM) algorithm is used to update the weight of the neural network dynamically, which improves the forecasting accuracy and realizes the dynamic multiple-step forecasting. The experimental results demonstrate that MulA-LMRBF is better than previous approaches in term of precision and is more suitable for multi-step forecasting. Pengcheng Zhang 0001, Wenrui Li 0002, Hareton K. N. Leung, Wei Song 0003 |
ICWS | 5 |
| 2017 | EHBDroid: beyond GUI testing for Android applicationsabstractWith the prevalence of Android-based mobile devices, automated testing for Android apps has received increasing attention. However, owing to the large variety of events that Android supports, test input generation is a challenging task. In this paper, we present a novel approach and an open source tool called EHBDroid for testing Android apps. In contrast to conventional GUI testing approaches, a key novelty of EHBDroid is that it does not generate events from the GUI, but directly invokes callbacks of event handlers. By doing so, EHBDroid can efficiently simulate a large number of events that are difficult to generate by traditional UI-based approaches. We have evaluated EHBDroid on a collection of 35 real-world large-scale Android apps and compared its performance with two state-of-the-art UI-based approaches, Monkey and Dynodroid. Our experimental results show that EHBDroid is significantly more effective and efficient than Monkey and Dynodroid: in a much shorter time, EHBDroid achieves as much as 22.3% higher statement coverage (11.1% on average) than the other two approaches, and found 12 bugs in these benchmarks, including 5 new bugs that the other two failed to find. Wei Song 0003, Xiangxing Qian, Jeff Huang 0001 |
ASE | 1 |
| 2017 | Scientific Workflow Mining in CloudsabstractComputing clouds have become the platform of choice for the deployment and execution of scientific workflows. Due to the uncertainty and unpredictability of scientific exploration, the execution plan for a scientific workflow may vary from the definition. It is therefore of great significance to be able to discover actual workflows from execution histories (event logs) to reproduce experimental results and to establish provenance. However, most existing process mining techniques focus on discovering control flow-oriented business processes in a centralized environment, and thus, they are mostly inapplicable to the discovery of data flow-oriented, unstructured scientific workflows in distributed cloud environments. In this paper, we present Scientific Workflow Mining as a Service (SWMaaS) to support both intra-cloud and inter-cloud scientific workflow mining. The approach is implemented as a ProM plug-in and is evaluated on event logs derived from real-world scientific workflows. Through experimental results, we demonstrate the effectiveness and efficiency of our approach. Wei Song 0003, Fangfei Chen, Hans-Arno Jacobsen, Xiaoxu Xia, Chunyang Ye, Xiaoxing Ma |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | Efficient Alignment Between Event Logs and Process ModelsabstractThe aligning of event logs with process models is of great significance for process mining to enable conformance checking, process enhancement, performance analysis, and trace repairing. Since process models are increasingly complex and event logs may deviate from process models by exhibiting redundant, missing, and dislocated events, it is challenging to determine the optimal alignment for each event sequence in the log, as this problem is NP-hard. Existing approaches utilize the cost-based A* algorithm to address this problem. However, scalability is often not considered, which is especially important when dealing with industrial-sized problems. In this paper, by taking advantage of the structural and behavioral features of process models, we present an efficient approach which leverages effective heuristics and trace replaying to significantly reduce the overall search space for seeking the optimal alignment. We employ real-world business processes and their traces to evaluate the proposed approach. Experimental results demonstrate that our approach works well in most cases, and that it outperforms the state-of-the-art approach by up to 5 orders of magnitude in runtime efficiency. Wei Song 0003, Xiaoxu Xia, Hans-Arno Jacobsen, Pengcheng Zhang 0001, Hao Hu 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2016 | Effa: a proM plugin for recovering event logsabstractWhile event logs generated by business processes play an increasingly significant role in business analysis, the quality of data remains a serious problem. Automatic recovery of dirty event logs is desirable and thus receives more attention. However, existing methods only focus on missing event recovery, or fall short of efficiency. To this end, we present Effa, a ProM plugin, to automatically recover event logs in the light of process specifications. Based on advanced heuristics including process decomposition and trace replaying to search the minimum recovery, Effa achieves a balance between repairing accuracy and efficiency. Xiaoxu Xia, Wei Song 0003, Fangfei Chen, Xuansong Li, Pengcheng Zhang 0001 |
Internetware | 2 |
| 2016 | Generating effective test cases based on satisfiability modulo theory solvers for service-oriented workflow applicationsabstractWeb Service Business Process Execution Language (WS-BPEL) is one of the most popular service-oriented workflow applications. The unique features (e.g. dead path elimination semantics and correlation mechanism) of WS-BPEL applications have raised enormous problems to its test case generation, especially in unit testing. Existing studies mainly assume that each path in the control flow graphs that correspond to WS-BPEL applications is feasible, which always yields imprecise test cases or complicates testing results. The current study tackles this problem based on satisfiability modulo theory solvers. First, a new coverage criterion is proposed to measure the quality of test sets for testing WS-BPEL applications. Second, decomposition algorithms are presented to obtain test paths that meet the proposed coverage criterion. Finally, this paper symbolically encodes each test path with several constraints by capturing the unique features of WS-BPEL. These constraints are solved and the test cases (test paths and test data) are obtained with the help of satisfiability modulo theory solvers to test WS-BPEL applications effectively. Experiments are conducted using our approach and other typical approaches (e.g. message-sequence generation-based approach and concurrent path analysis approach) with 10 WS-BPEL applications. Experimental results demonstrate that the test cases generated by our approach can avoid instantiating idle instance and expose more faults. Copyright © 2015 John Wiley & Sons, Ltd. Hongda Wang, Jianchun Xing, Qiliang Yang, Wei Song 0003 |
Softw. Test. Verification Reliab. | 4 |
| 2016 | Process Discovery from Dependence-Complete Event LogsabstractProcess mining, especially process discovery, has been utilized to extract process models from event logs. One challenge faced by process discovery is to identify concurrency effectively. State-of-the-art approaches employ activity orders in traces to undertake process discovery and they require stringent completeness notions of event logs. Thus, they may fail to extract appropriate processes when event logs cannot meet the completeness criteria. To address this problem, we propose in this paper a novel technique which leverages activity dependences in traces. Based on the observation that activities with no dependencies can be executed in parallel, our technique is in a position to discover processes with concurrencies even if the logs fail to meet the completeness criteria. That is, our technique calls for a weaker notion of completeness. We evaluate our technique through experiments on both real-world and synthetic event logs, and the conformance checking results demonstrate the effectiveness of our technique and its relative advantages compared with state-of-the-art approaches. Wei Song 0003, Hans-Arno Jacobsen, Chunyang Ye, Xiaoxing Ma |
IEEE Trans. Serv. Comput. | 1 |
| 2015 | Heuristic Recovery of Missing Events in Process LogsabstractEvent logs are of paramount significance for process mining and complex event processing. Yet, the quality of event logs remains a serious problem. Missing events of logs are usually caused by omitting manual recording, system failures, and hybrid storage of executions of different processes. It has been proved that the problem of minimum recovery based on a priori process specification is NP-hard. State-of-the-art approach is still lacking in efficiency because of the large search space. To address this issue, in this paper, we leverage the technique of process decomposition and present heuristics to efficiently prune the unqualified sub-processes that fail to generate the minimum recovery. We employ real-world processes and their incomplete sequences to evaluate our heuristic approach. The experimental results demonstrate that our approach achieves high accuracy as the state-of-the-art approach does, but it is more efficient. Wei Song 0003, Xiaoxu Xia, Hans-Arno Jacobsen, Pengcheng Zhang 0001, Hao Hu 0001 |
ICWS | 1 |
| 2015 | A Novel QoS Monitoring Approach Sensitive to Environmental FactorsabstractThe quality of service-oriented system relies heavily on the third-party service. Such reliance would result in many uncertainties, in consideration of the complex and changeable network environment. Hence, effective runtime monitoring technique is required by service-oriented system. Several monitoring approaches have been proposed. However, all of these approaches do not consider the influences of environmental factors such as the position of server and users, and the load at runtime. Ignoring these influences, which exist among monitoring process, may cause wrong monitoring results. In order to solve this problem, this paper proposes a novel QoS monitoring approach sensitive to environmental factors called wBSRM (weighted Bayesian Runtime Monitoring) based on weighted naive Bayesian and TF-IDF (Term Frequency-Inverse Document Frequency). The proposed approach measures influence of environmental factor by TF-IDF algorithm and then constructs weighted naïve Bayesian classifier by learning part of samples to classify monitoring results. Experiments are conducted based on both public network data set and randomly generated data set. The experimental results demonstrate that our approach is better than previous approaches. Pengcheng Zhang 0001, Hareton K. N. Leung, Wei Song 0003, Yu Zhou 0010 |
ICWS | 4 |
| 2015 | Qos-aware Automatic Web Service Composition Considering QoS CorrelationsabstractWeb service composition is the process of automatically arranging multiple services into workflow so as to supply complex user needs. With the rapid increase in the number of Web services, it's beyond the human ability to generate the composition result manually, which further indicates the importance of automatic service composition. Besides, it's essential that not only functional needs but also non-functional requirements need to be satisfied in the composition process. The QoS-aware automatic service composition has received considerable attention and made a lot of progress, but it's rare to consider the QoS correlations which are essential in actual application. Thus, in this paper, we take QoS correlations between services into account and propose a novel approach to address the QoS-aware automatic service composition problem. Evaluations show that, compared to the state of the art, our method can address QoS correlations between services and generate the service composition with optimal QoS values more efficiently. Hao Hu 0001, Wei Song 0003, Jidong Ge |
Internetware | 3 |
| 2014 | FuAET: a tool for developing fuzzy self-adaptive software systemsabstractHandling uncertainty in software self-adaptation has become an important and challenging issue. In our previous work, we proposed a fuzzy control based approach named Software Fuzzy Self-Adaptation (SFSA) to address fuzziness, a kind of uncertainty in software self-adaptation. However, our SFSA approach still lacks a tool to efficiently support the implementation process of SFSA. Existing tools for realizing self-adaptive applications does not directly deal with fuzziness in self-adaption loops. In this paper, we present the FuAET, a tool designed for building fuzzy self-adaptive software systems. The novelty of the tool is that it can not only provide a friendly GUI for editing and testing fuzzy self-adaptation strategies in intelligible domain specific language (DSL), but also automatically convert DSL-based fuzzy selfadaptation strategies into aspect-based programming code in native-language (NL), e.g., C++. This paper describes the design framework and implementation principles of FuAET, and then proposes a general development process using FuAET for reference to developers. Finally, we conduct an empirical study for evaluation of FuAET using an industrial control application. The results show that FuAET can automate the development of SFSA and ease the burden of the software engineers. Qiliang Yang, XianPing Tao, Hongwei Xie, Jianchun Xing, Wei Song 0003 |
Internetware | 5 |
| 2013 | COCO: consistency analysis of process-driven internetware applicationsabstractProcesses are an effective and efficient way to construct on-demand Internetware applications. It is often needed to evaluate whether two process-driven applications are consistent, or whether the implemented process conforms to the process specification. Most existing methods only return qualitative results (i.e., true or false), so slight inconsistencies may lead to a false result. To address this problem, based on activity constraints, we have presented a quantitative approach to process consistency analysis. In this paper, we focus on the implementation issues of our approach and introduce how to use our tool in practice. Wei Song 0003, Xiaoxing Ma, Qiliang Yang |
Internetware | 2 |
| 2013 | Behavioral Consistency Measurement and Analysis of WS-BPEL Processes
Wei Song 0003, Jianchun Xing, Qiliang Yang, Hongda Wang |
WAIM | 2 |
| 2013 | Fuzzy Self-Adaptation of Mission-Critical Software Under Uncertainty
Qiliang Yang, Jian Lu 0001, XianPing Tao, Xiaoxing Ma, Jianchun Xing, Wei Song 0003 |
J. Comput. Sci. Technol. | 6 |
| 2012 | Towards Dynamic Evolution of Service ChoreographiesabstractTo stay on the cutting edge, Web services ought to adapt themselves to the evolving business requirements and the changing environments. For a long-running service choreography, its member services may need to evolve even at run-time. However, inconsistencies or spurious results (e.g., unspecified receptions and deadlocks) may occur if these services evolve dynamically in an uncoordinated manner. To cope with this problem, we propose an approach that supports the dynamic evolution of choreographies. In our approach, two mechanisms are proposed to make sure that the choreography evolution can be conducted in an orderly fashion. First, an evolution protocol is proposed to support dynamic co-evolution of the member services in a choreography. Second, the proposed approach restricts choreography changes to one single service only if the relevant partner services can evolve simultaneously. A typical purchase order application is used to motivate our proposal and illustrate the viability of our approach. Wei Song 0003, Gongxuan Zhang, Yang Zou 0001, Qiliang Yang, Xiaoxing Ma |
APSCC | 1 |
| 2012 | A priority-based transaction commit protocol for composite web servicesabstractService composition provides an effective way to conduct cross-organizational business transactions. Some protocols and frameworks have been proposed to ensure ACID properties of transactional service requisitions from the perspective of a service requester. However, few of them have focused on how to optimize the services' profits from the perspective of a service provider. In this paper, we present a priority-based transaction commit protocol for composite Web services. For this protocol, Nash Bargaining Solution (NBS) is used to differentiate service requesters so that services (i.e., resources) can be allocated with different priorities. In this way, the profit of a service provider can be maximized with proportional fairness. For service requesters, the proposed protocol can guarantee the atomicity of service requisitions. Our experimental results reveal that the proposed protocol can significantly enhance the profit of service providers without violating atomicity of Web service transactions. Wei Song 0003, Xiaoxing Ma |
Internetware | 1 |
| 2011 | Global-Time-Offsets Based Checkpoint Selection for Dynamic Verification of Fixed-Time Constraints in Grid WorkflowsabstractIn a grid workflows, fixed-time constraints are often set at some activities to guarantee that the corresponding activities and finally the overall workflow can be timely completed. At the run-time stage, activity completion durations can vary which may lead to the violations of some fixed-time constraints. Therefore, the workflow needs to be monitored at the run-time so that some temporal violations can be identified timely and consequently some exception handling actions can be taken in time. However, existing temporal verification strategies conduct some unnecessary computations and comparisons which impact the efficiency of overall temporal verification. To address this problem, we present a global-time-offsets based temporal verification approach. With our approach, the time cost for dynamic verification of the fixed-time constraints will be reduced. The experimental result further demonstrates that our approach is more efficient than others. Hongda Wang, Wei Song 0003, Jianchun Xing, Qiliang Yang |
APSCC | 2 |
| 2011 | Refactoring and Publishing WS-BPEL Processes to Obtain More PartnersabstractWS-BPEL processes can facilitate service discovery when the services have multiple interfaces in certain order. Current approaches derive the abstract WS-BPEL processes directly from the corresponding executable ones by hiding or omitting the internal activities. However, these simple approaches may prevent the services from being found by valuable potential partners at service discovery stage. To address this problem, we propose a novel approach to refactoring the executable and abstract WS-BPEL processes for service discovery. We show the application of our approach through a typical travel agency service. Wei Song 0003, Xiaoxing Ma, Shing-Chi Cheung, Hao Hu 0001, Qiliang Yang, Jian Lu 0001 |
ICWS | 1 |
| 2010 | Toward a fuzzy control-based approach to design of self-adaptive softwareabstractSelf-adaptive software is expected to adjust itself attributes or structures at runtime in response to changes. Aiming at addressing some challenging problems such as difficult mathematically modeling software using the current control theoretical methods, we propose a novel fuzzy-control-based approach to achieve self-adaptive software, which is presented as framework of fuzzy self-adaptive software (FFSAS). In this framework, the general model, the implementation architecture, and the design methodology are put forward and discussed in detail. The fuzzy-control-based approach is evaluated with a news-website case study. Qiliang Yang, Jian Lu 0001, Juelong Li, Xiaoxing Ma, Wei Song 0003, Yang Zou 0001 |
Internetware | 5 |
| 2008 | Toward a Model-Based Approach to Dynamic Adaptation of Composite ServicesabstractFacing changing environments and evolving business rules, composite services ought to be adaptable, even at run-time. Existing mainstream service composition languages and execution engines exhibit insufficient support for variability and adaptability to cater for dynamic changes. Research efforts have been put on the extension of the languages and argumentation of the engines. However, how to ensure the correctness for the adaptation of a running composite service instance and minimize unnecessary re-execution of component services remains a challenge. To address this problem, we propose a model-based approach that allows run-time adaptation of composite services. It is based on an instance transfer mechanism that transfers an active instance of the old service composition schema to a appropriate state of the new schema. Algorithms are proposed to find the appropriate destination states of the transformation. After the migration, the suspended instances can resume their execution according to the new schema. An example based on a FindRoute composite service is also included. Wei Song 0003, Xiaoxing Ma, Wan-Chun Dou, Jian Lu 0001 |
ICWS | 1 |