Zhiliang Zhu 0001

dblp:55/1401-1 · DBLP profile ↗
← Back
65ranked-venue papers
4as first author
31since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 24 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 8 · 6 since 2021Computer networks · 8 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 1 since 2021Systems, architecture and hardware · 4Databases, data management, data science and information retrieval · 4 · 1 first-author
YearPublicationVenuePosition
2026 Enhanced neighborhood metric for spreadsheet fault prediction
Ying Wang 0038, Hai Yu 0001, Zhiliang Zhu 0001
Autom. Softw. Eng.4
2026 Multi-Image compressive encryption via a dynamic delayed feedback chaotic system and synchronized fractal diffusion
Zhiliang Zhu 0001, Wei Zhang 0150, Meng Xing, Hai Yu 0001
Expert Syst. Appl.2
2026 Selecting test cases to reveal defects in recurrent neural networks
Xiaohu Song, Hai Yu 0001, Zhiliang Zhu 0001
Neurocomputing5
2026 High-Sensitivity Chaotic System for Image Security: Applications to Multi-ROI Encryption
abstract
To ensure the security and privacy of images, this paper proposes an Improved 1D Sine–Logistic (I1DSL) chaotic system and its application in multiple Regions of Interest (ROI) image encryption. As the security foundation, the I1DSL chaotic system is employed to generate highly secure chaotic key streams. Performance analysis demonstrates that it exhibits superior dynamic behavior and a broader parameter space. During encryption, YOLOv10 is used to extract ROIs and generate keys accurately. For each ROI, a selective scrambling is applied to achieve a uniform and unpredictable pixel distribution. In image diffusion, a Multi-directional Dynamic Josephus (MDJ) matrix is introduced to realize synchronous scrambling and diffusion of multiple ROIs. Finally, a global random scrambling process is applied to ROIs of different sizes to further enhance the overall randomness and security of the encryption. Experimental results demonstrate that the proposed multi-ROI image encryption algorithm exhibits excellent performance in both security and efficiency, further validating the security of the I1DSL system.
Zhiliang Zhu 0001, Wei Zhang 0150, Xinhe Zhao, Hai Yu 0001
IEEE Internet Things J.2
2026 Dual-encoder semantics and hierarchical identity refinement for personalized image generation
Yingli Hou, Zhiliang Zhu 0001, Wei Zhang 0150, Hai Yu 0001
Knowl. Based Syst.2
2025 Diplomatist: What Do Cross-language Dependencies Reflect Software Ecosystem Health?
abstract
In large-scale software development, multilingual projects, those involving multiple interacting programming languages, have become increasingly common in both industry and the open-source community. Research indicates that cross-language dependencies in these projects can increase the like-lihood of risks, such as functionality defects and security vulnerabilities. While most existing studies focus on cross-language dependencies between host languages and specific guest languages (e.g., C/C++), interactions between host languages and a broader range of guest languages, as well as the broader impact of such dependencies on software ecosystems, remain underexplored.To address the above limitations, in this paper, we develop a technique, Diplomatist, to identify and analyze cross-language dependencies between host languages, such as Java, and guest languages, including JavaScript, Python, Ruby, PHP, and C/C++. Diplomatist automatically analyzes cross-language invocation APIs and constructs a large-scale knowledge repository to standardize code features for identifying library versions across various guest languages, enabling host languages to trace the guest language libraries they invoke. Evaluation shows that Diplomatist achieved an average precision of 88.9% and a recall of 91.5% on a high-quality benchmark, indicating its high accuracy in detecting cross-language dependencies. Using Diplomatist, we identified 435,258 Java libraries that indirectly or transitively depend on libraries from other ecosystems. Diplomatist provides a list of cross-language pivotal libraries that contribute to preserving the long-term health and sustainability of software ecosystems. Moreover, we conduct a case study to examine the impact of the risks introduced due to cross-language dependencies on programming language ecosystems, by analyzing a full-picture of the cross-language dependency graph. Our findings show that fragile projects or libraries can propagate security issues across ecosystems via these dependencies, impacting 13,739 downstream projects in the Maven ecosystem. We utilized Diplomatist to provide remediation suggestions to relevant project developers. Issue reports of some subjects have been confirmed by developers.
Fanyi Meng 0003, Ying Wang 0038, Chun Yong Chong, Hai Yu 0001, Zhiliang Zhu 0001
ASE5
2025 MCTS-Refined CoT: High-Quality Fine-Tuning Data for LLM-Based Repository Issue Resolution
abstract
LLMs demonstrate strong performance in automated software engineering, particularly for code generation and issue resolution. While proprietary models like GPT-4o achieve high benchmarks scores on SWE-bench, their API dependence, cost, and privacy concerns limit adoption. Open-source alternatives offer transparency but underperform in complex tasks, especially sub-100B parameter models. Although quality Chain-of-Thought (CoT) data can enhance reasoning, current methods face two critical flaws: (1) weak rejection sampling reduces data quality, and (2) inadequate step validation causes error accumulation. These limitations lead to flawed reasoning chains that impair LLMs’ ability to learn reliable issue resolution.The paper proposes MCTS-REFINE, an enhanced Monte Carlo Tree Search (MCTS)-based algorithm that dynamically validates and optimizes intermediate reasoning steps through a rigorous rejection sampling strategy, generating high-quality CoT data to improve LLM performance in issue resolution tasks. Key innovations include: (1) augmenting MCTS with a reflection mechanism that corrects errors via rejection sampling and refinement, (2) decomposing issue resolution into three subtasks—File Localization, Fault Localization, and Patch Generation—each with clear ground-truth criteria, and (3) enforcing a strict sampling protocol where intermediate outputs must exactly match verified developer patches, ensuring correctness across reasoning paths.Experiments on SWE-bench Lite and SWE-bench Verified demonstrate that LLMs fine-tuned with our CoT dataset achieve substantial improvements over baselines. Notably, Qwen2.5-72B-Instruct achieves 28.3%(Lite) and 35.0%(Verified) resolution rates, surpassing SOTA baseline SWE-Fixer-Qwen-72B with the same parameter scale, which only reached 24.7%(Lite) and 32.8%(Verified). Given precise issue locations as input, our fine-tuned Qwen2.5-72B-Instruct model achieves an impressive issue resolution rate of 43.8%(Verified), comparable to the performance of Deepseek-v3. We open-source our MCTS-REFINE framework, CoT dataset, and fine-tuned models to advance research in AI-driven software engineering.
Yibo Wang 0008, Zhihao Peng 0009, Ying Wang 0038, Zhao Wei, Hai Yu 0001, Zhiliang Zhu 0001
ASE6
2025 Demystifying Cross-Language C/C++ Binaries: A Robust Software Component Analysis Approach
abstract
Binary Software Composition Analysis (BSCA) is a technique for identifying the versions of third-party libraries (TPLs) used in compiled binaries, thereby tracing the dependencies and vulnerabilities of software components without access to their source code. However, existing BSCA techniques struggle with cross-language invoked C/C++ binaries in polyglot projects due to two key challenges: (1) interference from heterogeneous Foreign Function Interface (FFI) bindings that obscure distinctive TPL features and generate false positives during matching processes, and (2) the inherent complexity of composite binaries (fused binaries), particularly prevalent in polyglot development where multiple TPLs are frequently compiled into single executable units, resulting in blurred boundaries between libraries and substantially compromising version identification precision.We propose DeeperBin, a BSCA technique that addresses these challenges through a high-quality, large-scale feature database with four key advantages: (1) high scalability that is capable of analyzing 74,647 C/C++ TPL versions, (2) efficient noise filtering to remove FFI bindings and common functions, (3) automated extraction of version string regexes for 31,855 TPL versions, and (4) generation of distinctive version features using the Minimum Description Length (MDL) principle. Evaluated on 418 cross-language binaries, DeeperBin achieves 81.2% precision and 84.6% recall for TPL detection, outperforming state-of-the-art (SOTA) techniques by 14.1% and 23.2%, respectively. For version identification, it achieves 70.3% precision, a 12.6% improvement over state-of-the-art techniques. Ablation studies confirm the usefulness of FFI filtering and MDL-based features, boosting precision and recall by 17.1% and 18.8%. DeeperBin also maintains competitive efficiency, processing binaries in 364.3 seconds while supporting the largest feature database.
Meiqiu Xu, Ying Wang 0038, Xian Zhan, Shing-Chi Cheung, Hai Yu 0001, Zhiliang Zhu 0001
ASE7
2025 GILNet: Grouping interaction learning network for lightweight salient object detection
Yiru Wei, Zhiliang Zhu 0001, Hai Yu 0001, Wei Zhang 0150
Appl. Intell.2
2025 A cross dual branch guidance network for salient object detection
abstract
The effective integration of multi-level contextual information is crucial for deep learning-based salient object detection. However, most existing approaches either adopt the parallel structure or the progressive structure to predict salient objects, which still face challenges in consistently and accurately detecting salient objects of varying scales. In this paper, we propose a novel cross dual branch guidance network to effectively extract the rich semantic features and gradually enhance the saliency map scale-by-scale. Concretely, the parallel branch is guided by the progressive branch to obtain coarse location information of salient objects. In turn, the progressive branch is able to obtain uniform semantics and rich details to enhance saliency map with the guidance of the parallel branch. To obtain the dynamic receptive field, a dynamic sampling module (DSM) is introduced, which can dynamically adjust the sampling positions such that the spatial details of salient objects in complex scenes can be well recognized. In addition, we design a global context module (GCM) to explore the correlation between different parts of salient object or different salient objects, which is favorable for improving the completeness of saliency map. Experiments on five released benchmark datasets demonstrate the effectiveness and superiority of our proposed approach against other state-of-the-art methods.
Yiru Wei, Zhiliang Zhu 0001, Hai Yu 0001, Wei Zhang 0150
Eng. Appl. Artif. Intell.2
2025 Debloating Software Through Enhanced Static Analysis and Constraint Rules
abstract
ABSTRACT Introduction Java applications often bloat, consuming more resources than necessary. Existing bytecode debloating techniques have several limitations, such as compromised correctness caused by incomplete code collection, which leads to false positives. The debloating process is resource‐intensive because it relies significantly on the test coverage to identify unnecessary code. This method is particularly problematic for large‐enterprise applications, where executing a single test case can take hours or days. In addition, most available debloating tools are standalone utilities that require manual configuration to be integrated with the project‐building process. Methods In this study, we introduce an automated approach known as Trimming. Trimming achieves three main improvements: (1) it collects the necessary code using enhanced static analysis, improves the modeling of reflective calls, and infers instantiation objects. It also uses constraint rules to identify and process reference codes that contribute to bloating but cannot be removed without causing errors. (2) It is not based on test coverage. (3) It allows for direct debloating from the project build by implementing a Maven plugin that is configured within the POM.xml file of the bloated project. Results Our evaluation results show that when Java reflection analysis or constraint rules in TRIMMING were disabled, the effectiveness of debloating analysis dropped to 88.25% with a 95% confidence interval (CI) of [85.3%, 91.2%], and 35.0% with a 95% CI of [30.5%, 39.5%], respectively. Moreover, inferring instantiation objects reduced the program size by 4.9%. Trimming achieved a bytecode reduction rate of 24.1%, with all the debloated projects successfully compiling and passing test suites.
Xiaohu Song, Hai Yu 0001, Ying Wang 0038, Zhiliang Zhu 0001
Softw. Pract. Exp.4
2025 CLIP-GAN: Stacking CLIPs and GAN for Efficient and Controllable Text-to-Image Synthesis
abstract
Recent advances in text-to-image synthesis have captivated audiences worldwide, drawing considerable attention. Although significant progress in generating photo-realistic images through large pre-trained autoregressive and diffusion models, these models face three critical constraints: (1) The requirement for extensive training data and numerous model parameters; (2) Inefficient, multi-step image generation process; and (3) Difficulties in controlling the output visual features, requiring complexly designed prompts to ensure text-image alignment. Addressing these challenges, we introduce the CLIP-GAN model, which innovatively integrates the pretrained CLIP model into both the generator and discriminator of the GAN. Our architecture includes a CLIP-based generator that employs visual concepts derived from CLIP through text prompts in a feature adapter module. We also propose a CLIP-based discriminator, utilizing CLIP's advanced scene understanding capabilities for more precise image quality evaluation. Additionally, our generator applies visual concepts from CLIP via the Text-based Generator Block (TG-Block) and the Polarized Feature Fusion Module (PFFM) enabling better fusion of text and image semantic information. This integration within the generator and discriminator enhances training efficiency, enabling our model to achieve evaluation results not inferior to large pre-trained autoregressive and diffusion models, but with a 94% reduction in learnable parameters. CLIP-GAN aims to achieve the best efficiency-accuracy trade-off in image generation given the limited resource budget. Extensive evaluations validate the superior performance of the model, demonstrating faster image generation speed and the potential for greater stylistic diversity within the GAN model, while still preserving its smooth latent space.
Yingli Hou, Wei Zhang 0150, Zhiliang Zhu 0001, Hai Yu 0001
IEEE Trans. Multim.3
2025 A Hybrid Data Plane SDN Architecture Based on Cloud-Native for Segment Routing Over IPv6
abstract
SRv6 (Segment Routing over IPv6) serves as the cornerstone for next-generation internet solutions. In SRv6 networks, SDN architectures address critical challenges like multi-cloud SRv6 traffic engineering through deep integration with SRH parsing and Kubernetes CNI. Existing SDN implementations face two key limitations: 1) Inability to concurrently configure hybrid data planes combining VPP and Linux nodes, which diminishes SRv6’s cross-plane advantages; 2) Operational complexity from maintaining synchronization between disparate controllers, hindering unified southbound management of SRv6 nodes. To resolve these issues, we propose a redesigned SDN architecture for cloud-native networks featuring two innovations: 1) A Linux SR Plugin (LSP) southbound interface enabling coordinated SRv6 support across heterogeneous nodes; 2) Enhanced northbound API capabilities through key-value store integration with cloud platforms. The implemented LSP plugin allows transparent management of hybrid data planes via business-focused API operations. Our experimental validation demonstrates that the LSP successfully verifies all SRv6 configurations supported by both the Linux and VPP data planes while achieving ≥400 configurations/second throughput. These results confirm the LSP’s capability to deliver unified southbound control for SRv6 across heterogeneous network environments.
Donglei Yuan, Yuli Zhao, Hai Yu 0001, Zhiliang Zhu 0001
IEEE Trans. Netw. Serv. Manag.5
2024 Efficiently Trimming the Fat: Streamlining Software Dependencies with Java Reflection and Dependency Analysis
abstract
Numerous third-party libraries introduced into client projects are not actually required, resulting in modern software being gradually bloated. Software developers may spend much unnecessary effort to manage the bloated dependencies: keeping the library versions up-to-date, making sure that heterogeneous licenses are compatible, and resolving dependency conflict or vulnerability issues.
Xiaohu Song, Ying Wang 0038, Guangtai Liang, Qianxiang Wang, Zhiliang Zhu 0001
ICSE6
2024 CLUE: Customizing clustering techniques using machine learning for software modularization
abstract
Software clustering is often used as a remodularization and architecture recovery technique to help developers simplify software maintenance tasks and ease the burden of software comprehension. While the choice of clustering technique can significantly influence the outcomes of remodularization, it is noteworthy that existing works have yet to conduct an exhaustive exploration of the suitability of various clustering techniques for different software projects. Although many prior works introduce new clustering techniques, their validations often focus on specific domains, which may restrict the generalizability of their findings. In this paper, we conduct an empirical study aimed at understanding the impact of software features and architectural problems on the effectiveness of various software clustering techniques. Leveraging our empirical findings, we propose an approach, CLUE, which leverages Machine Learning (ML) models to customize a suitable software clustering technique for a given software. Our approach focuses on eight types of software clustering techniques and offers insights into their suitability based on features and architectural problems of software. This comprehensive analysis helps developers to select the suitable clustering technique that can achieve the best MoJoFM, c2ccvg, or TurboMQ value from the chosen pool of software clustering techniques for specific software. We evaluate CLUE by analyzing 100 open-source software projects. The experiment results demonstrate that CLUE achieves highly accurate clustering technique customization, with an accuracy exceeding 90%.
Fanyi Meng 0003, Ying Wang 0038, Chun Yong Chong, Hai Yu 0001, Zhiliang Zhu 0001
Internetware5
2024 Language-vision matching for text-to-image synthesis with context-aware GAN
Yingli Hou, Wei Zhang 0150, Zhiliang Zhu 0001, Hai Yu 0001
Expert Syst. Appl.3
2024 Online fountain code with an improved caching mechanism
abstract
Abstract The original online fountain codes discard a large number of symbols that do not meet the requirements at the decoder. To improve channel utilization, this article proposes a new online fountain code. In the completion phase, the proposed code improves the receiving rules of encoded symbols, that is, the encoded symbols discarded in the original online fountain codes are selectively cached. Moreover, an optimal degree selection strategy of encoded symbols is obtained in the proposed scheme. The valid degree range of the proposed strategy is also analyzed, leading to an upper bound of cached events which eventually limits the number of feedbacks. The theoretical analysis and simulation results reveal that the proposed scheme outperforms two state‐of‐the‐art online fountain codes in terms of overhead factors, number of feedback transmissions, and encoding/decoding efficiency.
Zhen Zhen, Yuli Zhao, Francis C. M. Lau 0002, Bochang Ma, Bin Zhang 0001, Hai Yu 0001, Zhiliang Zhu 0001
IET Commun.8
2024 Automated construction of reference model for software remodularization through software evolution
abstract
Abstract The undocumented evolution of a software project and its underlying architecture underscores the need to recover the architecture from the software's implementation‐level artifacts. Despite the existence of various software remodularization techniques, they often suffer from inaccuracies, and evaluating their effectiveness is challenging due to the absence of accurate “ground‐truth” architectures or reference models. Prior studies on reference model construction are time‐consuming and labor‐intensive as it heavily relies on manual analysis by domain experts. Besides, other existing approaches that directly utilize the directory or package structure of the latest version can be unreliable, lacking in‐depth analysis of the employed software structure. To address the above limitations, in this paper, we proposeAutomatedConstruction ofReferenceModel (ACRM), an approach for automatically constructing reference models by assigning weights to classes for various software projects using the metadata of all software versions and historical maintenance records. We evaluate ACRM through both quantitative and qualitative analyses. The experiment results provide quantitative validation and show that the generated reference models are reasonable, as confirmed by the relationship between proposed reference models and architectural smells or bugs. Furthermore, we conduct a survey among the practitioners from industry, to gain insights from practitioners' practices and further validate the generated reference models. The survey shows that, on average, 87% of the participants agree with the reference models generated by ACRM. Moreover, we propose an improved metric,wc2c, which analyzes the strengths and weaknesses of different types of software clustering techniques using the proposed reference models of the given software. Finally, we discuss the potential benefits of using ACRM in analyzed projects, particularly in terms of improving software quality, reducing maintenance costs, and enhancing developer productivity.
Fanyi Meng 0003, Hai Yu 0001, Chun Yong Chong, Ying Wang 0038, Zhiliang Zhu 0001
J. Softw. Evol. Process.5
2024 Expanding-Window Zigzag Decodable Fountain Codes for Scalable Multimedia Transmission
abstract
In this article, we present a coding method called expanding-window zigzag decodable fountain code with unequal error protection property (EWF-ZD UEP code) to achieve scalable multimedia transmission. The key idea of the EWF-ZD UEP code is to utilize bit-shift operation and expanding-window strategy to improve the decoding performance of the high-priority data without performance deterioration of the low-priority data. To provide more protection for the high-priority data, we precode the different importance level using LDPC codes of varying code rates. The generalized variable nodes of different importance levels are further grouped into several windows. Each window is associated with a selection probability and a bit-shift distribution. The combination of bit-shift and symbol exclusive-or operations is used to generate an encoded symbol. Theoretical and simulation results on input symbols of two importance levels reveal that the proposed EWF-ZD UEP code exhibits UEP property. With a small bit shift, the decoding delay for recovering high-priority input symbols is decreased without degrading the decoding performance of the low-priority input symbols. Moreover, according to the simulation results on scalable video coding, our scheme provides better basic video quality at a lower proportion of received symbols compared to three state-of-art UEP fountain codes.
Yuli Zhao, Francis C. M. Lau 0002, Hai Yu 0001, Zhiliang Zhu 0001, Bin Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Evolution-Aware Constraint Derivation Approach for Software Remodularization
abstract
Existing software clustering techniques tend to ignore prior knowledge from domain experts, leading to results (suggested big-bang remodularization actions) that cannot be acceptable to developers. Incorporating domain experts knowledge or constraints during clustering ensures the obtained modularization aligns with developers’ perspectives, enhancing software quality. However, manual review by knowledgeable domain experts for constraint generation is time-consuming and labor-intensive. In this article, we propose an evolution-aware constraint derivation approach, Escort , which automatically derives clustering constraints based on the evolutionary history from the analyzed software. Specifically, Escort can serve as an alternative approach to derive implicit and explicit constraints in situations where domain experts are absent. In the subsequent constrained clustering process, Escort can be considered as a framework to help supplement and enhance various unconstrained clustering techniques to improve their accuracy and reliability. We evaluate Escort based on both quantitative and qualitative analysis. In quantitative validation, Escort , using generated clustering constraints, outperforms seven classic unconstrained clustering techniques. Qualitatively, a survey with developers from five IT companies indicates that 89% agree with Escort ’s clustering constraints. We also evaluate the utility of refactoring suggestions from our constrained clustering approach, with 54% acknowledged by project developers, either implemented or planned for future releases.
Fanyi Meng 0003, Ying Wang 0038, Chun Yong Chong, Hai Yu 0001, Zhiliang Zhu 0001
ACM Trans. Softw. Eng. Methodol.5
2023 Can Machine Learning Pipelines Be Better Configured?
abstract
A Machine Learning (ML) pipeline configures the workflow of a learning task using the APIs provided by ML libraries. However, a pipeline’s performance can vary significantly across different configurations of ML library versions. Misconfigured pipelines can result in inferior performance, such as poor execution time and memory usage, numeric errors and even crashes. A pipeline is subject to misconfiguration if it exhibits significantly inconsistent performance upon changes in the versions of its configured libraries or the combination of these libraries. We refer to such performance inconsistency as a pipeline configuration (PLC) issue.
Yibo Wang 0008, Ying Wang 0038, Yue Yu 0001, Shing-Chi Cheung, Hai Yu 0001, Zhiliang Zhu 0001
ESEC/SIGSOFT FSE7
2023 Plumber: Boosting the Propagation of Vulnerability Fixes in the npm Ecosystem
abstract
Vulnerabilities are known reported security threats that affect a large amount of packages in thenpmecosystem. To mitigate these security threats, the open-source community strongly suggests vulnerable packages to timely publish vulnerability fixes and recommends affected packages to update their dependencies. However, there are still serious lags in the propagation of vulnerability fixes in the ecosystem. In our preliminary study on the latest versions of 356,283 activenpmpackages, we found that 20.0% of them can still introduce vulnerabilities via direct or transitive dependencies although the involved vulnerable packages have already published fix versions for over a year. Prior study by (Chinthanet et al. 2021) lays the groundwork for research on how to mitigate propagation lags of vulnerability fixes in an ecosystem. They conducted an empirical investigation to identify lags that might occur between the vulnerable package release and its fixing release. They found that factors such as the branch upon which a fix landed and the severity of the vulnerability had a small effect on its propagation trajectory throughout the ecosystem. To ensure quick adoption and propagation of a release that contains the fix, they gave several actionable advice to developers and researchers. However, it is still an open question how to design an effective technique to accelerate the propagation of vulnerability fixes. Motivated by this problem, in this paper, we conducted an empirical study to learn the scale of packages that block the propagation of vulnerability fixes in the ecosystem and investigate their evolution characteristics. Furthermore, we distilled the remediation strategies that have better effects on mitigating the fix propagation lags. Leveraging our empirical findings, we propose an ecosystem-level technique,Plumber, for deriving feasible remediation strategies to boost the propagation of vulnerability fixes. To precisely diagnose the causes of fix propagation blocking,Plumbermodels the vulnerability metadata, andnpmdependency metadata and continuously monitors their evolution. By analyzing a full-picture of the ecosystem-level dependency graph and the corresponding fix propagation statuses, it derives remediation schemes for pivotal packages. In the schemes,Plumberprovides customized remediation suggestions with vulnerability impact analysis to arouse package developers’ awareness. We appliedPlumberto generating 268 remediation reports for the identified pivotal packages, to evaluate its remediation effectiveness based on developers’ feedback. Encouragingly, 47.4% our remediation reports received positive feedback from many well-knownnpmprojects, such asTensorflow/tfjs,Ethers.js, andGoogleChrome/workbox. Our reports have boosted the propagation of vulnerability fixes into 16,403 root packages through 92,469 dependency paths. On average, each remediated package version is receiving 72,678 downloads per week by the time of this work.
Ying Wang 0038, Lin Pei, Yue Yu 0001, Chang Xu 0001, Shing-Chi Cheung, Hai Yu 0001, Zhiliang Zhu 0001
IEEE Trans. Software Eng.8
2023 Runtime Permission Issues in Android Apps: Taxonomy, Practices, and Ways Forward
abstract
Android introduces a new permission model that allows apps to request permissions at runtime rather than at the installation time since 6.0 (Marshmallow, API level 23). While this runtime permission model provides users with greater flexibility in controlling an app's access to sensitive data and system features, it brings new challenges to app development. First, as users may grant or revoke permissions at any time while they are using an app, developers need to ensure that the app properly checks and requests required permissions before invoking any permission-protected APIs. Second, Android's permission mechanism keeps evolving and getting customized by device manufacturers. Developers are expected to comprehensively test their apps on different Android versions and device models to make sure permissions are properly requested in all situations. Unfortunately, these requirements are often impractical for developers. In practice, many Android apps suffer from various runtime permission issues (ARP issues). While existing studies have explored ARP issues, the understanding of such issues is still preliminary. To better characterize ARP issues, we performed an empirical study using 135 Stack Overflow posts that discuss ARP issues and 199 real ARP issues archived in popular open-source Android projects on GitHub. Via analyzing the data, we observed 11 types of ARP issues that commonly occur in Android apps. For each type of issues, we systematically studied: (1) how they can be manifested, (2) how pervasive and serious they are in real-world apps, and (3) how they can be fixed. We also analyzed the evolution trend of different types of issues from 2015 to 2020 to understand their impact on the Android ecosystem. Furthermore, we conducted a field survey and in-depth interviews among the practitioners from open-source community and industry, to gain insights from practitioners’ practices and learn their requirements of tools that can help combat ARP issues. Finally, to understand the strengths and weaknesses of the existing tools that can detect ARP issues, we builtARPBench, an open benchmark consisting of 94 real ARP issues, and evaluated the performance of three available tools. The experimental results indicate that the existing tools have very limited supports for detecting our observed issue types and report a large number of false alarms. We further analyzed the tools’ limitations and summarized the challenges of designing an effective ARP issue detection technique. We hope that our findings can shed light on future research and provide useful guidance to practitioners.
Ying Wang 0038, Yibo Wang 0008, Yepang Liu 0001, Chang Xu 0001, Shing-Chi Cheung, Hai Yu 0001, Zhiliang Zhu 0001
IEEE Trans. Software Eng.8
2022 Insight: Exploring Cross-Ecosystem Vulnerability Impacts
abstract
Vulnerabilities, referred to as CLV issues, are induced by cross-language invocations of vulnerable libraries. Such issues greatly increase the attack surface of Python/Java projects due to their pervasive use of C libraries. Existing Python/Java build tools in PyPI and Maven ecosystems fail to report the dependency on vulnerable libraries written in other languages such as C. CLV issues are easily missed by developers. In this paper, we conduct the first empirical study on the status quo of CLV issues in PyPI and Maven ecosystems. It is found that 82,951 projects in these ecosystems are directly or indirectly dependent on libraries compiled from the C project versions that are identified to be vulnerable in CVE reports. Our study arouses the awareness of CLV issues in popular ecosystems and presents related analysis results.
Meiqiu Xu, Ying Wang 0038, Shing-Chi Cheung, Hai Yu 0001, Zhiliang Zhu 0001
ASE5
2022 Duplicated zigzag decodable fountain codes with the unequal error protection property
Yuli Zhao, Francis C. M. Lau 0002, Zhiliang Zhu 0001, Hai Yu 0001
Comput. Commun.4
2022 Weighted zigzag decodable fountain codes for unequal error protection
abstract
Abstract By combining bit‐shift and exclusive‐or operations, a weighted zigzag decodable fountain code is proposed to achieve an unequal error protection property. In the proposed scheme, the input symbols of different importance levels are first pre‐coded into variable nodes using low‐density parity‐check codes. Then, bit‐shift operations prior to exclusive‐or are performed on the non‐uniformly selected variable nodes to generate encoded symbols. An analysis of the erasure probabilities based on the and‐or tree is further introduced. Simulation results show that by appropriately choosing the maximum bit‐shift amount, the proposed weighted zigzag decodable fountain codes can successfully recover the more important input symbols prior to the less important ones.
Yuli Zhao, Francis C. M. Lau 0002, Bin Zhang 0001, Zhiliang Zhu 0001, Hai Yu 0001
IET Commun.5
2022 Devising optimal integration test orders using cost-benefit analysis
abstract
Integration testing is an integral part of software testing. Prior studies have focused on reducing test cost in integration test order generation. However, there are no studies concerning the testing priorities of critical classes when generating integration test orders. Such priorities greatly affect testing efficiency. In this study, we propose an effective strategy that considers both test cost and efficiency when generating test orders. According to a series of dynamic execution scenarios, the software is mapped into a multi-layer dynamic execution network (MDEN) model. By analyzing the dynamic structural complexity, an evaluation scheme is proposed to quantify the class testing priority with the defined class risk index. Cost—benefit analysis is used to perform cycle-breaking operations, satisfying two principles: assigning higher priorities to higher-risk classes and minimizing the total complexity of test stubs. We also present a strategy to evaluate the effectiveness of integration test order algorithms by calculating the reduction of software risk during their testing process. Experiment results show that our approach performs better across software of different scales, in comparison with the existing algorithms that aim only to minimize test cost. Finally, we implement a tool, ITOsolution, to help practitioners automatically generate test orders.
Fanyi Meng 0003, Ying Wang 0038, Hai Yu 0001, Zhiliang Zhu 0001
Frontiers Inf. Technol. Electron. Eng.4
2022 Will Dependency Conflicts Affect My Program's Semantics?
abstract
Java projects are often built on top of various third-party libraries. If multiple versions of a library exist on the classpath, JVM will only load one version and shadow the others, which we refer to asdependency conflicts. This would give rise tosemantic conflict(SC) issues, if the library APIs referenced by a project have identical method signatures but inconsistent semantics across the loaded and shadowed versions of libraries. SC issues are difficult for developers to diagnose in practice, since understanding them typically requires domain knowledge. Although adapting the existing test generation technique for dependency conflict issues,Riddle, to detect SC issues is feasible, its effectiveness is greatly compromised. This is mainly becauseRiddlerandomly generates test inputs, while the SC issues typically require specific arguments in the tests to be exposed. To address that, we conducted an empirical study of 316 real SC issues to understand the characteristics of such specific arguments in the test cases that can capture the SC issues. Inspired by our empirical findings, we propose an automated testing techniqueSensor, which synthesizes test cases using ingredients from the project under test to trigger inconsistent behaviors of the APIs with the same signatures in conflicting library versions. Our evaluation results show thatSensoris effective and useful: it achieved a$Precision$of 0.898 and a$Recall$of 0.725 on open-source projects and a$Precision$of 0.821 on industrial projects; it detected 306 semantic conflict issues in 50 projects, 70.4 percent of which had been confirmed as real bugs, and 84.2 percent of the confirmed issues have been fixed quickly.
Ying Wang 0038, Rongxin Wu, Ming Wen 0001, Yepang Liu 0001, Shing-Chi Cheung, Hai Yu 0001, Chang Xu 0001, Zhiliang Zhu 0001
IEEE Trans. Software Eng.9
2021 HERO: On the Chaos When PATH Meets Modules
abstract
Ever since its first release in 2009, the Go programming language (Golang) has been well received by software communities. A major reason for its success is the powerful support of library-based development, where a Golang project can be conveniently built on top of other projects by referencing them as libraries. As Golang evolves, it recommends the use of a new library-referencing mode to overcome the limitations of the original one. While these two library modes are incompatible, both are supported by the Golang ecosystem. The heterogeneous use of library-referencing modes across Golang projects has caused numerous dependency management (DM) issues, incurring reference inconsistencies and even build failures. Motivated by the problem, we conducted an empirical study to characterize the DM issues, understand their root causes, and examine their fixing solutions. Based on our findings, we developed Hero, an automated technique to detect DM issues and suggest proper fixing solutions. We applied Hero to 19,000 popular Golang projects. The results showed that Hero achieved a high detection rate of 98.5% on a DM issue benchmark and found 2,422 new DM issues in 2,356 popular Golang projects. We reported 280 issues, among which 181 (64.6%) issues have been confirmed, and 160 of them (88.4%) have been fixed or are under fixing. Almost all the fixes have adopted our fixing suggestions.
Ying Wang 0038, Liang Qiao 0002, Chang Xu 0001, Yepang Liu 0001, Shing-Chi Cheung, Na Meng 0001, Hai Yu 0001, Zhiliang Zhu 0001
ICSE8
2021 An ultrahigh-resolution image encryption algorithm using random super-pixel strategy
Wei Zhang 0150, Weijie Han, Zhiliang Zhu 0001, Hai Yu 0001
Multim. Tools Appl.3
2021 Superpixels With Content-Adaptive Criteria
abstract
Superpixels are widely used in computer vision applications. Most of the existing superpixel methods use established criteria to indiscriminately process all pixels, resulting in superpixel boundary adherence and regularity being unnecessarily inter-inhibitive. This study builds upon a previous work by proposing a new segmentation strategy that classifies image content into meaningful areas containing object boundaries and meaningless parts that include color-homogeneous and texture-rich regions. Based on this classification, we design two distinct criteria to process the pixels in different environments to achieve highly accurate superpixels in content-meaningful areas and keep the regularity of the superpixels in content-meaningless regions. Additionally, we add a group of weights when adopting the color feature, successfully reducing the undersegmentation error. The superior accuracy and the moderate compactness achieved by the proposed method in comparative experiments with several state-of-the-art methods indicate that the content-adaptive criteria efficiently reduce the compromise between boundary adherence and compactness.
Wei Zhang 0150, Hai Yu 0001, Zhiliang Zhu 0001
IEEE Trans. Image Process.4
2020 Watchman: monitoring dependency conflicts for Python library ecosystem
abstract
The PyPI ecosystem has indexed millions of Python libraries to allow developers to automatically download and install dependencies of their projects based on the specified version constraints. Despite the convenience brought by automation, version constraints in Python projects can easily conflict, resulting in build failures. We refer to such conflicts as Dependency Confict (DC) issues. Although DC issues are common in Python projects, developers lack tool support to gain a comprehensive knowledge for diagnosing the root causes of these issues. In this paper, we conducted an empirical study on 235 real-world DC issues. We studied the manifestation patterns and fixing strategies of these issues and found several key factors that can lead to DC issues and their regressions. Based on our findings, we designed and implemented Watchman, a technique to continuously monitor dependency conflicts for the PyPI ecosystem. In our evaluation, Watchman analyzed PyPI snapshots between 11 Jul 2019 and 16 Aug 2019, and found 117 potential DC issues. We reported these issues to the developers of the corresponding projects. So far, 63 issues have been confirmed, 38 of which have been quickly fixed by applying our suggested patches.
Ying Wang 0038, Ming Wen 0001, Yepang Liu 0001, Yibo Wang 0008, Zhenming Li, Hai Yu 0001, Shing-Chi Cheung, Chang Xu 0001, Zhiliang Zhu 0001
ICSE10
2020 A novel compressive sensing-based framework for image compression-encryption with S-box
Zhiliang Zhu 0001, Wei Zhang 0150, Hai Yu 0001, Yuli Zhao
Multim. Tools Appl.1
2020 Watershed-Based Superpixels With Global and Local Boundary Marching
abstract
Superpixels are widely used in computer vision applications, as they conserve the running costs of subsequent processing while preserving the original performance. In most of the existing algorithms, the boundary adherence and the compactness of superpixels are necessarily inter-inhibitive because the color/gradient information is balanced against the position constraints, and the set criteria define all pixels indiscriminately. In this paper, we present a two-phase superpixel segmentation method based on the watershed transformation. After designing a new approach for calculating the flooding priority, we propose a new strategy with two distinct criteria for global and local refinement of the boundary pixels. These criteria reduce the compromise between the boundary adherence and compactness. Unlike the indiscriminate standards, our method applies different treatments to pixels in different environments, preserving the color homogeneity in content-rich areas while improving the regularity of the superpixels in content-plain regions. The superior accuracy and computing time of our proposed method are verified in comparison experiments with several state-of-the-art methods.
Zhiliang Zhu 0001, Hai Yu 0001, Wei Zhang 0150
IEEE Trans. Image Process.2
2019 Could I have a stack trace to examine the dependency conflict issue?
abstract
Intensive use of libraries in Java projects brings potential risk of dependency conflicts, which occur when a project directly or indirectly depends on multiple versions of the same library or class. When this happens, JVM loads one version and shadows the others. Runtime exceptions can occur when methods in the shadowed versions are referenced. Although project management tools such as Maven are able to give warnings of potential dependency conflicts when a project is built, developers often ask for crashing stack traces before examining these warnings. It motivates us to develop Riddle, an automated approach that generates tests and collects crashing stack traces for projects subject to risk of dependency conflicts. Riddle, built on top of Asm and Evosuite, combines condition mutation, search strategies and condition restoration. We applied Riddle on 19 real-world Java projects with duplicate libraries or classes. We reported 20 identified dependency conflicts including their induced crashing stack traces and the details of generated tests. Among them, 15 conflicts were confirmed by developers as real issues, and 10 were readily fixed. The evaluation results demonstrate the effectiveness and usefulness of Riddle.
Ying Wang 0038, Ming Wen 0001, Rongxin Wu, Zhenwei Liu 0001, Shin Hwei Tan, Zhiliang Zhu 0001, Hai Yu 0001, Shing-Chi Cheung
ICSE6
2019 Efficient protection using chaos for Context-Adaptive Binary Arithmetic Coding in H.264/Advanced Video Coding
Zhiliang Zhu 0001, Wei Zhang 0150, Hai Yu 0001
Multim. Tools Appl.2
2018 Do the dependency conflicts in my project matter?
abstract
Intensive dependencies of a Java project on third-party libraries can easily lead to the presence of multiple library or class versions on its classpath. When this happens, JVM will load one version and shadows the others. Dependency conflict (DC) issues occur when the loaded version fails to cover a required feature (e.g., method) referenced by the project, thus causing runtime exceptions. However, the warnings of duplicate classes or libraries detected by existing build tools such as Maven can be benign since not all instances of duplication will induce runtime exceptions, and hence are often ignored by developers. In this paper, we conducted an empirical study on real-world DC issues collected from large open source projects. We studied the manifestation and fixing patterns of DC issues. Based on our findings, we designed Decca, an automated detection tool that assesses DC issues' severity and filters out the benign ones. Our evaluation results on 30 projects show that Decca achieves a precision of 0.923 and recall of 0.766 in detecting high-severity DC issues. Decca also detected new DC issues in these projects. Subsequently, 20 DC bug reports were filed, and 11 of them were confirmed by developers. Issues in 6 reports were fixed with our suggested patches.
Ying Wang 0038, Ming Wen 0001, Zhenwei Liu 0001, Rongxin Wu, Hai Yu 0001, Zhiliang Zhu 0001, Shing-Chi Cheung
ESEC/SIGSOFT FSE8
2018 SSCSMA-based random relay selection scheme for large-scale relay networks
Zhiliang Zhu 0001, Francis C. M. Lau 0002, Yuli Zhao, Hai Yu 0001
Comput. Commun.2
2018 Improved online fountain codes
abstract
Online fountain codes have been proven to require lower overhead and fewer feedbacks than growth codes for successful decoding. In an attempt to improve the intermediate symbol recovery rate, the authors propose sending a number of degree‐1 input symbols prior to the build‐up phase based on a simple application of probability theory. In addition, during the completion phase, received encoded symbols with three neighbouring white (un‐decoded) symbols are retained for decoding and updating the decoding graph later. The performance and characteristics of the proposed improved online fountain codes are compared to those of the original online fountain codes over an erasure channel. Simulation results reveal that the improved online fountain codes outperform the original fountain codes in terms of intermediate symbol recovery rate, average encoded symbols required to be generated by the sender, average feedback transmissions, and encoding/decoding efficiency.
Yuli Zhao, Francis C. M. Lau 0002, Hai Yu 0001, Zhiliang Zhu 0001
IET Commun.5
2018 Using reliability risk analysis to prioritize test cases
Ying Wang 0038, Zhiliang Zhu 0001, Fangda Guo, Hai Yu 0001
J. Syst. Softw.2
2018 Exploiting self-adaptive permutation-diffusion and DNA random encoding for secure and efficient image encryption
Junxin Chen 0001, Zhiliang Zhu 0001, Li-bo Zhang 0004, Yushu Zhang 0001, Benqiang Yang
Signal Process.2
2018 An image encryption scheme using self-adaptive selective permutation and inter-intra-block feedback diffusion
Dong-dai Liu, Wei Zhang 0150, Hai Yu 0001, Zhiliang Zhu 0001
Signal Process.4
2018 Automatic Software Refactoring via Weighted Clustering in Method-Level Networks
abstract
In this study, we describe a system-level multiple refactoring algorithm, which can identify the move method, move field, and extract class refactoring opportunities automatically according to the principle of “high cohesion and low coupling.” The algorithm works by merging and splitting related classes to obtain the optimal functionality distribution from the system-level. Furthermore, we present a weighted clustering algorithm for regrouping the entities in a system based on merged method-level networks. Using a series of preprocessing steps and preconditions, the “bad smells” introduced by cohesion and coupling problems can be removed from both the non-inheritance and inheritance hierarchies without changing the code behaviors. We rank the refactoring suggestions based on the anticipated benefits that they bring to the system. Based on comparisons with related research and assessing the refactoring results using quality metrics and empirical evaluation, we show that the proposed approach performs well in different systems and is beneficial from the perspective of the original developers. Finally, an open source tool is implemented to support the proposed approach.
Ying Wang 0038, Hai Yu 0001, Zhiliang Zhu 0001, Wei Zhang 0150, Yuli Zhao
IEEE Trans. Software Eng.3
2017 Anticipatory Runway Incursion Prevention Based on Inaccurate Position Surveillance Information
Hai Yu 0001, Zhiliang Zhu 0001, Jingde Cheng
ACIIDS (2)3
2017 Relative health index of wind turbines based on kernel density estimation
abstract
Reducing operation and the maintenance costs of wind turbines has become a primary issue of wind farm owners and operators. Since the supervisory control and data acquisition (SCADA) system has been widely used in wind farms, it is costeffective to use SCADA data to realize condition monitoring. To this end, this paper proposes a method to calculate health index of wind turbines based on SCADA data. This method uses kernel density estimatison (KDE), then calculate the relative health index (RHI) of all wind turbines in a certain wind farm. To show the effectiveness of the method, we apply our method to real SCADA data of several wind farms. The result shows that RHI can reflect the health status of each wind turbine. Therefore, it is convenient to maintain the turbines with low RHI, thus to save maintenance costs.
Xiwei Liu, Hai Yu 0001, Zhiliang Zhu 0001
IECON4
2017 Wind turbine gearbox condition monitoring based on extreme gradient boosting
abstract
Currently, supervisory control and data acquisition (SCADA) systems are deployed in most wind farms, with the low cost of data acquisition. However, SCADA data contains a lot of redundant and dirty information, which makes it difficult to observe the actual condition of Wind Turbine (WT) or WT's components directly. In this paper, a gearbox condition monitoring (CM) model of WT based on SCADA data is proposed. In the first stage, after data preprocessing, the prediction model of gearbox oil temperature is obtained based on healthy data with Extreme Gradient Boosting (XGBoost), and the absolute percentage error (APE) of oil temperature is the final observational variable. In the second stage, the Multivariate Quality Control Charts (MQCC) is used to generate the threshold to detect the fault symptoms, combined with the APE of healthy data. Afterwards, the CM framework is established, which is capable of identifying the abnormal state of gearboxes based on whether the APE of new data exceeds the threshold. Finally, the effectiveness of the gearbox CM model presented is demonstrated by examining 2 groups of WT from different wind farms in China.
Yuechen Wang, Zhiliang Zhu 0001
IECON2
2017 A distributed energy-efficient cooperative routing algorithm based on optimal power allocation
abstract
Power consumption is a key issue for wireless network design. Cooperative routing, an efficient scheme for saving transmission power in wireless networks, combines the advantage of the cooperative communication in the physical layer and the routing technology in the network layer. In this paper, we propose a distributed minimum power cooperative routing algorithm based on optimal power allocation, namely the OPAMPCR algorithm that realizes minimum total transmission power for each source-destination pair. Each cooperative route from a source to a destination is a concatenation of cooperative transmission links and direct transmission links. For each type of link mode, the minimum transmission power required for a target BER is calculated. Especially, the proposed algorithm first finds the relay node which consumes the minimum total transmission power for assisting the data transmission. If such a relay does not exist, the direct transmission mode is applied in current link; Otherwise, we adopt the cooperative transmission mode. Then, the proposed OPAMPCR algorithm applies the distributed Bellman-Ford shortest-path routing algorithm to find the route with least transmission power required. Simulation results reveal that the OPAMPCR algorithm outperforms the SNCP algorithm, the CASNCP algorithm and the MPCR algorithm with respect to the average total transmission power per route and the power saving. Moreover, the proposed OPAMPCR algorithm achieve less hops per route than the MPCR algorithm.
Yuli Zhao, Hai Yu 0001, Zhiliang Zhu 0001
IECON3
2016 Image encryption based on three-dimensional bit matrix permutation
Wei Zhang 0150, Hai Yu 0001, Yuli Zhao, Zhiliang Zhu 0001
Signal Process.4
2015 Prevention of Fault Propagation in Web Service: a Complex Network Approach
Ying Liu 0032, Shu Mao, Mingwei Zhang 0001, Guoqi Liu, Zhiliang Zhu 0001, Jingde Cheng
J. Web Eng.5
2015 Reusing the permutation matrix dynamically for efficient image cryptographic algorithm
Junxin Chen 0001, Zhiliang Zhu 0001, Chong Fu 0001, Hai Yu 0001, Yushu Zhang 0001
Signal Process.2
2014 Fast Single Image Super-Resolution via Self-Example Learning and Sparse Representation
abstract
In this paper, we propose a novel algorithm for fast single image super-resolution based on self-example learning and sparse representation. We propose an efficient implementation based on the K-singular value decomposition (SVD) algorithm, where we replace the exact SVD computation with a much faster approximation, and we employ the straightforward orthogonal matching pursuit algorithm, which is more suitable for our proposed self-example-learning-based sparse reconstruction with far fewer signals. The patches used for dictionary learning are efficiently sampled from the low-resolution input image itself using our proposed sample mean square error strategy, without an external training set containing a large collection of high- resolution images. Moreover, the l0-optimization-based criterion, which is much faster than l1-optimization-based relaxation, is applied to both the dictionary learning and reconstruction phases. Compared with other super-resolution reconstruction methods, our low- dimensional dictionary is a more compact representation of patch pairs and it is capable of learning global and local information jointly, thereby reducing the computational cost substantially. Our algorithm can generate high-resolution images that have similar quality to other methods but with an increase in the computational efficiency greater than hundredfold.
Zhiliang Zhu 0001, Fangda Guo, Hai Yu 0001, Chen Chen 0001
IEEE Trans. Multim.1
2013 Anticipatory Emergency Elevator Evacuation Systems
Yuichi Goto, Zhiliang Zhu 0001, Jingde Cheng
ACIIDS (1)3
2013 Personalized Quality Prediction for Dynamic Service Management Based on Invocation Patterns
Bin Zhang 0001, Claus Pahl, Lei Xu 0004, Zhiliang Zhu 0001
ICSOC5
2013 A correlation context-aware approach for composite service selection
abstract
SUMMARY Composite service selection is one of the core research issues in Web service composition. Because of the complex service correlation context, candidate services may perform differently when being used with other services. Presently, most service selection approaches ignore this issue, which makes the selected composite services less efficient than expected. To solve this problem, a service correlation context‐aware composite service selection approach is proposed on the basis of the concept of single‐entry single‐exit (SESE) region. The general process of our approach is as follows: (1) mining the SESE patterns that are frequently used together in the set of efficiently executed instances of a composite service; (2) dividing the process model of the composite service into SESE regions and generating the candidate SESE pattern set of each region, using the discovered SESE pattern set; and (3) optimizing composite service selection globally on the basis of QoS using divided regions as selection units and their candidate pattern sets as candidate service sets. Because SESE patterns are testified by large amount of efficiently executed instances, they have higher quality than the results of independent selection of services in an SESE region. Experimental results demonstrated that our approach can improve the quality of selected composite services effectively in the correlation context. Concurrency and Computation: Practice and Experience, 2012.© 2013 Wiley Periodicals, Inc.
Mingwei Zhang 0001, Chengfei Liu, Jian Yu 0002, Zhiliang Zhu 0001, Bin Zhang 0001
Concurr. Comput. Pract. Exp.4
2012 Introducing SaaS Capabilities to Existing Web-Based Applications Automatically
Jie Song 0001, Zhenxing Yan, Yubin Bao, Zhiliang Zhu 0001
APWeb5
2011 A SaaSify Tool for Converting Traditional Web-Based Applications to SaaS Application
abstract
Nowadays, SaaS is increasingly used by web-based applications for the benefits and profits it brings to both users and service providers. It is significative if service providers can automatically convert traditional applications into SaaS mode, a SaaSify tool is needed urgently. In this paper, we analyze and conclude the new challenges of automatically SaaSify web-based application, propose several key technologies for SaaSifying, and further propose SaaSify Flow Language (SFL) to model and implement SaaSify process, finally, we use a case study to show the effects of proposed tool, and the performance experiments prove that the proposed approach is efficient and effective.
Jie Song 0001, Zhenxing Yan, Guoqi Liu, Zhiliang Zhu 0001
IEEE CLOUD5
2011 SLA Based Dynamic Virtualized Resources Provisioning for Shared Cloud Data Centers
abstract
Cloud computing focuses on delivery of reliable, secure, sustainable, dynamic and scalable resources provisioning for hosting virtualized application services in shared cloud data centers. For an appropriate provisioning mechanism, we developed a novel cloud data center architecture based on virtualization mechanisms for multi-tier applications, so as to reduce provisioning overheads. Meanwhile, we proposed a novel dynamic provisioning technique and employed a flexible hybrid queueing model to determine the virtualized resources to provision to each tier of the virtualized application services. We further developed meta-heuristic solutions, which is according to different performance requirements of users from different levels. Simulation experiment results show that these proposed approaches can provide appropriate way to judiciously provision cloud data center resources, especially for improving the overall performance while effectively reducing the resource usage extra cost and maximizing the global profit of cloud infrastructure providers.
Zhiliang Zhu 0001, Jing Bi 0001, Haitao Yuan 0001
IEEE CLOUD1
2011 A Petri Net Based Hybrid Optimal Controller for Deadlock Prevention in Web Service Composition
abstract
In the process of web service composition, the check and prevention of semantic incompatibility is one of the most important issues. In this paper, a controlled Petri net (CtlPN)-based model for web service composition is proposed. Meanwhile, the optimal controller is constructed, such that the appropriate vectors of controllable place and arc are appended in the key transition which can lead to deadlock states. In addition, for the semantic incompatibility case, a policy based on appending optimal controller is presented. It is proved that our policy can be a good solution. Finally, the proposed controller is transformed as the activity of BPEL.
Jing Bi 0001, Zhiliang Zhu 0001, Haitao Yuan 0001, Yushun Fan, Ming Tie
ICWS2
2011 Long-Term Benefit Driven Adaptation in Service-Based Software Systems
abstract
Service-based software system (SBS) is a software system based on service-oriented architecture (SOA). Although often treated as a composite service, an SBS is proposed from a more practical point of view based on restricted service provisions. In the highly competitive market, just meeting such requirements seems not enough to get more customers for service providers, and they usually provide additional preferential policies, such as a special order "buy-two-get-one-free". However, most of current adaptation approaches focus on single transaction, which makes it hard to take full advantage of such preferential policies in reselecting substitutable services. In this paper, we try to make the adaptation decision and reselect services from a broader view, i.e. expand the computation domain from single transaction to the whole lifecycle of an SBS by considering all of the past, current and predicable future executions. We call it "long-term benefit" to distinguish benefit in current approaches and propose a long-term benefit driven adaptation approach in this paper. In our approach, services that would bring the max expected long-term benefit would be selected and substituted into current instance in once adaptation. As the long-term benefit is accumulated in several executions, i.e. it depends on a decision sequence, we model the decision making problem as a sequential decision problem, and describe a realization based on partially observable Markov decision process (POMDP) for maximizing the real income in providing an SBS as an example.
Jun Na, Bin Zhang 0001, Yan Gao 0001, Zhiliang Zhu 0001
ICWS5
2011 A chaos-based symmetric image encryption scheme using a bit-level permutation
Zhiliang Zhu 0001, Wei Zhang 0150, Kwok-Wo Wong, Hai Yu 0001
Inf. Sci.1
2010 Dynamic Provisioning Modeling for Virtualized Multi-tier Applications in Cloud Data Center
abstract
Dynamic provisioning is a useful technique for handling the virtualized multi-tier applications in cloud environment. Understanding the performance of virtualized multi-tier applications is crucial for efficient cloud infrastructure management. In this paper, we present a novel dynamic provisioning technique for a cluster-based virtualized multi-tier application that employ a flexible hybrid queueing model to determine the number of virtual machines at each tier in a virtualized application. We present a cloud data center based on virtual machine to optimize resources provisioning. Using simulation experiments of three-tier application, we adopt an optimization model to minimize the total number of virtual machines while satisfying the customer average response time constraint and the request arrival rate constraint. Our experiments show that cloud data center resources can be allocated accurately with these techniques, and the extra cost can be effectively reduced.
Jing Bi 0001, Zhiliang Zhu 0001, Ruixiong Tian, Qingbo Wang
IEEE CLOUD2
2010 A Web Service QoS Prediction Approach Based on Collaborative Filtering
abstract
With the increasing numbers of Web services and service users on World Wide Web, predicting QoS(Quality of Service) for users will greatly aid service selection and discovery. Due to the different backgrounds and experiences of users, they have different QoS experiences when interacting with the same service. Even two users who have similar experiences on some services can have diverging views when considering other services. This paper proposes an approach to predict QoS. It is based on not only other users' QoS experiences, but also the environment factor and user input factor. First bring forwards usage information feature model and calculate the similarity of two users based on the feature model. Then consider not only the historic information, but also environment and users' inputs, such as bandwidth and data size. Before calculating the user similarity, select a set of Web services that have the highest degree of similarity with the target service, not all of the services. The missing value can be calculated through the data of similar services. The results of the experiment prove that our approach is feasible and effective.
Bin Zhang 0001, Ying Liu 0032, Yan Gao 0001, Zhiliang Zhu 0001
APSCC5
2010 Two-Stage Adaptation for Dependable Service-Oriented System
abstract
The development of Service-Oriented Systems has gained a considerable momentum as a means for building distributed applications and business processes. It emphasizes the loosely coupled construction of services from independent providers over the network. As a consequence, the dependability of such systems strongly depends on their ability to self-adapt to changes in its execution environment, such as unreachable component services, or changed delivered QoS. In this paper, we propose a two-stage approach to realize a self-adaptive SOA system, aimed at the fulfillment of dependability requirements. Specifically, we divide the adaptation process into two stages, proactive adaptation and reactive adaptation, to implement self-protecting and self-healing respectively, and provide a methodology driving the system adaptation. To bring this approach to fruition, a prototype system A-ServiceMix is developed by extending Apache ServiceMix.
Jun Na, Bin Zhang 0001, Zhiliang Zhu 0001, Dancheng Li
ICSS4
2010 Web Service Publication and Discovery Architecture Based on JXTA
abstract
The publication and discovery of web service is one of the most important issues in Service-oriented architecture. The common Universal Description, Discovery, and Integration (UDDI) still has some defects, such as register information can be invalidated since it is not able to update actively, its centralized approach leads to a high cost of server, and service providers can't delete register information in time when they stop to provide service. In this paper, first, a novel web service publication and discovery architecture based on JXTA is presented, which is applicable for internal enterprise. Then we make use of P2P network to publish and discover service, and propose distributed concurrent discovery mechanism to accelerate discovery speed. We also adopt a detector to discontinuously detect service, and update service information timely as well as remove unavailable service in the architecture. Finally, we prove that the proposed novel architecture can offer a good solution as well as can effectively resolve enterprise internal web service publication and discovery.
Jingqi Wei, Dancheng Li, Jun Na, Jing Bi 0001, Zhiliang Zhu 0001
ICSS5
2010 Web Service Composition Based on QoS Rules
Mingwei Zhang 0001, Bin Zhang 0001, Ying Liu 0032, Jun Na, Zhiliang Zhu 0001
J. Comput. Sci. Technol.5