Liudong Xing

dblp:81/6453 · DBLP profile ↗
← Back
99ranked-venue papers
19as first author
22since 2021 · last 2026
0000-0003-1606-1644ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 39 · 6 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 23 · 4 first-author · 4 since 2021Security and privacy · 12 · 2 first-author · 2 since 2021Computer networks · 11 · 5 first-author · 3 since 2021Systems, architecture and hardware · 10 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CausaLM-Net: An LLM-guided causal graph and state-space learning framework for fault diagnosis in cloud native 5G base stations
Hongyan Dui, Jiabao Zhai, Wanyun Xia, Liudong Xing, Haidong Shao, Ning Wang 0002
Expert Syst. Appl.4
2026 Network Recovery From Cascading Failures in Cyber-Service-Coupled Manufacturing Internet of Things
abstract
Tightly interdependent cyber and service networks in the Manufacturing Internet of Things (MIoT) drive cascading failures to propagate in intertwined horizontal (intra-layer) and vertical (cross-layer) directions, greatly complicating post-cascading-failure recovery decisions. To address limitations of existing approaches that neglect cross-layer dependencies and struggle to simultaneously handle heterogeneous load patterns and multiple failure states, this paper proposes a coordinated recovery framework for cyber-service-coupled MIoT, termed the coupled reinforcement learning (Coupled-RL) mechanism. Specifically, the Coupled-RL-based recovery method equips two layer-specific recovery agents for the cyber and service networks and a lightweight coordinator that orchestrates cross-layer decision-making. This coordinator is designed to avoid infeasible and globally suboptimal plans: its Feasibility Module (FM) shares the set of repaired nodes between layers and filters out actions that violate cross-layer prerequisites, while its Prediction Module (PM) exchanges per-state maximal target Q-values across layers—these values are used to construct a coupled return and inject cross-layer foresight into the Bellman update process of the agents. A weighted coupled return function and an alternating decision procedure further enable decentralized policy coordination. Extensive experiments demonstrate that the proposed recovery method effectively addresses the two-layer network coordination problem in MIoT during cascading failure recovery. Additionally, a comparative analysis between the proposed algorithm and existing ones is conducted to verify its superiority.
Jiayu Qian, Xiuwen Fu, Liudong Xing, Rui Peng 0001
IEEE Internet Things J.3
2026 Digital Twin-Enabled Smart Operation and Maintenance Framework With Generative AI Design of Intelligent Manufacturing Systems
abstract
Digital twin with generative artificial intelligence (AI)-enabled maintenance optimization serves as an essential foundation for the performance of intelligent manufacturing systems (IMS). However, existing models often fail to simultaneously consider both reliability and cost. In an IMS, reliability guarantees stable system operation and consistent product quality, while cost control enables enterprises to optimize resource use, enhance productivity, and lower operating costs. Together, these metrics determine the overall effectiveness of the system and the competitiveness of the enterprise. To address the research gap, this study proposes a maintenance optimization method that jointly considers reliability and cost. In particular, a novel reliability assessment method is developed, incorporating both physical failures modeled and functional outputs that account for imperfect quality inspection. Moreover, considering rework and imperfect quality inspection, a cost analysis is performed for various operation modes of IMS. Further, a novel adaptive multi-objective particle swarm optimization with maintenance priority constraints (AMOPSO-P) method is developed to conduct the IMS control decision-making process, optimizing reliability and cost. Finally, to validate the proposed algorithm, we conduct a case study of China United Equipment Group on control decisions for a three-stage, four-station servo valve manufacturing system using simulations.
Hongyan Dui, Hengbo Wang, Liudong Xing
IEEE Trans. Reliab.3
2026 Optimizing Mission Abort Policy With Threshold Voting of Imperfect Shock Detectors
abstract
Shock count is a key parameter used in designing mission abort policies (MAPs) for systems executing their operations under random shock conditions. Existing models mostly assume a perfect mechanism of detecting shocks. In practice, the shock detection system may fail to detect shocks that have occurred (false negative) or flag non-existent shocks (false positive), both leading to wrong shock count and misleading MAP designs. This work's contribution lies in modeling a single-attempt mission system with a fault-tolerant shock detection system that applies threshold voting among multiple imperfect detectors to contribute to the mission abort decision based on shock count and system operation time. A probabilistic approach is put forward for assessing mission performance of the considered system in the form of task success probability (TSP), survival probability of system (SPS), and expected losses of mission (ELM). An ELM minimization problem is further formulated and solved, which aims to determine the optimal tri-parametric MAP, achieving a balance between TSP and SPS. We analyze a drone-based surveillance system to showcase the suggested model. We also examine the impact of key parameters (cost, shock occurrence rate and detection probability) on mission performance metrics and on the best-obtained MAPs, leading to important managerial recommendations.
Gregory Levitin, Liudong Xing
IEEE Trans. Reliab.2
2026 Phased UAV Swarm Mission Reliability Assessment Considering Cascading Failures
abstract
In many critical real-world scenarios (e.g., emergency response, environmental protection, infrastructure surveillance), based on the Internet of Things (IoT) technology, multiple UAV swarms work together to acquire, transmit, and process information, subsequently taking action to provide the intended service. Collaborations and interactions among UAVs create interdependencies that can facilitate the propagation of failures, potentially leading to high-impact cascading effects. Leveraging the hierarchical architecture of IoT, this work examines the cascading failure mechanism triggered by the relay malfunction and its impact on mission reliability of a heterogeneous multi-phase UAV swarm system. To analyze mission reliability under cascading effects, a separable and analytical modeling method is put forward, which decomposes the original problem into independent, reduced reliability problems for efficient analysis. A detailed case study on a two-phase rescue mission showcases the model's ability to capture failure propagation and assess its impact. The suggested approach offers an efficient means of quantifying cascading effects, enabling mission planners to identify vulnerabilities and improve overall resilience strategies. Furthermore, based on the proposed mission reliability analysis method, resource allocation problems are formulated and solved to balance cost and mission reliability in practical mission design.
Junxing Ren, Liudong Xing
IEEE Trans. Reliab.2
2026 Task-Oriented Reliability Modeling and Analysis of Federated Learning-Enabled Intelligent Manufacturing Systems
abstract
Intelligent manufacturing systems progressively advance toward highly collaborative and distributed decision making. Federated learning (FL)-enabled intelligent manufacturing systems facilitate efficient data processing and intelligent decision making by deploying artificial intelligence models at the edge and integrating model aggregation mechanisms. However, anomalies occurring in the local devices on which the local models depend can propagate through the aggregation process, leading to model contamination and performance degradation, thereby compromising the task reliability of the entire system. To address this issue, this article proposes a reliability modeling approach for FL-enabled intelligent manufacturing systems (FL-IMSs). In the proposed model, we characterize the performance degradation process of the local model caused by terminal device failures and its impact on task reliability. This process includes the effect of data quality degradation triggered by device failures on model performance, as well as failure propagation caused by intermodel dependencies. Furthermore, to assess the impact of model performance variations on practical production tasks, a task-oriented reliability metric is introduced. Simulation and experimental results demonstrate that the proposed modeling approach effectively captures local model performance degradation and task reliability in FL-IMSs under terminal device failure conditions.
Xiaoluoteng Song, Xiuwen Fu, Liudong Xing, Rui Peng 0001
IEEE Trans. Reliab.3
2025 Optimizing Power Resilience Performance of Intelligent Solar Photovoltaic System for Smart Energy Management Considering Reliability and Cost
abstract
Due to being nonpolluting and renewable, intelligent solar photovoltaic (PV) technology is widely used to provide electricity and becomes a cornerstone to sustainable energy and smart energy management. Different from existing studies that improve the PV efficiency by changing cell materials, this article proposes a novel system reliability and cost model of enhancing the PV power resilience performance from the perspective of optimizing the number of PV panels. Specifically, a multiobjective planning model is proposed, which determines the optimum number of spare parts for PV panels maximizing the output power resilience while maximizing the system reliability and minimizing the cost. The reliability measures the probability of stable operation of a PV panel considering the no-power output state. The cost factor encompasses negative cost of environmental benefits, resource cost, operation and maintenance cost, and penalty cost. Experiments are performed on fifty sets of Pareto optimal solutions in summer and winter cases to illustrate effectiveness of the proposed method by using a ground-mounted PV project in Zhongwei City, China.
Hongyan Dui, Yaohui Lu, Liudong Xing
IEEE Trans. Reliab.3
2025 Dynamic Reliability Assessment Model for IoT-Enabled Smart Offshore Wind Farm
abstract
Offshore wind farm is one of the most promising applications in the Internet of Things (IoT), due to being energy-renewable and resources-unlimited. However, the reliability monitoring and maintenance models of power equipment based on communication paths and sensors are still immature in the smart offshore wind farm (SOWF). Based on the hierarchical architecture and end-to-end communication, a dynamic reliability assessment model (DRAM) is proposed for SOWFs. First, based on the IoT hierarchy, a four-stage network is developed to represent the relationship or dependencies between diverse devices in a complex SOWF. Second, a two-layer DRAM with forward monitoring (FM) and lateral protection (LP) is proposed. The FM encompasses a sensor network-based state-monitoring phase (monitoring weather data like temperature and wind speed), and a data-monitoring phase (monitoring the reliability-related data like reception power and data processing speed). The LP includes a signal-protection mode (LP-I) ensuring that virtual machines read the data and issue protection orders before turbine failures to minimize losses, and a radius-maintenance model (LP-II) performing maintenance of the failed turbine nodes. Simulation results show that the optimal maintenance strategy based on DRAM outperforms the benchmark maintenance method for traditional wind grids.
Hongyan Dui, Xinmin Wu, Liudong Xing
IEEE Trans. Reliab.4
2025 Reliability Modeling and Analysis of Digital Twin-Driven Cyber-Physical Manufacturing Systems
abstract
With the advancement of digitalization and intelligentization in manufacturing systems, digital twin-driven cyber-physical manufacturing systems (DT-driven CPMSs) have emerged as a key technology for enabling smart manufacturing. Existing studies have primarily focused on the applications of DT technology, but have not fully addressed the reliability challenges arising from equipment degradation and sudden failures during system operation. To address this challenge, we propose an interdependent network model for DT-driven CPMSs that integrates real-time sensing and control feedback dependencies across the cyber layer, physical layer, and virtual decision space. The model emphasizes the characterization of data dependencies between devices under sensing-control dependencies, including production data support and collaborative production dependencies. Based on the proposed system model, we further develop a system reliability model. By incorporating the routing-driven characteristics of data in the cyber layer and the material supply-demand relationships among equipment in the physical layer, the proposed reliability model enables the joint modeling of long-term equipment degradation and sudden failure propagation under sensing-control dependencies within the system. Experimental results demonstrate that the proposed model can effectively capture system reliability behavior under these challenging operational conditions. Further analysis reveals that although the cyber layer constitutes a key bottleneck for system reliability, the physical layer is more effective in regulating it. Specifically, the average gain in system reliability achieved through redundancy enhancement in the physical layer reaches 0.82, which is significantly higher than the 0.39 gain achieved in the cyber layer.
Xiuwen Fu, Dingyi Zheng, Liudong Xing, Rui Peng 0001
IEEE Trans. Reliab.3
2025 Reliability Modeling and Assessment of Internet-of-Things in Smart Manufacturing Systems: A Modular Petri Net Approach
abstract
The Internet of Things (IoT) represents a transformative convergence of traditional manufacturing systems with advanced information technologies, collectively referred to as smart manufacturing. The interconnected nature of IoT facilitates real-time data collection and analysis, optimizing production processes and improving operational efficiency. However, the increased complexity and interdependence of IoT systems pose significant challenges in reliability modeling and assessment. This article introduces a novel reliability model that comprehensively integrates factors such as degradation of physical systems and information networks, along with their interactive impacts on system performance and reliability. A modular Petri net approach is developed to efficiently assess reliability of IoT systems by leveraging a structured framework to model the intricate interdependencies within IoT. The modular nature of the proposed approach enables targeted analysis and scalability enhancements, addressing the critical need for models that can adapt to the evolving landscape of IoT in smart manufacturing systems. A vehicle manufacturing system example is introduced to demonstrate the proposed approach. The results reveal the distinct impact pathways of the physical system and information network on overall system reliability. Statistical analysis across various system configurations shows that modifying the architecture of the information network can lead to an average improvement of 20.67% in system reliability.
Yu Liu 0006, Liudong Xing, Hong-Zhong Huang
IEEE Trans. Reliab.3
2025 Modeling and Analysis of Cascading Failures in Industrial Internet of Things Considering Sensing-Control Flow and Service Community
abstract
Cascading failures are a critical factor affecting the reliability of industrial Internet of things (IIoT) systems. Establishing a realistic cascading failure model is of significant importance for researching and improving the reliability of IIoT. However, existing research on cascading failure modeling for IIoT lacks in-depth exploration of the actual characteristics of industrial scenarios, making it difficult to accurately characterize the cascading failure process in IIoT. In this work, based on the cyber–service coupling characteristic of IIoT systems, we establish a realistic interdependent network model, taking into full consideration the sensing-control data flow, the service community structure, and the diverse coupling patterns. On this basis, a cascading failure model for IIoT is developed, considering the routing-driven characteristic of the cyber network and the production–supply relationships among various manufacturing units in the service network. Extensive experiments are conducted to verify the rationality of the proposed model, and some meaningful findings are also obtained.
Dingyi Zheng, Xiuwen Fu, Liudong Xing, Rui Peng 0001
IEEE Trans. Reliab.4
2024 Mission Aborting Policies and Multiattempt Missions
abstract
The state of the art in the recently emerged and rapidly developing field of mission aborting and multiattempt missions is briefly discussed. The research aims to develop optimal rules for interrupting a mission and activating system rescue procedures (and, if needed, subsequent attempts to complete the mission) that balance the probabilities of mission success and system loss or minimize the cost of losses associated with the mission.
Gregory Levitin, Liudong Xing
IEEE Trans. Reliab.2
2023 Reliability Theory and Practice for Unmanned Aerial Vehicles
abstract
Due to rapid advancements on the Internet of Things (IoT), unmanned aerial vehicles (UAVs), also known as drones, are transforming numerous military and civil application areas. UAVs aim to enhance the production efficiency, ensure safety, and reduce risk, particularly protecting the human workforce in the case of harsh and dangerous environments. Due to the mission-critical, business-critical, or safety-critical nature of the UAV applications, it is pivotal that UAVs perform reliably to deliver the required service during the intended mission time. Therefore, reliability is one of the essential requirements for designing and operating UAVs. This article presents a critical review of UAV reliability literature in both theoretical and practical research, pinpointing failure causes and reliability challenges of UAV systems, classifying and reflecting on the reliability modeling, analysis, and design methods for UAV systems and key subsystems. Some open research problems and opportunities are also discussed to highlight potential new challenges for designing reliable and resilient UAVs and UAV-assisted IoT systems.
Liudong Xing, Barry W. Johnson
IEEE Internet Things J.1
2023 Reliability Analysis of Dynamic Load-Sharing Systems With Constrained and Changing Component Performances
abstract
Considerable research efforts have been expended in modeling load-sharing systems. The existing models, however, have various limitations, such as being limited to the exponential time-to-failure distribution, constant component performances, or performances without constraints. In this article, we make contributions by modeling a dynamic load-sharing system (DLSS), where the performance of each component is dynamic according to prespecified load-sharing principles and is limited by its capacity constraint. Moreover, the capacity constraint of a component can reduce due to degradations. In the proposed model, increasing failure rates are also involved since the surviving components must share the load of the failed component and continue working with increasing stresses. When the desired performance for a component exceeds the limitation, the entire system fails. An extended Markov process (EMP) method is proposed for evaluating the reliability of the considered DLSS with nonrepairable components. The proposed analytical method is flexible in handling arbitrary component time-to-failure distributions and in handling diverse load allocation mechanisms. Numerical studies of a power transmission system and a water transmission system are provided to validate the proposed method and its advantages. Effects of several model parameters are also investigated through case studies.
Heping Jia, Liudong Xing, Yi Ding 0001, Dunnan Liu
IEEE Trans. Syst. Man Cybern. Syst.2
2022 Efficient Analysis of Resource Availability for Cloud Computing Systems to Reduce SLA Violations
abstract
Resource availability of cloud computing systems is vital for today's information infrastructures. However, the constituent computing nodes are not fault free. The availability requirements on the delivered computing resources should be defined clearly by using a formal, contractual agreement, known as the Service Level Agreement (SLA), between providers and customers. To reduce the risk of various SLA violations, an efficient analysis of resource availability is important. This paper proposes a new analytical approach based on multi-valued decision diagrams (MDD) for the efficient resource availability analysis of cloud computing systems with heterogeneous, multi-state computing nodes. Particularly, a novel and efficient MDD construction method is presented to generate compact MDD models encoding different amounts of cumulative computing resources. Two detailed case studies are performed to illustrate basics and application of the proposed approach to reduce SLA violations and guarantee the availability requirements on the delivered computing resources. Benchmark studies are further conducted to show efficiency of the proposed MDD-based approach as compared with the continuous-time Markov chains-based method and the universal generation function-based method.
Yuchang Mo, Liudong Xing
IEEE Trans. Dependable Secur. Comput.2
2022 Reliability versus Vulnerability of N-Version Programming Cloud Service Component With Dynamic Decision Time Under Co-Resident Attacks
abstract
The virtual machine (VM) co-resident architecture of cloud computing enables simultaneous provision of multiple services to different users, but also makes these services vulnerable to co-resident attacks. For example, by establishing side channels, a malicious attacker can access and even corrupt services performed by other VMs co-residing on the same server as the attacker's VM (AVM). We model a threshold-voting-basedN-version programming service component with multiple independent versions simultaneously performing the same requested service to enhance the service reliability. However, the reliability enhancement can be greatly hindered by the co-resident attack, which may corrupt an adequate number of versions leading to a wrong output. We formulate and solve constrained optimization problems that determine the number of service component versions and the voting threshold to balance two conflicting service performance metrics: reliability (service component success probability) and vulnerability (service corruption attack success probability). Two cases respectively having certain and uncertain knowledge about the attacker's power in terms of the number of AVMs are considered. We also investigate impacts of different model parameters on the service performance as well as on solutions to the considered optimization problems through examples.
Gregory Levitin, Liudong Xing, Yanping Xiang
IEEE Trans. Serv. Comput.2
2022 Mission Aborting in n-Unit Systems With Work Sharing
abstract
Mission aborting has recently attracted great attention, where mission abort rules (MARs) have been modeled and optimized for different types of technological systems aiming to effectively mitigate the risk of system losses. However, none of the existing works have considered systems with multiple work-sharing units. This article makes advancements in the state of the art by modeling condition-based MARs for nonrepairable work-sharing systems that must perform a specified amount of work during the primary mission (PM). The MAR considered presumes aborting the PM to prevent considerable damage when the number of available units reduces to a certain number${k}$while the amount of work accomplished in the PM is less than${L}$(${k}$). After the PM abortion, a rescue procedure (RP) is executed by the remaining units to survive the system. Dynamic operating conditions during PM and RP are considered. A probabilistic model-based numerical algorithm is proposed to evaluate several performance metrics, including mission success probability, RP success probability, damage avoidance probability, and expected cost of losses. The MAR optimization problem for minimizing the expected cost of losses is formulated and solved using the genetic algorithm. An example of a chemical reactor system is provided to demonstrate the proposed algorithm as well as the benefit of MARs optimized in comparison to the “no abort” policy and the most conservative abort policy (the PM aborts upon the failure of any unit). Effects of several model parameters on the system performance metrics and optimization solutions are also examined through examples.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Syst. Man Cybern. Syst.2
2022 Co-Residence Data Theft Attacks on N-Version Programming-Based Cloud Services With Task Cancelation
abstract
Powered by virtualization, the cloud computing has brought good merits of cost effective and on-demand resource sharing among many users. On the other hand, cloud users face security risks from co-residence attacks when using this virtualized platform. Particularly, a malicious attacker may create side channels to steal data from a target user’s virtual machine (VM) that co-resides with the attacker’s VM on the same physical server. This article models a cloud service undergoing the co-residence data theft attacks. The threshold-voting-based${N}$-version programming (NVP) is implemented to improve the service reliability, where multiple service component versions (SCVs) are activated in parallel to perform the requested service. The final output is determined upon receiving a threshold number of identical outputs from the SCVs, immediately followed by canceling all outstanding SCVs to reduce expenses. Probabilistic models are first introduced to evaluate performance metrics of the considered service, including the data theft probability, service success probability, expected service operation time, and expected utility. Optimization problems are further solved to find the optimal number of SCVs maximizing the expected utility. Interactions among different model parameters and VM allocation policies, as well as their effects on the considered performance metrics and on the optimization solutions are studied through examples.
Gregory Levitin, Liudong Xing, Yanping Xiang
IEEE Trans. Syst. Man Cybern. Syst.2
2021 Cascading Failures in Internet of Things: Review and Perspectives on Reliability and Resilience
abstract
In the Internet of Things (IoT), various devices operate collaboratively in collecting data, relaying information to one another, and processing information intelligently. Due to interactions and dependencies between the IoT devices, the malfunction of one device may trigger a cascade of unexpected and often undesired state changes of other devices, introducing or accelerating catastrophic cascading failures. Understanding the causes of cascading failures and modeling their behavior and effects is crucial for guaranteeing the reliability of IoT systems and delivering the desired quality of service. This article systematically reviews cascading failures modeling and reliability analysis methodologies, as well as mitigation strategies for building the resilience of IoT systems against cascading failures. The review covers diverse IoT applications, from smart grids to smart homes, from sensor networks to IoT cloud computing, and from transportation networks to interdependent infrastructure networks. Opportunities and open research issues are also discussed in relation to restrictions of the current cascading failure models and methods, and potential new technologies and complexity of the constantly evolving IoT systems.
Liudong Xing
IEEE Internet Things J.1
2021 Efficient Analysis of Repairable Computing Systems Subject to Scheduled Checkpointing
abstract
To improve the success probability of a mission execution, scheduled checkpointing is often implemented to save completed portions of the mission task so that a system can resume the mission execution effectively after its restoration whenever the system failure occurs. This paper considers a repairable computing system subject to the scheduled checkpointing. The checkpointing intervals are deterministic, but can be even or uneven. The system repair time is fixed while the system time-to-failure can follow any arbitrary type of distributions. The maximum number of repairs is specified by a certain threshold value. A multi-valued decision diagram (MDD)-based analytical approach is proposed to evaluate the exact success probability of a mission execution for the considered repairable system. The proposed approach enables generating a compact mission MDD model where identical subMDD models can be merged to improve computational efficiency and reduce storage requirement. The MDD model, once being constructed, can be reused for system reliability evaluations using different input parameter values. A benchmark study is presented to show the efficiency of proposed MDD approach. A case study is performed to illustrate the application of the proposed MDD approach to facilitate decision making about proper system design and parameter selection.
Yuchang Mo, Liudong Xing, Yi-Kuei Lin, Wenzhong Guo
IEEE Trans. Dependable Secur. Comput.2
2021 Defending N-Version Programming Service Components against Co-Resident Attacks in IoT Cloud Systems
abstract
The real innovation of Internet of Things (IoT) can be spurred only when being combined with cloud computing, a paradigm that allows numerous users to simultaneously access configurable resources and services. However, serious vulnerability concerns have arisen from the virtual machine co-resident architecture of the IoT cloud. Specifically, co-resident attacks can be launched, where an attacker can access and corrupt a user's sensitive data/software by co-locating their virtual machines on the same physical server. Various solutions have been suggested in literature to mitigate negative effects of the co-resident attacks in the cloud environment. However, to the best of our knowledge no work has been performed for studying co-resident attacks in cloud systems withN-version programming (NVP), a popular redundancy technique for enhancing survivability of critical cloud service components. This paper makes original contributions by modeling IoT cloud system services implementing the NVP component redundancy, and evaluating the corruption probability of the NVP service component. Further, users’ policies on choosing the optimal number of service component versions are investigated through formulating and solving a new set of optimization problems with the objective to minimize the expected cost of losses of a cloud service provider. As demonstrated through examples, these policies can effectively help defend the NVP service component against the co-resident attacks in the cloud system.
Liudong Xing, Gregory Levitin, Yanping Xiang
IEEE Trans. Serv. Comput.1
2021 Economic Design of a Linear Consecutively Connected System Considering Cost and Signal Loss
abstract
Linear multistate consecutively connected systems (LMCCSs) have been widely applied in telecommunications. An LMCCS usually has several nodes arranged in sequence along a line, where connecting elements (CEs) are deployed at each node to provide connections to the following nodes. Many researchers have studied the reliability modeling and optimization of LMCCSs. However, most of the existing works on LMCCSs have focused on the uncertainty in connection ranges of CEs; none of them have considered signal loss during the transmission. In practice, a signal emitted from a node may neither completely reach nor completely not reach the destination node. In other words, only a fraction of the signal may reach the destination node whereas the rest is lost. This article makes new contributions by proposing a model that evaluates the expected signal fraction receivable by the sink node in an LMCCS subject to signal loss. Moreover, we solve the optimal design policy problem, which co-determines CEs allocation and nodes building to minimize the system cost while meeting certain constraints on system reliability and expected receivable signal fraction. Three examples are provided to illustrate the proposed model.
Kaiye Gao, Xiangbin Yan, Rui Peng 0001, Liudong Xing
IEEE Trans. Syst. Man Cybern. Syst.4
2020 Reliability in Internet of Things: Current Status and Future Perspectives
abstract
The Internet of Things (IoT) aims to transform the human society toward becoming intelligent, convenient, and efficient with potentially enormous economic and environmental benefits. Reliability is one of the main challenges that must be addressed to enable this revolutionized transformation. Based on the layered IoT architecture, this article first identifies reliability challenges posed by specific enabling technologies of each layer. This article then presents a systematic synthesis and review of IoT reliability-related literature. Reliability models and solutions at four layers (perception, communication, support, and application) are reflected and classified. Despite the rich body of works performed, the IoT reliability research is still in its early stage. Challenging research problems and opportunities are then discussed in relation to current underexplored behaviors and future new aspects of evolving IoT system complexity and dynamics.
Liudong Xing
IEEE Internet Things J.1
2020 Modeling and Analyzing Linear Wireless Sensor Networks With Backbone Support
abstract
Rapid advancement in micro-electromechanical techniques leads to the wide application of wireless sensor networks (WSNs). In a linear WSN (LWSN), all sensor nodes are arranged in a straight line to monitor health status of some linear infrastructure structure such as bridges, highways, pipelines, etc. To enhance reliability of the infrastructure monitoring services, LWSNs are often designed to incorporate a limited number of backbone nodes for transferring or relaying information, leading to a more complex hybrid structure. In this paper, a multivalued decision diagram (MDD)-based analytical approach is proposed to evaluate performance of an LWSN system with backbone nodes. Particularly, we model and analyze the probability that the hybrid LWSN performs at a particular performance level, which is characterized by the number of sensor nodes being able to reach the base station. A single compact MDD model is constructed by sharing all isomorphic submodel structures involved in different performance levels. The MDD model, once being constructed, can be reused for evaluation using different failure time distributions or mission time. A case study is presented to substantiate the application of the proposed MDD approach for developing the optimal backbone node allocation strategy to guarantee the reliability requirement on the infrastructure monitoring services.
Yuchang Mo, Liudong Xing, Jianhui Jiang
IEEE Trans. Syst. Man Cybern. Syst.2
2019 Correlation Modeling and Resource Optimization for Cloud Service With Fault Recovery
abstract
Energy-efficient cloud computing has recently attracted much attention, where not only performance but also energy consumption are important metrics to be considered for designing rational resource scheduling strategies. Most of existing approaches for achieving energy efficient computing focus on connecting these two metrics and balancing the tradeoff between them, which however is inadequate because another important factor reliability is not considered. In fact, both virtual machine (VM) failures and server failures inevitably interrupt execution of a cloud service, and eventually result in spending more time and consuming more energy on completing the cloud service. Therefore, reliability significantly affects service performance and energy consumption, and thus they should not be handled separately. Connecting these correlated metrics is essential for making more precise evaluation and further for developing rational cloud resource scheduling strategies. In this paper, we present a correlated modeling approach applying Semi-Markov models, the Laplace-Stieltjes transform (LST), a Bayesian approach to analyze reliability-performance (R-P) and reliability-energy (R-E) correlations for cloud services using a retrying fault recovery mechanism. A recursive method is also proposed for modeling the correlations for cloud services using a check-pointing fault recovery mechanism. The proposed correlation models can be used to calculate the expected service time and energy consumption for completing a cloud service. Moreover, the models can contribute to analyzing the expected performance-energy tradeoff. We formulate the expected performance-energy optimization problem by describing performance and energy consumption metrics as functions of assigned CPU frequencies. Finally, we use a derivation approach to determine Pareto optimal solutions for the formulated optimization problem. Illustrative examples are provided.
Xiwei Qiu, Yuan-Shun Dai, Yanping Xiang, Liudong Xing
IEEE Trans. Cloud Comput.4
2019 Optimal Spot-Checking for Collusion Tolerance in Computer Grids
abstract
Many grid-computing systems adopt voting-based techniques to resist sabotage. However, these techniques become ineffective in grid systems subject to collusion behavior, where some malicious resources can collectively sabotage a job execution by returning identical wrong results. Spot-checking has been used to detect and tackle the collusive issue by sending randomly chosen resources a certain number of spotter jobs with known correct results to estimate resource credibility based on the returned result. This paper makes original contributions by formulating and solving a new spot-checking optimization problem for grid systems subject to collusion attacks, with the objective to minimize probability of the genuine task failure (PGTF, i.e., the wrong output probability) while meeting an expected overhead constraint. The problem solution contains an optimal combination of task distribution policy parameters, including the number of deployed spotter tasks, the number of resources tested by each spotter task, and the number of resources assigned to perform the genuine task. The optimization procedure encompasses a new iterative method for evaluating system performance metrics of PGTF and expected overhead in terms of the total number of task assignments. Both fixed and uncertain attack parameters are considered. Illustrative examples are provided to demonstrate the proposed optimization problem and solution methodology.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Dependable Secur. Comput.2
2018 Optimization of dynamic spot-checking for collusion tolerance in grid computing
Gregory Levitin, Liudong Xing, Barry W. Johnson, Yuan-Shun Dai
Future Gener. Comput. Syst.2
2018 Performability Analysis of Large-Scale Multi-State Computing Systems
abstract
Modern computing systems typically use a large number of independent, non-identical computing nodes to perform a set of coordinated computations in parallel. The computing system and its constituent computing nodes often exhibit more than two performance levels or states corresponding to different computing powers. This paper models and evaluates performability of largescale multi-state computing systems, which is the probability that a computing system performs at a particular performance level. The heterogeneity in the constituent components of different nodes (due to factors such as different model generations, model suppliers, and operating environments) makes performability analysis difficult and challenging. In this paper a specification method for system performance level (SPL) is first introduced. A multi-valued decision diagram (MDD) based approach is then proposed for performability analysis of multi-state computing systems consisting of nodes with different state occupation probabilities, which encompasses novel and efficient MDD model generation procedures. Example and benchmark studies are performed to show that the proposed approach can offer efficient performability analysis of large-scale computing systems.
Yuchang Mo, Lirong Cui, Liudong Xing, Zhao Zhang 0002
IEEE Trans. Computers3
2018 Performability Analysis of k-to-l-Out-of-n Computing Systems Using Binary Decision Diagrams
abstract
Modern computing systems typically utilize a large number of computing nodes to perform coordinated computations in parallel or simultaneously. They can exhibit multiple performance states or levels due to statuses or failures of their consistent nodes. Performability analysis is concerned with assessing the probability that the computing system performs at a particular performance level. In the context of performability analysis, these computing systems can be modeled using k-to-l-out-of-n structures. This paper proposes new analytical methods based on binary decision diagrams (BDD) for the performability analysis of large computing systems with unrepairable computing nodes. A new and efficient BDD algorithm that makes full uses of the special k -to-l-out-of-n structure is first proposed for systems with computing node having identical computing powers. New simplification rules are further proposed to generate compact and canonical BDD models for systems with heterogeneous computing nodes characterized by different computing powers. Ordering heuristic is also explored to further reduce the size of BDD models. Examples are provided to illustrate the proposed BDD-based performability analysis methodology as well as its efficiency in analyzing large-scale computing systems.
Yuchang Mo, Liudong Xing, Joanne Bechta Dugan
IEEE Trans. Dependable Secur. Comput.2
2018 Mission Abort Policy in Heterogeneous Nonrepairable 1-Out-of-N Warm Standby Systems
abstract
Many real-world critical systems, such as aircraft and human space flight systems, utilize mission aborts to enhance the survivability of the system. Specifically, the mission objectives of these systems can be aborted in cases where a certain malfunction condition is met, and a rescue or recovery procedure is then initiated for system survival. Traditional system reliability models typically cannot address the effects of mission aborts, and thus are not applicable to analyzing systems subject to mission abort requirements. In this paper, we first develop a numerical methodology to model and evaluate mission success probability and system survivability of 1-out-of-N warm standby systems subject to constant or adaptive mission abort policies. The system components are heterogeneous, characterized by different performances and different types of time-to-failure distributions. Based on the proposed evaluation method, we make another new contribution by formulating and solving the optimal mission abort problem, as well as a combined optimization problem that identifies the mission abort policy and component activation sequence maximizing mission success probability while achieving the desired level of system survivability. Efficiencies of constant and adaptive mission abort policies are compared through examples. Examples also demonstrate the tradeoff between system survivability and mission success probability due to the utilization of a mission abort policy. Such a tradeoff analysis can help identify optimal decisions on system mission abort and standby policies, promoting safe and reliable operation of warm standby systems.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Reliab.2
2018 System Reliability Modeling Considering Correlated Probabilistic Competing Failures
abstract
A combinatorial system reliability modeling method is proposed to consider the effects of correlated probabilistic competing failures caused by the probabilistic-functional-dependence (PFD) behavior. PFD exists in many real-world systems, such as sensor networks and computer systems, where functions of some system components (referred to as dependent components) rely on functions of other components (referred to as triggers) with certain probabilities. Competitions exist in the time domain between a trigger failure and propagated failures of corresponding dependent components, causing a twofold effect. On one hand, if the trigger failure happens first, an isolation effect can take place preventing the system function from being compromised by further dependent component failures. On the other hand, if any propagated failure of the dependent components happens before the trigger failure, the propagation effect takes place and can cause the entire system to fail. In addition, correlations may exist due to the shared trigger or dependent components, which make system reliability modeling more challenging. This paper models effects of correlated, probabilistic competing failures in reliability analysis of nonrepairable binary-state systems through a combinatorial procedure. The proposed method is demonstrated using a case study of a relay-assisted wireless body area network system in healthcare. Correctness of the method is verified using Monte-Carlo simulations.
Liudong Xing, Honggang Wang 0001, David W. Coit
IEEE Trans. Reliab.2
2018 Optimizing Dynamic Performance of Multistate Systems With Heterogeneous 1-Out-of-N Warm Standby Components
abstract
This paper models and optimizes dynamic performance of multistate systems with a general series parallel structure. Each system component is a 1-out-of-N warm standby configuration of heterogeneous functional elements, which can be characterized by different time-to-failure distributions, performances, and costs. The entire system must satisfy a random demand specified by a time-dependent distribution. An iterative algorithm is developed for determining performance stochastic processes of particular components. A universal generating function technique is used for evaluating expected system availability and unsupplied demand over a particular mission time for the considered system. Two types of optimization problems are then identified and solved, with the objective of finding component structures and element activation sequences to maximize system availability, or minimize unsupplied system demand, or minimize total cost. Optimization results can facilitate the optimal decision on design and operation of multistate series parallel systems. A practical example of a power station coal transportation system is provided to illustrate application of the proposed methodology and optimization problems.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Syst. Man Cybern. Syst.2
2018 Optimizing Computational Mission Operation by Periodic Backups and Preventive Replacements
abstract
This paper models a warm standby system where a single element is online performing a specified mission task (e.g., a computing task) and subject to corrective replacement (CR) by an available standby element upon its failure. During the mission, preventive replacements (PRs) are also performed to renew the aged or worn online operating element before its actual failure according to a predetermined policy. In addition, to facilitate an effective restoration of system function in case of CR or PR happening, backups are also performed periodically so that the mission task can be resumed from the last successful backup point instead of from scratch. The mission succeeds if the specified mission task is accomplished; in other words, the mission fails when no operating elements remain prior to the mission task completion. In this paper, we make new contributions by first proposing an event transition-based numerical method to evaluate mission performance indices of the considered standby system subject to periodic backups, CR and PR. Mission success probability (MSP), expected mission completion time, expected mission operation cost (EMC), and expected uncompleted work fraction are evaluated. Based on the suggested evaluation algorithm, we make another contribution by formulating and solving optimization problems that help to determine the optimal backup-PR policy or the optimal combination of element activation sequencing and backup-PR policy to maximize MSP or minimize EMC. Influence of element performance and reliability parameters, data backup and retrieval complexity parameters on the optimal operation policy is investigated. Findings from this paper can guide the optimal decision making on policies related to element sequencing, backups as well as preventive maintenance planning, contributing toward reliable and cost-effective design and operation of standby computing systems.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Syst. Man Cybern. Syst.2
2017 Optimal data partitioning in cloud computing system with random server assignment
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
Future Gener. Comput. Syst.2
2017 Reliability Modeling of Mesh Storage Area Networks for Internet of Things
abstract
With advances in Internet of Things (IoT), intelligent data sensors are being added to more and more devices that interact with human's daily life in areas, such as medical services, smart grids, and financial services. IoT has made big contributions to data growth, requiring highly reliable data storage solutions. Storage area networks (SANs) are one of such solutions. To meet high reliability and availability requirements, SANs have to provide fault tolerance through redundancy to minimize or eliminate system downtime, thus preventing business discontinuity due to catastrophic events. Mesh is one of the common SAN topologies that have been applied to implement a fault tolerant SAN in practice. In this paper, failure behavior of a mesh SAN is modeled using a dynamic fault tree (DFT) in the case of perfect links, or a network graph in the case of imperfect links. Based on the constructed DFT or network graph model, reliability of the mesh SAN is evaluated using a binary decision diagram-based method. Results obtained from the case study can provide insights into the behavior of general mesh SAN systems, providing guidelines in the reliable design and operation of SANs.
Liudong Xing, Massarrah Tannous, Vinod Vokkarane, Honggang Wang 0001
IEEE Internet Things J.1
2017 Dynamic Checkpointing Policy in Heterogeneous Real-Time Standby Systems
abstract
This paper models 1-out-of-N standby computing systems with a dynamic checkpointing policy. The system performs a real-time mission task that has to be accomplished within an allowed mission time. During the mission, to facilitate an effective failure recovery the system undergoes checkpointing procedures according to a policy that dynamically determines a checkpointing frequency based on the activated element and remaining work for completing the mission. System elements are heterogeneous; they can follow different, arbitrary types of time-to-failure distributions, have different performance and wait in different standby modes before their activation. A new numerical algorithm based on state space event transitions is first proposed to evaluate mission success probability of the real-time standby systems considered in this work. Additional new contributions are made by formulating and solving optimal dynamic checkpointing policy problems, as well as an integrated optimization problem that finds the optimal combination of checkpointing policy and element activation sequence maximizing mission success probability. Advantages of using the dynamic checkpointing policy over fixed even checkpoints are demonstrated through examples. Examples and results are also provided to illustrate effects of different mission and element parameters on mission success probability as well as on the optimal dynamic checkpointing policy.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai, Vinod Vokkarane
IEEE Trans. Computers2
2017 Optimal Periodic Inspections and Activation Sequencing Policy in Standby Systems With Condition-Based Mode Transfer
abstract
This paper models a hybrid standby system subject to periodic inspections and condition-based standby mode transfers during a mission. At the beginning of the mission only one element is online and operating. The second element waits in a hot standby mode being ready to replace the failed online element at any time. Other elements wait in less-stressful and less-costly warm standby mode. During the mission periodic inspections are performed for checking conditions of the online and hot standby elements and subsequently triggering necessary mode transfer(s) of available warm standby element(s) to replace the failed hot standby element and/or online element. We suggest an efficient numerical method to assess availability and expected total mission cost (including standby cost, operation cost and mode transfer cost of system elements, inspection cost, system interruption or idle cost) of the considered system. The algorithm is flexible and applicable to arbitrary type of time-to-failure distributions. Then we formulate and solve new optimization problems that identify the optimal combination of inter-inspection interval and element activation sequence to minimize expected total mission cost while satisfying a certain constraint on system availability. As illustrated through examples, the optimization results can facilitate cost-effective and availability-aware planning of system inspection and operation.
Yuan-Shun Dai, Gregory Levitin, Liudong Xing
IEEE Trans. Reliab.3
2017 Preventive Replacements in Real-Time Standby Systems With Periodic Backups
abstract
This paper models a real-time warm standby system that has to accomplish a specified amount of task by a hard deadline. The system is subject to corrective replacements (CRs) upon failure of its operating element. It can also be renewed according to a predetermined schedule through preventive replacements (PRs). To facilitate an effective recovery of system operation after replacements, periodic backups are performed so that warm standby elements, upon being activated, can take over the mission task from the last backup point instead of from scratch. This paper presents a novel integrated model that considers effects of periodic backups, CRs and PRs in analyzing and optimizing real-time warm standby systems. Mission success probability and expected mission completion time are evaluated. Impacts of different mission and element parameters on mission success probability, optimal backup and PR policies, and optimal element activation sequence are investigated. It is shown that in warm standby systems with periodic backups and tight deadlines, PRs can improve the mission success probability even when they take the same time as CRs. When the maximum allowed mission time exceeds a certain level, PRs become ineffective and the optimal policy can involve only periodic backups.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Reliab.2
2017 Optimization of Component Allocation/Distribution and Sequencing in Warm Standby Series-Parallel Systems
abstract
Existing works on redundancy allocation problems have typically focused on active or cold standby redundancies or a mix of them; little research is dedicated to warm standby systems but with an assumption of allocating the same choice of components within each subsystem. Motivated by the fact that components with different costs and failure time distributions from different vendors can be available for the design of the same subsystem in practice, this paper advances the state-of-the-art by presenting a solution methodology to determine combined optimal design configuration and optimal operation of heterogeneous warm standby series-parallel systems. Particularly, based on a proposed numerical reliability evaluation algorithm, two combined optimization problems (component allocation and sequencing problem, and component distribution and sequencing problem) are formulated and solved. Necessity and significance of the proposed methodology are illustrated through examples. Efficiency of the methodology is also successfully demonstrated on large warm standby series-parallel systems containing 14 subsystems of different choices of components.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Reliab.2
2017 Reliability Versus Expected Mission Cost and Uncompleted Work in Heterogeneous Warm Standby Multiphase Systems
abstract
This paper considers 1-out-of-N : G warm standby (WS) systems subject to multiple-phased mission requirements, where performance, failure behavior, operation cost, replacement time, and cost of system elements can vary from phase to phase due to changing working conditions and stress levels. The system succeeds if its elements can accomplish a specified mission task within the maximum allowed time. A numerical algorithm is proposed for evaluating mission indices, including system reliability, expected mission cost, and expected uncompleted work of the heterogeneous, dynamic standby system considered. The influence of the maximum allowed mission time on mission indices is investigated. As demonstrated through examples, the proposed algorithm can also facilitate a determination of mission interruption based on reliability or cost-oriented criteria. Based on the proposed algorithm, an unconstrained optimal standby element sequencing problem is then formulated and solved for heterogeneous WS multiphase systems. The problem is to find the optimal activation sequence of system elements maximizing system reliability, or minimizing expected mission cost, or minimizing expected uncompleted work. The constrained optimization problems of minimizing expected mission cost subject to providing a desired level of mission reliability or uncompleted work are also solved. Illustrative examples of these optimization problems and their applications in estimating the role of each element in the mission success are presented.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Syst. Man Cybern. Syst.2
2017 Optimal Distribution of Nonperiodic Full and Incremental Backups
abstract
Data backups play an important role in effective recovery of system operations, especially, those in computing and Information Technology-related applications. By saving copies of information associated with the completed part of a mission task, a failed system, upon being repaired, can resume its operation from the latest backup point instead of having to repeat the entire work from scratch. This paper considers a repairable, real-time system performing a sequence of nonperiodic full backup (FB) and incremental backup (IB) procedures during its mission. The mission succeeds if a specified amount of work can be accomplished by the system within a deadline. New contributions of this paper are twofold. First, a numerical recursive method is proposed to model effects of the nonperiodic, mixed backup strategy in assessing mission success probability, and expected completion time of the considered system. The method is applicable to any distribution of FBs and IBs and has no limitation on system time-to-failure distributions. Second, the optimal backup policy problem is formulated and solved, which finds the distribution of FBs and IBs maximizing the mission success probability. Examples are provided to illustrate the proposed methodology as well as influence of different parameters on the optimal solution. Advantages of adopting nonperiodic backups over periodic ones are also illustrated.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Syst. Man Cybern. Syst.2
2016 Optimization of Full versus Incremental Periodic Backup Policy
abstract
This paper models repairable computing systems performing a mission that is successful if the system can accomplish a specified amount of work within the allowed mission time or deadline. During the mission the system is subject to a sequence of full and incremental data backup procedures to facilitate an effective system recovery and avoid repeating the entire mission work from the very beginning when a system failure happens. The repair time is fixed while the system time-to-failure can follow any arbitrary type of distributions. This paper makes novel contributions by first developing a new numerical algorithm to evaluate mission success probability and expected completion time of the considered repairable real-time computing systems subject to mixed full and incremental backups. Correctness of the proposed evaluation algorithm is verified using Monte Carlo simulations. We make another new contribution by formulating and solving the backup schedule optimization problem that finds the full and incremental backup frequencies maximizing the mission success probability. Through illustrative examples, effects of different parameters (including the system time-to-failure distribution parameter, maximum allowed mission time, data backup and retrieval times, storage availability, repair time and efficiency) on the mission success probability and expected completion time as well as on the optimal backup schedule solution are investigated.
Gregory Levitin, Liudong Xing, Qingqing Zhai, Yuan-Shun Dai
IEEE Trans. Dependable Secur. Comput.2
2016 Reliability Evaluation of Network Systems with Dependent Propagated Failures Using Decision Diagrams
abstract
In a network system, a propagated failure (PF) is a failure originating from a network component that can cause extensive damages to other network components or even the failure of the entire system. Existing works on PFs have mostly assumed the deterministic effect from a component PF, i.e., a fixed subset of system components is affected whenever the PF occurs. However, in many real-world systems, the components may have different levels of protection, and the effect of damage from a component PF can be dependent upon the status of other components within the same system or the occurrence order of component failures. This paper proposes a new analytical method based on multi-valued decision diagrams (MDDs) for the reliability analysis of network systems with dependent propagation effects. Particularly, new MDD modeling procedures are proposed for considering different types of dependent PF effects introduced by different protection levels. After the system MDD is generated using a new MDD combination algorithm to efficiently handle the dependent PF effects, methods for computing the network reliability and component importance measures are presented. The detailed analysis of an example network system subjected to dependent PFs is presented to illustrate the basics and application of the proposed method. It is shown that the proposed MDD-based method generates smaller model size and thus presents lower computational complexity in the model generation and evaluation than the existing Markov method and separable method.
Yuchang Mo, Liudong Xing, Farong Zhong, Zhao Zhang 0002
IEEE Trans. Dependable Secur. Comput.2
2016 Heterogeneous Non-Repairable Warm Standby Systems With Periodic Inspections
abstract
Motivated by practical applications (for example production systems, flow transmission systems, and transportation systems), this paper considers 1-out-of- N: G warm standby systems subject to periodic inspections. Periodic inspections are performed to detect failures of system components. If an online operating component failure is detected during inspection, an available warm standby component is activated to take over the mission task. The mission succeeds if system components can complete a pre-specified amount of work. A state space event transition based numerical algorithm is first suggested for evaluating important mission performance indices including mission success probability, expected mission cost, expected mission time, expected mission completion delay over a desired time, and expected uncompleted work of the considered standby system. Based on the proposed evaluation algorithm, inspection interval optimization problems are formulated and solved, which find the optimal value of the inspection interval to minimize the expected total mission cost subject to providing a desired level of mission success probability, expected completion time, and delay. Further, combined component sequencing and inspection interval optimization problems are solved, for minimizing the expected total mission cost, or maximizing the mission success probability. As illustrated through examples, the proposed methodology can also facilitate solving dynamic optimization problems that update the optimal solution after each component failure, leading to further improvements of various mission indices.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Reliab.2
2016 Cold Standby Systems With Imperfect Backup
abstract
Cold standby redundancy is a widely-applied design strategy to achieve high system reliability in numerous applications. To facilitate an effective system recovery in case of an online element failure occurring, the backup mechanism is typically implemented which enables a standby element to take over the mission task from a backup point instead of re-executing the entire mission task from the very beginning. However, the backup mechanism is not perfectly reliable in practice, and effect of its failure on the system reliability can be non-monotonic and is correlated to other system and element parameters. In this paper a new numerical approach is first proposed to evaluate reliability and expected mission completion time for 1-out-of- N: G cold standby systems subject to imperfect even backups. The system elements are non-repairable during the mission. Based on the proposed evaluation algorithm, effects of backup system reliability in connection with other system parameters including backup frequency, data backup and retrieval times, replacement failure probability, replacement time, and number of system elements are investigated through examples. Further, the optimal backup frequency and initiation sequencing problem is formulated and solved for heterogeneous cold standby systems, providing solutions that can maximize mission reliability or minimize expected mission completion time depending on design requirements.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Reliab.2
2016 Heterogeneous Warm Standby Multi-Phase Systems With Variable Mission Time
abstract
While the majority of existing works on modeling and optimizing standby systems focus on single-phased missions, only a few works consider standby systems with phased-mission requirements, and these works are only applicable to restricted cold-standby configurations with assumptions of negligible replacement times and fixed phase durations. This paper makes new theoretical contributions by suggesting a general model of heterogeneous 1-out-of- N: G warm standby phased-mission systems (PMS) with dynamic phase durations. The model takes into account diverse, phase-dependent performances and time-to-failure distributions of system elements. It also considers non-zero replacement times specific for activated elements, and for the mission phase when the replacement occurs. Both cold and hot standby PMSs are special cases of the proposed model. An algorithm for evaluating the mission reliability and expected completion time is first presented. The algorithm is applicable to any type of time-to-failure distribution for system elements. The heterogeneous standby PMS can demonstrate non-coherent behavior where the reduction of reliability of some elements can cause the increase of the entire mission reliability. Then it is demonstrated that the sequence of the element activation affects the mission reliability and its expected completion time. Hence the element sequence optimization problem is further formulated and solved for the heterogeneous warm standby PMS. Illustrative examples of mission reliability and expected completion time analysis and optimization are presented.
Gregory Levitin, Liudong Xing, Amir Ehsani Zonouz, Yuan-Shun Dai
IEEE Trans. Reliab.2
2016 A Hierarchical Correlation Model for Evaluating Reliability, Performance, and Power Consumption of a Cloud Service
abstract
Cloud computing is a new emerging technology aimed at large-scale resource sharing and service-oriented computing. To achieve the efficient use of cloud resources for supporting a cloud service, many important factors need to be considered, particularly, reliability, performance, and power consumption of the cloud service. Evaluation of these metrics is essential for further designing rational resource scheduling strategies. However, these metrics are closely related; they do affect one another. The cloud system should consider correlations among the metrics to make more precise evaluation. Most of the existing approaches and models handle these metrics separately, and thus they cannot be used to study the correlations. This paper presents a new hierarchical correlation model for analyzing and evaluating these correlated metrics, which encompasses Markov models, queuing theory, and a Bayesian approach. Various distinctive characteristics of the cloud system are investigated and captured in the model, such as multiple virtual machines (VMs) hosted on the same server, common cause failures of co-located VMs caused by server failures, and logical mapping mechanisms for multicore CPUs. Moreover, for evaluating and balancing the tradeoff between performance and power consumption, a tradeoff parameter and a pure profit optimization model are developed based on the presented correlation model. Numerical examples are provided.
Xiwei Qiu, Yuan-Shun Dai, Yanping Xiang, Liudong Xing
IEEE Trans. Syst. Man Cybern. Syst.4
2015 Mission Reliability, Cost and Time for Cold Standby Computing Systems with Periodic Backup
abstract
Life critical applications like space missions and flight controls require their computing systems to be equipped with some fault-tolerance mechanism to meet stringent reliability requirements by performing the intended function even in the case of element failures. Such benefit, however, cannot come without extra time as well as extra overhead and capital costs. This paper for the first time considers the modeling and evaluation of mission reliability, expected mission time and cost simultaneously for 1-out-of-$N$: G non-repairable cold standby computing systems subject to periodic backup actions. Based on the suggested numerical evaluation method, the optimal backup frequency problems are formulated and solved, providing the optimal number of backup operations during the mission to maximize the system reliability or to minimize the mission cost or time. In the case of non-identical system elements, the optimal standby element sequencing problem arises as the order in which the system elements are initiated can impact the system reliability and mission cost and time greatly; such problems are formulated and solved for the 1-out-of-$N$: G cold standby computing systems with periodic backups. Furthermore, a combined optimization problem is considered, where a combination of the element initiation sequence and backup frequency providing the best combination of mission reliability, cost, and time is found. The proposed methodology can facilitate a reliability-cost-time tradeoff study in the practical design of cold standby systems, thus assist in making the optimal decision on the system's standby and backup policy. Examples are provided for illustrating the considered problems and suggested solution methodology.
Gregory Levitin, Liudong Xing, Barry W. Johnson, Yuan-Shun Dai
IEEE Trans. Computers2
2015 Effect of Failure Propagation on Cold vs. Hot Standby Tradeoff in Heterogeneous 1-Out-of-N: G Systems
abstract
This paper considers 1-out-of- N:G heterogeneous fault-tolerant systems that are designed with a mix of hot and cold standby redundancies to achieve the tradeoff between restoration and operation costs of standby elements. In such systems, the way in which the elements are distributed between hot and cold standby groups and the initiation sequence of all the cold standby elements can greatly affect the system reliability and mission cost. Therefore, it is significant to solve the optimal standby element distributing and sequencing problem (SE-DSP). The failure that occurs in a system element can propagate, causing the outage of other system elements, which complicates the solution to the SE-DSP problem. In this paper, we first propose a numerical method for evaluating the reliability and expected mission cost of 1-out-of- N:G systems with mixed hot and cold redundancy types and propagated failures. Two different failure propagation modes are considered: an element failure causing the outage of all the system elements, and an element failure causing the outage of only working or hot standby elements but not cold standby elements. A genetic algorithm is utilized as an optimization tool for solving the formulated SE-DSP problem, leading to a solution that can minimize the expected mission cost of the system while providing a desired level of the system reliability. Effects of the failure propagation probability on the system reliability, expected mission cost, as well as the optimization results are investigated. The suggested methodology can facilitate a reliability-cost tradeoff study of the considered systems, thus assisting in optimal decision making regarding the system's standby policy. Examples are provided for illustrating the considered problem as well as the proposed solution methodology.
Gregory Levitin, Liudong Xing, Hanoch Ben-Haim, Yuan-Shun Dai
IEEE Trans. Reliab.2
2015 Reliability of Non-Coherent Warm Standby Systems With Reworking
abstract
In this paper we model and analyze non-repairable 1-out-of- N : G warm standby systems subject to periodic backups and dynamic reworking. Particularly, in such systems, a standby element must redo some portion of already performed work by the failed online element before taking over the mission task, which makes the actual mission time dynamic. The considered systems are widely used in applications such as computing and manufacturing, but have not been well studied in reliability theory. In this work, we make new contributions by suggesting a numerical algorithm to evaluate the reliability of the considered warm standby systems. It is revealed that these systems are non-coherent, where the system reliability has non-monotonic dependence on the reliability of individual elements. Numerical examples further show that the non-coherency phenomenon is more distinguished for elements initiated earlier than those initiated later in the warm standby list. Example results also imply that placing highly unreliable elements at the end of the warm standby waiting list, or even removing them from the system planning, can enhance the reliability of a warm standby system subject to reworking. Findings from this work can guide the reliability design of the considered warm standby systems in practice.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Reliab.2
2015 Reliability and Mission Cost of 1-Out-of-N: G Systems With State-Dependent Standby Mode Transfers
abstract
The paper proposes a new fault tolerant system design model, in particular a 1-out-of- N:G hybrid redundant system with standby elements subject to state-dependent standby mode transfers. Specifically, in such systems, one standby element always resides in the hot standby mode, and thus is ready to replace the failed online element at any time to make the system dependable. If the online operating element or the hot standby element fails, one of the warm standby elements is immediately transferred to the hot standby mode. A numerical algorithm is first suggested for evaluating the reliability and the expected mission cost of the considered system. The algorithm is based on a discrete approximation of element time-to-failure distributions, and can work with any type of distribution. Furthermore, based on the suggested algorithm, the problem of optimal sequencing of standby elements initiation is formulated and solved. The objective of this optimization problem is to minimize the expected mission cost associated with elements' standby and operation expenses, as well as the mode transfer expenses, while meeting a certain system reliability constraint. Illustrative examples are provided.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Reliab.2
2015 Non-Homogeneous 1-Out-of-N Warm Standby Systems With Random Replacement Times
abstract
In standby systems, when an online working element fails, a replacement procedure is initiated to activate a standby unit which will take over the mission task to sustain the system function. Existing works on standby systems have mostly assumed that such replacement procedure takes a negligible or fixed amount of time. This assumption is not practical in many real-world systems, where the replacement procedure can take times that are random and different for different standby elements. This paper makes novel contributions by considering the effects of the random replacement times in analyzing and optimizing 1-out-of- N: G non-repairable warm standby systems. The system elements are not necessarily identical; different elements can have different time-to-failure distributions, different performance levels, and different replacement time distributions. The system is considered failed if the elements cannot complete the specified amount of work (mission task) within the maximum allowed mission time. A numerical algorithm is first proposed to simultaneously evaluate the mission reliability and expected mission completion time of the considered warm standby system. Influences of different mission and element parameters on the mission reliability and expected completion time are investigated. It is revealed that the considered warm standby systems exhibit non-coherent behavior as mission reliability may increase with the decrease of an element's reliability. Due to heterogeneity of the system elements, the order in which the elements are initiated and replaced can affect the mission reliability and actual completion time significantly. Therefore, based on the suggested numerical evaluation algorithm, we further formulate and solve the optimal element replacement sequencing problem for the considered warm standby system subject to random replacement times. Examples are given to demonstrate the considered problems and proposed methodology.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Reliab.2
2015 Heterogeneous 1-Out-of-N Warm Standby Systems With Dynamic Uneven Backups
abstract
In this paper, mission reliability, expected mission completion time, and cost of non-repairable 1-out-of- N: G warm standby sparing systems subject to uneven backup actions are modeled and optimized. The backup actions are used to facilitate the data recovery process in the case of an online operating element failure, which enables an activated standby element to take over the mission task through subsequent data retrievals. Both data backup and retrieval times are dynamic, and physically dependent on the amount of work performed. The system elements are not necessarily identical; each element can be characterized by a different time-to-failure distribution, a different performance, and a different level of readiness to take over the system task during the warm standby mode. An iterative numerical method is first proposed to simultaneously evaluate mission reliability, expected mission completion time, and the cost of the considered heterogeneous warm standby systems. Due to the non-monotonic effect of the backup distribution on the mission reliability, time, and cost, we formulate and solve the optimal backup distribution problem considering different combinations of optimization objectives and constraints. In the case of system elements being non-identical, their activation order can influence the mission reliability, expected mission completion time, and mission cost significantly. Therefore, we also formulate and solve the optimal element sequencing problem for the considered system. Furthermore, new integrated optimization problems are formulated and addressed. The integrated optimization aims to identify the optimal combination of backup distribution and element activation order that maximizes the mission reliability, or minimizes the expected mission time or mission cost. As shown through examples, the proposed methodology can implement a tradeoff analysis among the three mission requirements of reliability, cost, and completion time, leading to the optimal decision on both backup and standby policies of warm standby systems.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Reliab.2
2015 Multi-Valued Decision Diagram-Based Reliability Analysis of k-out-of-n Cold Standby Systems Subject to Scheduled Backups
abstract
To improve the system reliability while conserving the limited system resources, cold standby sparing is often used. In computing tasks, because active components fail randomly, and the standby component has to pick up the mission task whenever required, scheduled backups are often implemented to save the completed portions of the task. The backups can facilitate an effective system recovery where the standby component can take over the mission task from the last backup point instead of resuming the mission task from the very beginning. This paper considers a k-out-of- n cold standby system subject to scheduled backups, where k components are online and operating, with the remaining components waiting in the unpowered, cold standby mode. Whenever an online component fails, a cold standby component is activated to take over the mission task from the last backup point. The backup intervals are deterministic, but can be even or uneven. As the component may fail due to an imperfect switching from the standby state to the fully powered up state, the switching failure is also considered in the system model. A multi-valued decision diagram (MDD)-based analytical approach is proposed to evaluate the reliability of the considered system, and its complexity is analyzed. The proposed method is applicable to systems with non-identical components following arbitrary lifetime distributions. Examples are given to illustrate the MDD-based method. The correctness and efficiency of the proposed method are verified using Monte Carlo simulations.
Qingqing Zhai, Liudong Xing, Rui Peng 0001, Jun Yang 0018
IEEE Trans. Reliab.2
2015 Optimal Backup Distribution in 1-out-of-N Cold Standby Systems
abstract
This paper considers nonrepairable 1-out-of-N: G cold standby (CS) systems subject to uneven backup actions as well as dynamic backup and retrieval times. In such systems, only one element is online and operates with the rest of the elements waiting in the unpowered CS mode. The operating element performs data backup actions when certain fractions of the mission task are accomplished. The backup actions facilitate data recovery in case of an operating element failure, which allows an activated standby element to take over the task through data retrieval. Backup distribution can have a nonmonotonic effect on mission reliability, time, and cost, leading to the optimal backup distribution problem. In this paper, we first suggest a numerical method to model and evaluate mission reliability, expected time, and cost simultaneously for the considered CS systems with uneven backup actions and dynamic backup and retrieval times. Based on the suggested evaluation method and genetic algorithm, the optimal backup distribution problem is then formulated and solved with the objective to minimize the expected mission cost subject to meeting certain levels of mission reliability and expected mission time. Examples show that the proposed methodology can facilitate a tradeoff study between mission reliability, and time and cost, which assists in the optimal decisionmaking for the backup policy used in the practical design of CS systems.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Syst. Man Cybern. Syst.2
2015 Optimal Design of Hybrid Redundant Systems With Delayed Failure-Driven Standby Mode Transfer
abstract
Standby redundancy is a design technique that has been widely adopted to enhance system reliability and achieve fault tolerance in many critical applications. In this paper, we consider a 1-out-of-N: G hybrid standby redundant system with unrepairable elements being subject to delayed failure-driven standby mode transfers. Specifically, in the considered system, all the standby elements are initially in a warm standby mode (WSM) but can be transferred to a hot standby mode (HSM) so as to be ready to replace the online operating element when it fails. The WSM to HSM transfer is performed with a fixed time delay after no element resides in HSM either due to the element failure or because the element leaves the HSM to replace the failed online element for operation. A new iterative numerical algorithm is first proposed for evaluating reliability and expected mission cost (relevant to elements' standby expense, operation expense, as well as mode transfer expense) of the considered hybrid standby system. The algorithm has no restriction on element time-to-failure distribution types. Based on the proposed evaluation algorithm, we further formulate and solve a new optimization problem that finds the optimal delay and optimal sequence of standby elements with the objective to minimize expected mission cost while satisfying a certain level of mission reliability constraint. Examples are presented to demonstrate applications of the proposed methodology.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Syst. Man Cybern. Syst.2
2015 A joint resource allocation-channel coding design based on distributed source coding
abstract
Wireless sensor networks WSNs have found a wide variety of applications recently. However, the challenges in WSNs still remain in improving the sensor energy efficiency and information quality distortion reduction of the sensing data transmissions. In this paper, we propose a novel cross-layer design of resource allocation and channel coding to protect distributed source coding DSC-based data transmission. Resource allocation strategies include rate adaptation and automatic repeat-request retransmissions. Our proposed joint design of resource allocation, channel coding, and DSC can improve the network energy efficiency and information quality while meeting the data transmission latency requirements. Further, we investigate how the resource allocation enables the network to achieve unequal error protection among correlated DSC streams. Our simulation studies demonstrate that the proposed joint design significantly improves the DSC-based data transmission quality and the network energy efficiency. Copyright © 2013 John Wiley & Sons, Ltd.
Sasan Khoshroo, Honggang Wang 0001, Liudong Xing, Dayalan Kasilingam
Wirel. Commun. Mob. Comput.3
2014 Trust-aware privacy evaluation in online social networks
abstract
While personal data privacy is threatened by online social networks, researchers are seeking for privacy protection tools and methods to assist online social network providers and users. In this paper, we aim to address this problem by investigating how to quantitatively evaluate the privacy risk, as a function of people's awareness of privacy risks as well as whether their friends can be trusted to protect their personal data. We present a trust-aware privacy evaluation framework, called TAPE. Simulations are performed to illustrate the key concepts and calculations in TAPE, as well as demonstrate the advantages of TAPE.
Yongbo Zeng, Yan Lindsay Sun, Liudong Xing, Vinod Vokkarane
ICC3
2014 Minimum Mission Cost Cold-Standby Sequencing in Non-Repairable Multi-Phase Systems
abstract
This paper considers the optimal cold standby element sequencing problem (SESP) for 1-out-of- n: G heterogeneous non-repairable cold-standby systems that accomplish multi-phase missions. Given a fixed set of element choices, the objective of the optimal system design is to select the initiation sequence of the system elements so as to minimize the expected mission cost while providing a desired level of system reliability. It is assumed that during different mission phases the elements are exposed to different stresses, which affects their time-to-failure distributions. The startup and exploitation costs of system elements are also phase dependent. We suggest an algorithm for evaluating the mission reliability and expected mission cost based on a discrete approximation of time-to-failure distributions of the system elements. A genetic algorithm is used as an optimization tool for solving the formulated SESP for multi-phase cold-standby systems. Examples are given to illustrate the considered problem and the proposed solution methodology.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Reliab.2
2014 A Multiple-Valued Decision Diagram Based Method for Efficient Reliability Analysis of Non-Repairable Phased-Mission Systems
abstract
Many practical systems are phased-mission systems (PMSs), where the mission consists of multiple, consecutive, and non-overlapping phases of operation. An accurate reliability analysis of a PMS must consider statistical dependence of component states across phases, as well as dynamics in system configurations, success criteria, and component behavior. This paper proposes a new method based on multiple-valued decision diagrams (MDDs) for the reliability analysis of a non-repairable binary-state PMS. Due to its multi-valued logic nature, the MDD model has recently been applied to the reliability analysis of multistate systems. In this work, we present a novel way to adapt MDDs for the reliability analysis of systems with multiple phases. Examples show how the MDD models are generated and evaluated to obtain the mission reliability measures. Performance of the MDD-based method is compared with an existing binary decision diagram (BDD)-based method for PMS analysis. Empirical results show that the MDD-based method can offer lower computational complexity as well as a simpler model construction and improved evaluation algorithms over those used in the BDD-based method.
Yuchang Mo, Liudong Xing, Suprasad V. Amari
IEEE Trans. Reliab.2
2014 Combinatorial Reliability Analysis of Imperfect Coverage Systems Subject to Functional Dependence
abstract
Functional dependence occurs when the failure of one component causes other components within the same system to become inaccessible or unusable. It is one of the dynamic behaviors that have been recognized in the dynamic fault tree analysis, where a dynamic gate called FDEP was designed to model such behavior. Traditional approaches to handling functional dependence in the reliability analysis of fault-tolerant systems with imperfect fault coverage are mainly based on Markov models, which are often computationally intensive, and even intractable due to the well-known state space explosion problem. In addition, the Markov-based approaches are typically restricted to exponential time-to-failure distributions for system components. In this paper, a combinatorial, separable method based on the divide-and-conquer paradigm and total probability theorem is proposed for addressing the above problems. The proposed method obviates the use of inefficient Markov models, offering exact, computationally-efficient solutions to the reliability analysis of imperfect coverage systems subject to functional dependencies. The proposed method is applicable to the analysis of large systems with any arbitrary time-to-failure distributions. Several case studies are given to illustrate the application and advantages of the proposed method.
Liudong Xing, Brock A. Morrissette, Joanne Bechta Dugan
IEEE Trans. Reliab.1
2014 Structure Optimization of Nonrepairable Phased Mission Systems
abstract
System structure optimization is a well-studied problem in the field of reliability engineering, which aims to achieve the best possible reliability versus cost solutions for the system design. This problem is usually solved for systems that do not change their task and configuration during the mission. However, many practical systems are phased-mission systems (PMS), where the mission involves multiple, consecutive, and nonoverlapping phases of operation. An accurate analysis of PMS must consider the dynamics in system configuration, success criteria, and element behavior as well as statistical dependence of element states across phases. In this paper, we propose a method for solving the structure optimization problem of multistate PMS consisting of nonidentical nonrepairable binary elements. The system configuration and demand as well as the failure distributions of system elements can change from phase to phase. The proposed approach is based on a recursive algorithm for the reliability evaluation of PMS and a genetic algorithm for the structure optimization of PMS. The method is illustrated using an example of a six-phased mission running on an airborne distributed computing system.
Yuan-Shun Dai, Gregory Levitin, Liudong Xing
IEEE Trans. Syst. Man Cybern. Syst.3
2014 Mission Cost and Reliability of 1-out-of- $N$ Warm Standby Systems With Imperfect Switching Mechanisms
abstract
In this paper, mission cost and reliability of 1-out-of-N: G nonrepairable warm standby systems with imperfect switching are modeled and analyzed using an iterative method. A general switching structure is considered, which consists of an overall fault detection mechanism (FDM) and a set of individual switches (one for each standby element). The failure of the FDM prevents the replacement of the failed online element by any standby element while the failure of an individual switch only makes the corresponding standby element unavailable. The entire mission fails either when the FDM fails before the failure of the online element during the mission, or when all the system elements have failed or become unavailable before the mission completion. Based on the proposed algorithm for the mission cost and reliability analysis, the optimal element sequencing problem is further formulated and solved for 1-out-of-N: G nonrepairable warm standby systems with nonidentical elements and imperfect switching mechanisms. The objective of the problem is to find the optimal initiation sequence of system elements that can minimize the expected mission cost while providing a certain level of system reliability. Examples are given to illustrate the considered problem and the proposed solution methodology.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Syst. Man Cybern. Syst.2
2014 MDD-Based Method for Efficient Analysis on Phased-Mission Systems With Multimode Failures
abstract
Many practical systems are phased-mission systems with multimode failures (MFPMSs) where the mission consists of multiple nonoverlapping phases of operation, and the system components may assume more than one failure mode. In MFPMSs, dependence arises among different phases and among different failure modes of the same component, which makes the reliability analysis of MFPMSs difficult. This paper proposes a new analytical method based on multivalued decision diagrams (MDDs) for the reliability analysis of nonrepairable MFPMSs. MDDs have recently been applied to the reliability analysis of single-phase systems with multiple component states. In this paper, we make the new contribution by proposing a novel way to adapt MDDs for the reliability analysis of systems with multiple phases and multimode failures. Examples show how the MDD models are generated and evaluated to obtain the mission reliability measures. Performance of the MDD-based method is compared with an existing binary decision diagram (BDD)-based method for MFPMS analysis through several examples and a comprehensive benchmark study. Empirical results show that the proposed MDD-based method can offer lower computational complexity and simpler model construction and evaluation algorithms than the BDD-based method, and it can be effectively applied to large practical cases.
Yuchang Mo, Liudong Xing, Joanne Bechta Dugan
IEEE Trans. Syst. Man Cybern. Syst.2
2013 Application Communication Reliability of Wireless Sensor Networks Supporting K-coverage
abstract
Application communication in wireless sensor networks (WSN) depends on two important factors: acquisition of sensed data from a specific area, and network connectivity that concerns the reliable delivery of the observed data from sensor nodes to the sink node. In this paper, we consider the application communication reliability (ACR) of WSN supporting K-coverage in the presence of shadowing for a specific monitored area. The analytical evaluation of ACR involves two steps. We first identify all the K-coverage sets. Then, we evaluate the communication reliability of delivering the observed data from sensor nodes within the identified K-coverage sets to the sink node. Two single-path routing algorithms, shortest-path distance algorithm and shortest-path hop algorithm, are considered for evaluating the communication reliability during the second step; their performances in terms of ACR and energy consumptions are compared through an empirical analysis of several examples. Different scenarios are considered to evaluate the impact of node density, channel condition and different monitored areas on ACR. Simulation and analytical results show that WSN using the shortest-path distance algorithm is more reliable than that using the shortest-path hop algorithm in most cases, but WSN using the shortest-path hop algorithm consumes less energy for delivering the sensed data to the sink node.
Amir Ehsani Zonouz, Liudong Xing, Vinod Vokkarane, Yan Lindsay Sun
DCOSS2
2013 Editorial of special section on advanced in high performance, algorithm, and framework for future computing
Taeshik Shon, Shiuh-Jeng Wang, Lei Shu 0001, Liudong Xing
J. Supercomput.4
2013 Reliability of Series-Parallel Systems With Random Failure Propagation Time
abstract
This paper presents an algorithm for evaluating the reliability and performance distribution of complex non-repairable series-parallel multi-state systems with common cause failures caused by propagation of failures in system elements. The failure propagation can have a selective effect, which means that the failures originating from different elements can cause failures of different subsets of system elements. The failure propagation time is assumed to be a random value with a given distribution. The suggested algorithm is based on the universal generating function approach, and a generalized reliability block diagram method (recursive aggregation of pairs of elements and their replacement by an equivalent one). The evaluation procedure is repeated for each combination of elements affected by the common cause failures. Illustrative examples are provided.
Gregory Levitin, Liudong Xing, Hanoch Ben-Haim, Yuan-Shun Dai
IEEE Trans. Reliab.2
2013 Optimal Allocation of Connecting Elements in Phase Mission Linear Consecutively-Connected Systems
abstract
Many continuous transportation systems and communication networks can be modeled as a linear consecutively connected system (LCCS) consisting of n+1 linearly ordered nodes. Some of these nodes contain connection elements (CEs) with different characteristics. Each CE i in its working state provides a connection between the node to which it belongs and L(i) next nodes, the set of nodes that follow the given node. If the system contains any node not connected with any previous node, then it fails. We consider multi-phase LCCS that should perform a sequence of transmission tasks. In each phase, it should provide continuous connection along a specific path of nodes. Different paths can contain the same nodes, which creates statistical dependence across the phases. Different nodes are characterized by specific conditions expressed by different failure acceleration factors affecting the CEs located at these nodes. Thus, an accurate reliability analysis of a multi-phase LCCS must consider the statistical dependencies of CE states across phases, and dynamics in path structures as well as in the failure acceleration factors. In this paper, we propose a method for optimal allocation of CEs in multi-phase LCCS. The method is based on a recursive algorithm for exact system reliability evaluation, and the genetic algorithm for the optimal allocation of CEs to nodes of LCCS. The proposed approach is illustrated using a practical example of wireless sensor networks.
Gregory Levitin, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Reliab.2
2013 Algorithm for Reliability Evaluation of Nonrepairable Phased-Mission Systems Consisting of Gradually Deteriorating Multistate Elements
abstract
Reliability analysis of phased-mission systems (PMS) must consider the statistical dependences of element states across different phases as well as changes in system configuration, success criteria, and component behavior. This paper proposes a recursive method for the exact reliability evaluation of PMS consisting of nonidentical independent nonrepairable multistate elements. The method is based on conditional probabilities and the branch-and-bound principle. It is invariant to changes in system structure, demand, and the elements' state transition rates among phases. The main advantage of this method is that it does not require the composition of decision diagrams and can be fully automated. Both analytical and numerical examples are presented to illustrate the application and advantages of the proposed method. The computational performance of the proposed algorithm is illustrated through comprehensive experimentation on the CPU running time of the algorithm.
Gregory Levitin, Suprasad V. Amari, Liudong Xing
IEEE Trans. Syst. Man Cybern. Syst.3
2013 Reliability of Nonrepairable Phased-Mission Systems With Common Cause Failures
abstract
Phased-mission systems (PMSs) are systems supporting missions characterized by multiple, consecutive, and nonoverlapping phases of operation. Examples of PMSs abound in many practical applications such as aerospace, nuclear power, and airborne weapon systems. Reliability analysis of a PMS must consider statistical dependence of component states across different phases, as well as dynamics in system structure functions and component behavior. In this paper, we propose a recursive method for exact reliability evaluation of a binary-state or multistate PMS consisting of nonidentical, binary, and nonrepairable elements. The system elements can fail individually or due to common-cause failures (CCFs) caused by some external factors. The proposed method is based on the branch-and-bound principle, and can be fully automated. The method is applicable to PMSs with nonoverlapping or overlapping sets of elements that can fail as a result of CCFs. The method is illustrated using both analytical and numerical examples.
Gregory Levitin, Liudong Xing, Suprasad V. Amari, Yuan-Shun Dai
IEEE Trans. Syst. Man Cybern. Syst.2
2013 Reliability of Systems Subject to Failures With Dependent Propagation Effect
abstract
This paper suggests a method for the reliability analysis of binary-state systems subject to component failures with dependent propagation effect. A propagated failure originating from a system component causes extensive damage to the rest of the system. The level of the damage can be dependent upon the status of other system components and the order of the component failures. Such dependent propagation effect typically takes place in systems with some protection mechanism (e.g., firewalls, filters, and antivirus programs) or systems subject to functional dependence behavior where the failure of a system component, referred to as a trigger, causes other components within the same system to become inaccessible or isolated from the system. A combinatorial and analytical method is proposed for addressing the dependent propagation effect in the system reliability analysis. Basics and application of the proposed method are illustrated through analyses of an example of a network system in two different scenarios for propagation effects and an example of a memory system with multiple trigger events.
Liudong Xing, Gregory Levitin, Chaonan Wang, Yuan-Shun Dai
IEEE Trans. Syst. Man Cybern. Syst.1
2012 Design and evaluation of small-world wireless ad-hoc networks under rayleigh fading
abstract
Small-world phenomenon is an important property of many complex networks possessing small average shortest path lengths and high clustering coefficients. On the other side, wireless ad-hoc networks are highly clustered having large average shortest path length caused by the locality of their physical connections. Since the average shortest path length (number of hops) traversed by data packets has a tight interplay with the level of quality of service experienced in the network, efforts have recently focused on the creation of small-world property in wireless networks. In this study, the concept of small world is investigated in the context of wireless ad-hoc networks in a Rayleigh fading environment. In particular, we prove that the extension of few wireless links to farther located nodes (using the concept of rewiring) can reduce the average shortest path length to a decent extent, even in the presence of severe channel fading. Simulation results show that a reduction of up to 23.3% in clustering coefficient and up to 25.6% in average shortest path length can be achieved for a low-density network in the presence of fading.
Amir Ehsani Zonouz, Navid Tadayon, Sonia Aïssa, Liudong Xing
GLOBECOM4
2012 Recursive Algorithm for Reliability Evaluation of Non-Repairable Phased Mission Systems With Binary Elements
abstract
Many practical systems are phased-mission systems (PMS), where the mission consists of multiple, consecutive, and non-overlapping phases of operation. An accurate reliability analysis of a PMS must consider the statistical dependencies of component states across phases, as well as dynamics in system configurations, success criteria, and component behavior. In this paper, we propose a method for exact reliability evaluation of arbitrary binary or multi-state PMS consisting of non-identical binary non-repairable elements. The method is invariant to changes in system structure and demand among missions, and takes into account the time-varying and phase-dependent failure rates and associated cumulative damage effects. The proposed method is based on conditional probabilities, and an efficient recursive formula to compute these probabilities based on branch and bound. The main advantage of this method is that it does not require composition of decision diagrams, and can be fully automated. The method is illustrated using both an analytical example, and a numerical example.
Gregory Levitin, Liudong Xing, Suprasad V. Amari
IEEE Trans. Reliab.2
2012 Reliability Analysis of Nonrepairable Cold-Standby Systems Using Sequential Binary Decision Diagrams
abstract
Many real-world systems, particularly those with limited power resources, are designed with cold-standby redundancy for achieving fault tolerance and high reliability. Cold-standby units are unpowered and, thus, do not consume any power until needed to replace a faulty online component. Cold-standby redundancy creates sequential dependence between the online component and standby components; in particular, a standby component can start to work and then fail only after the online component has failed. Traditional approaches to handling the cold-standby redundancy are typically state-space-based or simulation-based or inclusion/exclusion-based methods. Those methods, however, have the state-space explosion problem and/or require long computation time particularly when results with a high degree of accuracy are desired. In this paper, we propose an analytical method based on sequential binary decision diagrams (SBDD) for combinatorial reliability analysis of nonrepairable cold-standby systems. Different from the simulation-based methods, the proposed approach can generate exact system reliability results. In addition, the system SBDD model and reliability evaluation expression, once generated, are reusable for the reliability analysis with different component failure parameters. The approach has no limitation on the type of time-to-failure distributions for the system components or on the system structure. Application and advantages of the proposed approach are illustrated through several case studies.
Liudong Xing, Ola Tannous, Joanne Bechta Dugan
IEEE Trans. Syst. Man Cybern. Part A1
2011 An Integrated Biometric-Based Security Framework Using Wavelet-Domain HMM in Wireless Body Area Networks (WBAN)
abstract
In this paper, we proposed an integrated biometric-based security framework for wireless body area networks, which takes advantage of biometric features shared by body sensors deployed at different positions of a person's body. The data communications among these sensors are secured via the proposed authentication and selective encryption schemes that only require low computational power and less resources (e.g., battery and bandwidth). Specifically, a wavelet-domain Hidden Markov Model (HMM) classification is utilized by considering the non-Gaussian statistics of ECG signals for accurate authentication. In addition, the biometric information such as ECG parameters is selected as the biometric key for the encryption in the framework. Our experimental results demonstrated that the proposed approach can achieve more accurate authentication performance without extra requirements of key distribution and strict time synchronization.
Honggang Wang 0001, Hua Fang 0001, Liudong Xing, Min Chen 0003
ICC3
2011 Consequence Oriented Self-Healing and Autonomous Diagnosis for Highly Reliable Systems and Software
abstract
Computing software and systems have become increasingly large and complex. As their dependability and autonomy are of great concern, self-healing is an ongoing challenge. This paper presents an innovative model and technology to realize the self-healing function under the real-time requirement. The proposed approach, different from existing technologies, is based on a new concept defined as consequence-oriented diagnosis and healing. Derived from the new concept, a prototype model for proactive self-healing actions is presented. Then, a hybrid diagnosis tool is proposed that takes advantages from the Multivariate Decision Diagram, Fuzzy Logic, and Neural Networks, achieving an efficient, effective, accurate, and intelligent result. The consequence-oriented diagnosis and self-healing function is also implemented. The experimental results exhibit that the innovative system is very effective and precise in predicting the consequence, and in preventing resulting software and system failures.
Yuan-Shun Dai, Yanping Xiang, Yan-Fu Li, Liudong Xing, Gewei Zhang
IEEE Trans. Reliab.4
2011 Reliability Analysis of Multistate Phased-Mission Systems With Unordered and Ordered States
abstract
Multistate phased-mission systems (MS-PMS) are multistate systems subject to multiple, consecutive, and nonoverlapping phases of operation. The challenges in analyzing MS-PMS reside in the dynamic system configuration, failure criteria, and component state transition behavior in different phases, as well as thes-dependence across different phases and among different states of a given component. Existing methods for the reliability analysis of MS-PMS are either based on monolithic Markov models that suffer from the well-known state explosion problem, or using a hierarchical strategy that can only handle ordered component states. This paper presents integrated modeling approaches for the reliability analysis of repairable MS-PMS with both ordered and unordered component states. The proposed methods integrate efficient decision diagram models for representing the system structure function and incorporating the unordered/ordered component states at the system level, and Markov models for describing dependence and transition behaviors at the component level. The application and advantages of the proposed approaches are illustrated through a case study in which the reliability for a sequence of tasks in a multistate distributed computing system is analyzed.
Akhilesh Shrestha, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Syst. Man Cybern. Part A2
2011 Combinatorial Algorithm for Reliability Analysis of Multistate Systems With Propagated Failures and Failure Isolation Effect
abstract
This paper considers the reliability analysis of multistate systems (MSSs) subject to propagated failure with global effect (PFGE) and failure isolation effect. The PFGE can be caused by an imperfect fault coverage despite the presence of fault-tolerant mechanism or by a destructive effect of failures that originate from some system components on other components. The failure isolation effect is caused by functional dependence among system components, where the failure of some component can prevent the propagation of failures that originate from other components within the same system. Existing approaches for simultaneously addressing PFGE and failure isolation are limited to binary-state systems in which the system and its components exhibit two and only two states: operation or failure. In practice, however, many systems are MSS in which the system and/or its components may exhibit multiple performance levels corresponding to different states ranging from perfect operation to complete failure. In this paper, a separable and combinatorial methodology is proposed for evaluating the reliability of MSS subject to both PFGE and the failure isolation effect. The proposed method has no limitation on the type of time-to-failure distributions for the system components and is applicable to MSS with any arbitrary system structure. Application and advantages of the proposed method are illustrated through a detailed analysis of an example of a multistate memory system.
Liudong Xing, Gregory Levitin
IEEE Trans. Syst. Man Cybern. Part A1
2010 Performability Analysis of Multistate Computing Systems Using Multivalued Decision Diagrams
abstract
A distinct characteristic of multistate systems (MSS) is that the systems and/or their components may exhibit multiple performance levels (or states) varying from perfect operation to complete failure. MSS can model behaviors such as shared loads, performance degradation, imperfect fault coverage, standby redundancy, limited repair resources, and limited link capacities. The nonbinary state property of MSS and their components as well as dependencies existing among different states of the same component make the analysis of MSS difficult. This paper proposes efficient algorithms for analyzing MSS using multivalued decision diagrams (MDD). Various reliability, availability, and performability measures based on state probabilities or failure frequencies are considered. The application and advantages of the proposed algorithms are demonstrated through two examples. Furthermore, experimental results on a set of benchmark examples are presented to illustrate the advantages of the proposed MDD-based method for the performability analysis of MSS, as compared to the existing methods.
Suprasad V. Amari, Liudong Xing, Akhilesh Shrestha, Jennifer Akers, Kishor S. Trivedi
IEEE Trans. Computers2
2010 An Efficient Multistate Multivalued Decision Diagram-Based Approach for Multistate System Sensitivity Analysis
abstract
Multistate systems (MSS) are systems in which the system and its components are characterized by multiple states or performance levels. Component importance or sensitivity analysis facilitates the identification of vulnerabilities within the system, and aids in the quantification of criticalities of the system components. Multistate component importance analysis poses unique challenges to existing methods that are primarily based on binary-state applications. This paper presents an analytical method based on multistate multivalued decision diagrams (MMDD) for multistate component importance analysis. The contribution of this work is two-fold: 1) a novel, efficient algorithm for directly generating an MMDD model from multistate capacity network specifications without inefficient enumeration of multistate minimal path or cut vectors; and 2) an efficient, exact MMDD-based approach for evaluating MSS reliability and importance measures. The advantages of the proposed method are illustrated through a comparison with existing methods, and through detailed analyses of three case studies.
Akhilesh Shrestha, Liudong Xing, David W. Coit
IEEE Trans. Reliab.2
2010 Decision Diagram Based Methods and Complexity Analysis for Multi-State Systems
abstract
Decision diagrams are graphical structures based on Shannon's decomposition. They have been extensively used for representing and manipulating logic functions in areas such as circuit verification, compact Markov chain representation, and symbolic model checking. However, their applicability in reliability modeling and analysis has only been recently studied. Moreover, the study had been mostly restricted to binary-state systems in which both the system and its components are either operational, or failed. Nevertheless, many practical systems are multi-state systems (MSS) in which both the system and its components may reside at multiple (more than two) performance levels (or states), varying from perfect operation to complete failure. This paper presents three forms of decision diagrams for the modeling and analysis of MSS: binary decision diagrams, logarithmically encoded binary decision diagrams, and multi-valued decision diagrams. We present both separated, and shared methods based on these decision diagrams. Comprehensive complexity analysis, and performance comparisons among these methods, are conducted with both mathematical, and empirical approaches.
Akhilesh Shrestha, Liudong Xing, Yuan-Shun Dai
IEEE Trans. Reliab.2
2010 Automated Modeling of Dynamic Reliability Block Diagrams Using Colored Petri Nets
abstract
Computer system reliability is conventionally modeled and analyzed using techniques such as fault tree analysis and reliability block diagrams (RBDs), which provide static representations of system reliability properties. A recent extension to RBDs, called dynamic RBDs (DRBD), defines a framework for modeling the dynamic reliability behavior of computer-based systems. However, analyzing a DRBD model in order to locate and identify design errors, such as a deadlock error or faulty state, is not trivial when done manually. A feasible approach to verifying it is to develop its formal model and then analyze it using programmatic methods. In this paper, we first define a reliability markup language that can be used to formally describe DRBD models. Then, we present an algorithm that automatically converts a DRBD model into a colored Petri net. We use a case study to illustrate the effectiveness of our approach and demonstrate how system properties of a DRBD model can be verified using an existing Petri net tool. Our formal modeling approach is compositional; thus, it provides a potential solution to automated verification of DRBD models.
Ryan Robidoux, Haiping Xu, Liudong Xing, MengChu Zhou
IEEE Trans. Syst. Man Cybern. Part A3
2009 A Data Transmission Mechanism for Survivable Sensor Networks
abstract
In wireless sensor networks (WSN), packets being sent over the transmission channel could get lost or corrupted due to factors such as channel noises, interferences, and node failures. Using the traditional retransmission mechanism to recover the lost/erroneous packets is costly or even impossible for WSN because sensor nodes are severely resource constrained and data transmission is the most energy intensive activity in WSN. An alternative approach is to apply forward error correcting (FEC) codes to achieve reliable data transmission. Nevertheless, FEC mechanisms do not provide security. To overcome this problem, an integrated mechanism has been proposed that systematically combines Reed-Solomon FEC codes and multiple versions of cryptographic algorithms to achieve both reliable and secure data transmission, leading to survivable WSN. In this paper, queuing modeling is applied to evaluate the performance of the integrated data transmission mechanism, and several case studies are performed to illustrate the application of the mechanism.
Ruiping Ma, Liudong Xing, Tongdan Jin, Tailiang Song
NAS2
2009 A New Decision-Diagram-Based Method for Efficient Analysis on Multistate Systems
abstract
Multistate systems can model many practical systems in a wide range of real applications. A distinct characteristic of these systems is that the systems and their components may assume more than two levels of performance (or states), varying from perfect operation to complete failure. The nonbinary property of multistate systems and their components makes the analysis of multistate systems difficult. This paper proposes a new decision-diagram-based method, called multistate multivalued decision diagrams (MMDD), for the analysis of multistate systems with multistate components. Examples show how the MMDD models are generated and evaluated to obtain the system-state probabilities. The MMDD method is compared with the existing binary decision diagram (BDD)-based method. Empirical results show that the MMDD method can offer less computational complexity and simpler model evaluation algorithm than the BDD-based method.
Liudong Xing, Yuan-Shun Dai
IEEE Trans. Dependable Secur. Comput.1
2009 Incorporating Common-Cause Failures Into the Modular Hierarchical Systems Analysis
abstract
This paper considers the problem of evaluating the reliability of hierarchical systems subject to common-cause failures (CCF); and dynamic failure behavior such as spares, functional dependence, priority dependence, and dependence caused by multi-phased operations. We present a separable solution that has low computational complexity, and which is easy to integrate into existing analytical methods. The resulting approach is applicable to Markov analyses, and combinatorial models for the modular analysis of the system reliability. We illustrate the approach, and the advantages of the proposed approach, through the detailed analyses of two examples of dynamic hierarchical systems subject to CCF.
Liudong Xing, Akhilesh Shrestha, Leila Meshkat, Wendai Wang
IEEE Trans. Reliab.1
2008 A Logarithmic Binary Decision Diagram-Based Method for Multistate System Analysis
abstract
Multistate systems (MSS) are systems in which both the systems, and/or their components may exhibit multiple performance levels or states. MSS can model complex behaviors such as shared loads, performance degradation, imperfect fault coverage, standby redundancy, and limited repair resources. The non-binary state property of MSS, and their components makes the analysis of MSS challenging. In this paper, we propose efficient logarithmically-encoded binary decision diagram (LBDD)-based methods for analysing MSS. The application and advantages of the proposed LBDD-based approaches, as compared to the existing binary decision diagram-based approaches, are demonstrated through the analyses of practical MSS examples, and a set of benchmark examples.
Akhilesh Shrestha, Liudong Xing
IEEE Trans. Reliab.2
2008 An Efficient Binary-Decision-Diagram-Based Approach for Network Reliability and Sensitivity Analysis
abstract
Reliability and sensitivity analysis is a key component in the design, tuning, and maintenance of network systems. Tremendous research efforts have been expended in this area, but two practical issues, namely, imperfect coverage (IPC) and common-cause failures (CCF), have generally been missed or have not been fully considered in existing methods. In this paper, an efficient approach for fully incorporating both IPC and CCF into network reliability and sensitivity analysis is proposed. The challenges are to allow multiple failure modes introduced by IPC and to cope with multiple dependent faults caused by CCF simultaneously in the analysis. Our methodology for addressing the aforementioned challenges is to separate the consideration of both IPC and CCF from the combinatorics of the solution, which is based on reduced ordered binary decision diagrams (ROBDD). Due to the nature of the ROBDD and the separation of IPC and CCF from the solution combinatorics, our approach has a low computational complexity and is easy to implement. A sample network system is analyzed to illustrate the basics and advantages of our approach. A software tool that we developed for fault-tolerant network reliability and sensitivity analysis is also presented.
Liudong Xing
IEEE Trans. Syst. Man Cybern. Part A1
2007 Efficient Analysis of Systems with Multiple States
abstract
A multistate system is a system in which both the system and its components may exhibit multiple performance levels (or states) varying from perfect operation to complete failure. Examples abound in real applications such as communication networks and computer systems. Analyzing the probability of the system being in each state is essential to the design and tuning of dependable multistate systems. The difficulty in analysis arises from the non-binary state property of the system and its components as well as dependence among those multiple states. This paper proposes a new model called multistate multivalued decision diagrams (MMDD) for the analysis of multistate systems with multistate components. The computational complexity of the MMDD-based approach is low due to the nature of the decision diagrams. An example is analyzed to illustrate the application and advantages of the approach.
Liudong Xing
AINA1
2007 Incorporating Modular Imperfect Coverage into Dynamic Hierarchical Systems Analysis
abstract
This paper deals with the evaluation of the reliability of a dynamic hierarchical system. Reliability of the system is calculated precisely by incorporating modular imperfect coverage model for both static and dynamic subsystems. Modular imperfect coverage model is analyzed in detail using formal methods to debug the specification errors and remove any ambiguity. A modular and hierarchical decomposition method that combines Markov Analysis and a separable binary decision diagrams based combinatorial method is applied to the reliability analysis of dynamic hierarchical systems subject to modular imperfect coverage. We illustrate our approach by analyzing an example hierarchical computer system.
Prashanthi Boddu, Liudong Xing
DASC2
2007 MBDD versus MMDD for Multistate Systems Analysis
abstract
Many combinatorial reliability models assume binary designation of states for both systems and their components. In many real applications, however, systems and their components may have more than two states (or levels of performance) varying from perfect operation to complete failure. In this paper, we present two combinatorial decision diagram based models that have been proposed for the analysis of multistate systems: multistate binary decision diagrams (MBDD) based approach and multistate multivalued decision diagrams (MMDD) based approach. And we conduct an empirical performance comparison between those two methods in terms of model size and computational complexity via two illustrative examples.
Akhilesh Shrestha, Liudong Xing, Yuan-Shun Dai
DASC2
2007 Node-Replacement Policies to Maintain Threshold-Coverage in Wireless Sensor Networks
abstract
With the rapid deployment of wireless sensor networks, there are several new sensing applications with specific requirements. Specifically, target tracking applications are fundamentally concerned with the area of coverage across a sensing site in order to accurately track the target. We consider the problem of maintaining a minimum threshold-coverage in a wireless sensor network, while maximizing network lifetime and minimizing additional resources. We assume that the network has failed when the sensing coverage falls below the minimum threshold-coverage. We develop three node-replacement policies to maintain threshold-coverage in wireless sensor networks. These policies assess the candidature of each failed sensor node for replacement. Based on different performance criteria, every time a sensor node fails in the network, our replacement policies either replace with a new sensor or ignore the failure event. The node-replacement policies replace a failed node according to a node weight. The node weight is assigned based on one of the following parameters: cumulative reduction of sensing coverage, amount of energy increase per node, and local reduction of sensing coverage. We also implement a first-fail-first-replace policy and a no-replacement policy to compare the performance results. We evaluate the different node-replacement polices through extensive simulations. Our results show that given a fixed number of replacement sensor nodes, the node-replacement policies significantly increase the network lifetime and the quality of coverage, while keeping the sensing-coverage about a pre-set threshold.
Sachin Parikh, Vinod Vokkarane, Liudong Xing, Dayalan Kasilingam
ICCCN3
2007 Reliability Evaluation of Phased-Mission Systems With Imperfect Fault Coverage and Common-Cause Failures
abstract
This paper proposes efficient methods to assess the reliability of phased-mission systems (PMS) considering both imperfect fault coverage (IPC), and common-cause failures (CCF). The IPC introduces multimode failures that must be considered in the accurate reliability analysis of PMS. Another difficulty in analysis is to allow for multiple CCF that can affect different subsets of system components, and which can occur$s$-dependently. Our methodology for resolving the above difficulties is to separate the consideration of both IPC and CCF from the combinatorics of the binary decision diagram-based solution, and adjust the input and output of the program to generate the reliability of PMS with IPC and CCF. According to the separation order, two equivalent approaches are developed. The applications and advantages of the approaches are illustrated through examples. PMS without IPC and/or CCF appear as special cases of the approaches.
Liudong Xing
IEEE Trans. Reliab.1
2006 Fault-Intrusion Tolerant Techniques in Wireless Sensor Networks
abstract
Wireless sensor networks (WSN) consist of a large number of tiny sensor devices that have limited power and limited sensing, computation, and wireless communications capabilities. Sensor nodes usually operate in unattended and even harsh environments, and as a result, sensor nodes are prone to failures and are vulnerable to malicious attacks. Therefore, for reliable and secure computation and communication in WSN, fault tolerance and intrusion tolerance become two essential attributes that should be designed into WSN. In this paper, we study state-of-the-art fault tolerance and intrusion tolerance techniques for WSN and propose a new fault-intrusion tolerant routing mechanism called MVMP (multi-version multi-path) for WSN that will support highly reliable and secure sensor networks. Follow-up work on the MVMP to be carried out in the near future is also discussed
Ruiping Ma, Liudong Xing, Howard E. Michel
DASC2
2006 Infrastructure Communication Reliability of Wireless Sensor Networks
abstract
In this paper we consider the problem of modeling and evaluating the infrastructure communication reliability (ICR) of wireless sensor networks (WSN). Our approach is progressive and based on reduced ordered binary decision diagrams (ROBDD). Due to the nature of the ROBDD and the reduction scheme used in the progressive process, the approach is computationally efficient and is easy to implement. We also incorporate the consideration of common cause failures (CCF) into the reliability analysis of WSN. We illustrate the application and advantages of our approach by working through the analysis of an example WSN
Akhilesh Shrestha, Liudong Xing, Hong Liu 0019
DASC2
2006 QoS reliability of hierarchical clustered wireless sensor networks
abstract
We consider the problem of reliability modeling and analysis of hierarchical clustered wireless sensor networks (WSN) in this paper. We propose reliability measures that integrate the conventional connectivity-based network reliability with the sensing coverage measure indicating the quality of service (QoS) of the WSN. And we propose a progressive approach for evaluating such coverage-oriented QoS reliability. Our approach is unique in that we analyze a more practical and generic hierarchical clustered WSN. We also incorporate the consideration of some dynamic dependencies, like the standby spare behavior, by proposing a modular scheme
Liudong Xing, Akhilesh Shrestha
IPCCC1
2004 Comments on PMS BDD generation in 'A BDD-based algorithm for Reliability Analysis of phased-mission systems'
abstract
This paper discusses some of the difficulties in the paper: "A BDD-based algorithm for reliability analysis of phased-mission systems" by X. Zang, H. Sun, and K. S. Trivedi.
Liudong Xing, Joanne Bechta Dugan
IEEE Trans. Reliab.1
2004 A separable ternary decision diagram based analysis of generalized phased-mission reliability
abstract
This paper considers the reliability analysis of a generalized phased-mission system (GPMS) with two-level modular imperfect coverage. Due to the dynamic behavior & the statistical dependencies, generalized phased-mission systems offer big challenges in reliability modeling & analysis. A new family of decision diagrams called ternary decision diagrams (TDD) is proposed for use in the resulting separable approach to the GPMS reliability evaluation. Compared with existing methods, the accuracy of our solution increases due to the consideration of modular imperfect coverage; the computational complexity decreases due to the nature of the TDD, and the separation of mission imperfect coverage from the solution combinatorics. In this paper, the TDD-based separable approach is presented, and compared with existing methods for analyzing the GPMS reliability. An example generalized phased-mission system is analyzed to illustrate the advantages of our approach.
Liudong Xing, Joanne Bechta Dugan
IEEE Trans. Reliab.1
2003 Reliability Analysis of Fault-Tolerant Systems with Common-Cause Failures
abstract
This paper proposes two separable approaches for analyzing reliability of fault-tolerant systems that are possibly subject to both common-cause failures (CCF) and imperfect fault coverage (IPC). Almost all the existing reliability assessment methods either fail to consider CCF or fail to consider IPC. Both cases result in exaggerated system reliability. Our methods provide accurate reliability by incorporating both CCF and IPC into the reliability modeling for an arbitrary system. Further, the two methods have low computational complexity in that they separate the consideration of both IPC and CCF from the combinatorics of the solution. The methods are illustrated with a concrete analysis of a sample network system. A general illustration of those approaches in the context of two important classes of reliability problems: networks and fault trees, is also provided.
Liudong Xing
DSN1
2002 Analysis of generalized phased-mission system reliability, performance, and sensitivity
abstract
This paper proposes a generalized phased-mission system (GPMS) analysis methodology called GPMS-CPR that has high computation efficiency and is easy to implement. GPMS-CPR can evaluate a wider range of more practical systems with less restrictive mission requirements, while offering more human-friendly performance indices such as multilevel grading as compared to the previous PMS approaches. GPMS-CPR also accounts for imperfect coverage. This method implements an exciting synthesis of several approaches into a single methodology. Further, it extends the methodology to allow for the analysis of sensitivity of each component for each phase as well as for the entire phased mission, with respect to the multilevel reliability for the GPMS. The conventional phase-OR PMS appear as a special case this GPMS. The advantages of this approach are in the low computational complexity, broad applicability, easy implementation, and high integration (reliability, performance, and sensitivity for GPMS). The methodology is illustrated with examples.
Liudong Xing, Joanne Bechta Dugan
IEEE Trans. Reliab.1