VLDB 2026 Research / reviewers in the wild / expert
Gregory Levitin
dblp:22/3589
· DBLP profile ↗
67ranked-venue papers
55as first author
7since 2021 · last 2026
0000-0002-2107-8291ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 41 · 35 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 16 · 13 first-author · 3 since 2021Systems, architecture and hardware · 6 · 4 first-authorSecurity and privacy · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimizing Mission Abort Policy With Threshold Voting of Imperfect Shock DetectorsabstractShock count is a key parameter used in designing mission abort policies (MAPs) for systems executing their operations under random shock conditions. Existing models mostly assume a perfect mechanism of detecting shocks. In practice, the shock detection system may fail to detect shocks that have occurred (false negative) or flag non-existent shocks (false positive), both leading to wrong shock count and misleading MAP designs. This work's contribution lies in modeling a single-attempt mission system with a fault-tolerant shock detection system that applies threshold voting among multiple imperfect detectors to contribute to the mission abort decision based on shock count and system operation time. A probabilistic approach is put forward for assessing mission performance of the considered system in the form of task success probability (TSP), survival probability of system (SPS), and expected losses of mission (ELM). An ELM minimization problem is further formulated and solved, which aims to determine the optimal tri-parametric MAP, achieving a balance between TSP and SPS. We analyze a drone-based surveillance system to showcase the suggested model. We also examine the impact of key parameters (cost, shock occurrence rate and detection probability) on mission performance metrics and on the best-obtained MAPs, leading to important managerial recommendations. Gregory Levitin, Liudong Xing |
IEEE Trans. Reliab. | 1 |
| 2024 | Mission Aborting Policies and Multiattempt MissionsabstractThe state of the art in the recently emerged and rapidly developing field of mission aborting and multiattempt missions is briefly discussed. The research aims to develop optimal rules for interrupting a mission and activating system rescue procedures (and, if needed, subsequent attempts to complete the mission) that balance the probabilities of mission success and system loss or minimize the cost of losses associated with the mission. Gregory Levitin, Liudong Xing |
IEEE Trans. Reliab. | 1 |
| 2022 | Reliability versus Vulnerability of N-Version Programming Cloud Service Component With Dynamic Decision Time Under Co-Resident AttacksabstractThe virtual machine (VM) co-resident architecture of cloud computing enables simultaneous provision of multiple services to different users, but also makes these services vulnerable to co-resident attacks. For example, by establishing side channels, a malicious attacker can access and even corrupt services performed by other VMs co-residing on the same server as the attacker's VM (AVM). We model a threshold-voting-basedN-version programming service component with multiple independent versions simultaneously performing the same requested service to enhance the service reliability. However, the reliability enhancement can be greatly hindered by the co-resident attack, which may corrupt an adequate number of versions leading to a wrong output. We formulate and solve constrained optimization problems that determine the number of service component versions and the voting threshold to balance two conflicting service performance metrics: reliability (service component success probability) and vulnerability (service corruption attack success probability). Two cases respectively having certain and uncertain knowledge about the attacker's power in terms of the number of AVMs are considered. We also investigate impacts of different model parameters on the service performance as well as on solutions to the considered optimization problems through examples. Gregory Levitin, Liudong Xing, Yanping Xiang |
IEEE Trans. Serv. Comput. | 1 |
| 2022 | Optimal Preventive Replacement for Cold Standby Systems With Elements Exposed to Shocks During Operation and Task TransfersabstractThis article considers heterogeneous, cold standby systems performing missions with the fixed amount of work when a failure of an operating element results in a mission failure. A system is operating in a random environment modeled by the Poisson process of shocks. Each shock decreases the remaining lifetime of an operating element and, therefore, its preventive replacement (PR) is scheduled on experiencing the predetermined number of shocks. An important feature of the discussed model is that the failure can also occur during these PRs (task transfers) with two elements involved. The duration of the task transfer depends on the time from the start of a mission. The recursive equations for obtaining the mission success probability are derived and the corresponding numerical algorithm is developed. The number of shocks triggering elements’ replacements is obtained as a solution of the formulated optimization problem. The numerical example with the detailed analysis for a set of virtual machines operating in a cloud computing environment is presented. Gregory Levitin, Maxim Finkelstein, Yanping Xiang |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2022 | Mission Aborting in n-Unit Systems With Work SharingabstractMission aborting has recently attracted great attention, where mission abort rules (MARs) have been modeled and optimized for different types of technological systems aiming to effectively mitigate the risk of system losses. However, none of the existing works have considered systems with multiple work-sharing units. This article makes advancements in the state of the art by modeling condition-based MARs for nonrepairable work-sharing systems that must perform a specified amount of work during the primary mission (PM). The MAR considered presumes aborting the PM to prevent considerable damage when the number of available units reduces to a certain number${k}$while the amount of work accomplished in the PM is less than${L}$(${k}$). After the PM abortion, a rescue procedure (RP) is executed by the remaining units to survive the system. Dynamic operating conditions during PM and RP are considered. A probabilistic model-based numerical algorithm is proposed to evaluate several performance metrics, including mission success probability, RP success probability, damage avoidance probability, and expected cost of losses. The MAR optimization problem for minimizing the expected cost of losses is formulated and solved using the genetic algorithm. An example of a chemical reactor system is provided to demonstrate the proposed algorithm as well as the benefit of MARs optimized in comparison to the “no abort” policy and the most conservative abort policy (the PM aborts upon the failure of any unit). Effects of several model parameters on the system performance metrics and optimization solutions are also examined through examples. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2022 | Co-Residence Data Theft Attacks on N-Version Programming-Based Cloud Services With Task CancelationabstractPowered by virtualization, the cloud computing has brought good merits of cost effective and on-demand resource sharing among many users. On the other hand, cloud users face security risks from co-residence attacks when using this virtualized platform. Particularly, a malicious attacker may create side channels to steal data from a target user’s virtual machine (VM) that co-resides with the attacker’s VM on the same physical server. This article models a cloud service undergoing the co-residence data theft attacks. The threshold-voting-based${N}$-version programming (NVP) is implemented to improve the service reliability, where multiple service component versions (SCVs) are activated in parallel to perform the requested service. The final output is determined upon receiving a threshold number of identical outputs from the SCVs, immediately followed by canceling all outstanding SCVs to reduce expenses. Probabilistic models are first introduced to evaluate performance metrics of the considered service, including the data theft probability, service success probability, expected service operation time, and expected utility. Optimization problems are further solved to find the optimal number of SCVs maximizing the expected utility. Interactions among different model parameters and VM allocation policies, as well as their effects on the considered performance metrics and on the optimization solutions are studied through examples. Gregory Levitin, Liudong Xing, Yanping Xiang |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2021 | Defending N-Version Programming Service Components against Co-Resident Attacks in IoT Cloud SystemsabstractThe real innovation of Internet of Things (IoT) can be spurred only when being combined with cloud computing, a paradigm that allows numerous users to simultaneously access configurable resources and services. However, serious vulnerability concerns have arisen from the virtual machine co-resident architecture of the IoT cloud. Specifically, co-resident attacks can be launched, where an attacker can access and corrupt a user's sensitive data/software by co-locating their virtual machines on the same physical server. Various solutions have been suggested in literature to mitigate negative effects of the co-resident attacks in the cloud environment. However, to the best of our knowledge no work has been performed for studying co-resident attacks in cloud systems withN-version programming (NVP), a popular redundancy technique for enhancing survivability of critical cloud service components. This paper makes original contributions by modeling IoT cloud system services implementing the NVP component redundancy, and evaluating the corruption probability of the NVP service component. Further, users’ policies on choosing the optimal number of service component versions are investigated through formulating and solving a new set of optimization problems with the objective to minimize the expected cost of losses of a cloud service provider. As demonstrated through examples, these policies can effectively help defend the NVP service component against the co-resident attacks in the cloud system. Liudong Xing, Gregory Levitin, Yanping Xiang |
IEEE Trans. Serv. Comput. | 2 |
| 2019 | Optimal Spot-Checking for Collusion Tolerance in Computer GridsabstractMany grid-computing systems adopt voting-based techniques to resist sabotage. However, these techniques become ineffective in grid systems subject to collusion behavior, where some malicious resources can collectively sabotage a job execution by returning identical wrong results. Spot-checking has been used to detect and tackle the collusive issue by sending randomly chosen resources a certain number of spotter jobs with known correct results to estimate resource credibility based on the returned result. This paper makes original contributions by formulating and solving a new spot-checking optimization problem for grid systems subject to collusion attacks, with the objective to minimize probability of the genuine task failure (PGTF, i.e., the wrong output probability) while meeting an expected overhead constraint. The problem solution contains an optimal combination of task distribution policy parameters, including the number of deployed spotter tasks, the number of resources tested by each spotter task, and the number of resources assigned to perform the genuine task. The optimization procedure encompasses a new iterative method for evaluating system performance metrics of PGTF and expected overhead in terms of the total number of task assignments. Both fixed and uncertain attack parameters are considered. Illustrative examples are provided to demonstrate the proposed optimization problem and solution methodology. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2018 | Optimization of dynamic spot-checking for collusion tolerance in grid computing
Gregory Levitin, Liudong Xing, Barry W. Johnson, Yuan-Shun Dai |
Future Gener. Comput. Syst. | 1 |
| 2018 | Mission Abort Policy in Heterogeneous Nonrepairable 1-Out-of-N Warm Standby SystemsabstractMany real-world critical systems, such as aircraft and human space flight systems, utilize mission aborts to enhance the survivability of the system. Specifically, the mission objectives of these systems can be aborted in cases where a certain malfunction condition is met, and a rescue or recovery procedure is then initiated for system survival. Traditional system reliability models typically cannot address the effects of mission aborts, and thus are not applicable to analyzing systems subject to mission abort requirements. In this paper, we first develop a numerical methodology to model and evaluate mission success probability and system survivability of 1-out-of-N warm standby systems subject to constant or adaptive mission abort policies. The system components are heterogeneous, characterized by different performances and different types of time-to-failure distributions. Based on the proposed evaluation method, we make another new contribution by formulating and solving the optimal mission abort problem, as well as a combined optimization problem that identifies the mission abort policy and component activation sequence maximizing mission success probability while achieving the desired level of system survivability. Efficiencies of constant and adaptive mission abort policies are compared through examples. Examples also demonstrate the tradeoff between system survivability and mission success probability due to the utilization of a mission abort policy. Such a tradeoff analysis can help identify optimal decisions on system mission abort and standby policies, promoting safe and reliable operation of warm standby systems. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Reliab. | 1 |
| 2018 | Optimizing Dynamic Performance of Multistate Systems With Heterogeneous 1-Out-of-N Warm Standby ComponentsabstractThis paper models and optimizes dynamic performance of multistate systems with a general series parallel structure. Each system component is a 1-out-of-N warm standby configuration of heterogeneous functional elements, which can be characterized by different time-to-failure distributions, performances, and costs. The entire system must satisfy a random demand specified by a time-dependent distribution. An iterative algorithm is developed for determining performance stochastic processes of particular components. A universal generating function technique is used for evaluating expected system availability and unsupplied demand over a particular mission time for the considered system. Two types of optimization problems are then identified and solved, with the objective of finding component structures and element activation sequences to maximize system availability, or minimize unsupplied system demand, or minimize total cost. Optimization results can facilitate the optimal decision on design and operation of multistate series parallel systems. A practical example of a power station coal transportation system is provided to illustrate application of the proposed methodology and optimization problems. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2018 | Optimizing Computational Mission Operation by Periodic Backups and Preventive ReplacementsabstractThis paper models a warm standby system where a single element is online performing a specified mission task (e.g., a computing task) and subject to corrective replacement (CR) by an available standby element upon its failure. During the mission, preventive replacements (PRs) are also performed to renew the aged or worn online operating element before its actual failure according to a predetermined policy. In addition, to facilitate an effective restoration of system function in case of CR or PR happening, backups are also performed periodically so that the mission task can be resumed from the last successful backup point instead of from scratch. The mission succeeds if the specified mission task is accomplished; in other words, the mission fails when no operating elements remain prior to the mission task completion. In this paper, we make new contributions by first proposing an event transition-based numerical method to evaluate mission performance indices of the considered standby system subject to periodic backups, CR and PR. Mission success probability (MSP), expected mission completion time, expected mission operation cost (EMC), and expected uncompleted work fraction are evaluated. Based on the suggested evaluation algorithm, we make another contribution by formulating and solving optimization problems that help to determine the optimal backup-PR policy or the optimal combination of element activation sequencing and backup-PR policy to maximize MSP or minimize EMC. Influence of element performance and reliability parameters, data backup and retrieval complexity parameters on the optimal operation policy is investigated. Findings from this paper can guide the optimal decision making on policies related to element sequencing, backups as well as preventive maintenance planning, contributing toward reliable and cost-effective design and operation of standby computing systems. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2017 | Optimal data partitioning in cloud computing system with random server assignment
Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
Future Gener. Comput. Syst. | 1 |
| 2017 | Dynamic Checkpointing Policy in Heterogeneous Real-Time Standby SystemsabstractThis paper models 1-out-of-N standby computing systems with a dynamic checkpointing policy. The system performs a real-time mission task that has to be accomplished within an allowed mission time. During the mission, to facilitate an effective failure recovery the system undergoes checkpointing procedures according to a policy that dynamically determines a checkpointing frequency based on the activated element and remaining work for completing the mission. System elements are heterogeneous; they can follow different, arbitrary types of time-to-failure distributions, have different performance and wait in different standby modes before their activation. A new numerical algorithm based on state space event transitions is first proposed to evaluate mission success probability of the real-time standby systems considered in this work. Additional new contributions are made by formulating and solving optimal dynamic checkpointing policy problems, as well as an integrated optimization problem that finds the optimal combination of checkpointing policy and element activation sequence maximizing mission success probability. Advantages of using the dynamic checkpointing policy over fixed even checkpoints are demonstrated through examples. Examples and results are also provided to illustrate effects of different mission and element parameters on mission success probability as well as on the optimal dynamic checkpointing policy. Gregory Levitin, Liudong Xing, Yuan-Shun Dai, Vinod Vokkarane |
IEEE Trans. Computers | 1 |
| 2017 | Optimal Periodic Inspections and Activation Sequencing Policy in Standby Systems With Condition-Based Mode TransferabstractThis paper models a hybrid standby system subject to periodic inspections and condition-based standby mode transfers during a mission. At the beginning of the mission only one element is online and operating. The second element waits in a hot standby mode being ready to replace the failed online element at any time. Other elements wait in less-stressful and less-costly warm standby mode. During the mission periodic inspections are performed for checking conditions of the online and hot standby elements and subsequently triggering necessary mode transfer(s) of available warm standby element(s) to replace the failed hot standby element and/or online element. We suggest an efficient numerical method to assess availability and expected total mission cost (including standby cost, operation cost and mode transfer cost of system elements, inspection cost, system interruption or idle cost) of the considered system. The algorithm is flexible and applicable to arbitrary type of time-to-failure distributions. Then we formulate and solve new optimization problems that identify the optimal combination of inter-inspection interval and element activation sequence to minimize expected total mission cost while satisfying a certain constraint on system availability. As illustrated through examples, the optimization results can facilitate cost-effective and availability-aware planning of system inspection and operation. Yuan-Shun Dai, Gregory Levitin, Liudong Xing |
IEEE Trans. Reliab. | 2 |
| 2017 | Preventive Replacements in Real-Time Standby Systems With Periodic BackupsabstractThis paper models a real-time warm standby system that has to accomplish a specified amount of task by a hard deadline. The system is subject to corrective replacements (CRs) upon failure of its operating element. It can also be renewed according to a predetermined schedule through preventive replacements (PRs). To facilitate an effective recovery of system operation after replacements, periodic backups are performed so that warm standby elements, upon being activated, can take over the mission task from the last backup point instead of from scratch. This paper presents a novel integrated model that considers effects of periodic backups, CRs and PRs in analyzing and optimizing real-time warm standby systems. Mission success probability and expected mission completion time are evaluated. Impacts of different mission and element parameters on mission success probability, optimal backup and PR policies, and optimal element activation sequence are investigated. It is shown that in warm standby systems with periodic backups and tight deadlines, PRs can improve the mission success probability even when they take the same time as CRs. When the maximum allowed mission time exceeds a certain level, PRs become ineffective and the optimal policy can involve only periodic backups. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Reliab. | 1 |
| 2017 | Optimization of Component Allocation/Distribution and Sequencing in Warm Standby Series-Parallel SystemsabstractExisting works on redundancy allocation problems have typically focused on active or cold standby redundancies or a mix of them; little research is dedicated to warm standby systems but with an assumption of allocating the same choice of components within each subsystem. Motivated by the fact that components with different costs and failure time distributions from different vendors can be available for the design of the same subsystem in practice, this paper advances the state-of-the-art by presenting a solution methodology to determine combined optimal design configuration and optimal operation of heterogeneous warm standby series-parallel systems. Particularly, based on a proposed numerical reliability evaluation algorithm, two combined optimization problems (component allocation and sequencing problem, and component distribution and sequencing problem) are formulated and solved. Necessity and significance of the proposed methodology are illustrated through examples. Efficiency of the methodology is also successfully demonstrated on large warm standby series-parallel systems containing 14 subsystems of different choices of components. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Reliab. | 1 |
| 2017 | Reliability Versus Expected Mission Cost and Uncompleted Work in Heterogeneous Warm Standby Multiphase SystemsabstractThis paper considers 1-out-of-N : G warm standby (WS) systems subject to multiple-phased mission requirements, where performance, failure behavior, operation cost, replacement time, and cost of system elements can vary from phase to phase due to changing working conditions and stress levels. The system succeeds if its elements can accomplish a specified mission task within the maximum allowed time. A numerical algorithm is proposed for evaluating mission indices, including system reliability, expected mission cost, and expected uncompleted work of the heterogeneous, dynamic standby system considered. The influence of the maximum allowed mission time on mission indices is investigated. As demonstrated through examples, the proposed algorithm can also facilitate a determination of mission interruption based on reliability or cost-oriented criteria. Based on the proposed algorithm, an unconstrained optimal standby element sequencing problem is then formulated and solved for heterogeneous WS multiphase systems. The problem is to find the optimal activation sequence of system elements maximizing system reliability, or minimizing expected mission cost, or minimizing expected uncompleted work. The constrained optimization problems of minimizing expected mission cost subject to providing a desired level of mission reliability or uncompleted work are also solved. Illustrative examples of these optimization problems and their applications in estimating the role of each element in the mission success are presented. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2017 | Optimal Distribution of Nonperiodic Full and Incremental BackupsabstractData backups play an important role in effective recovery of system operations, especially, those in computing and Information Technology-related applications. By saving copies of information associated with the completed part of a mission task, a failed system, upon being repaired, can resume its operation from the latest backup point instead of having to repeat the entire work from scratch. This paper considers a repairable, real-time system performing a sequence of nonperiodic full backup (FB) and incremental backup (IB) procedures during its mission. The mission succeeds if a specified amount of work can be accomplished by the system within a deadline. New contributions of this paper are twofold. First, a numerical recursive method is proposed to model effects of the nonperiodic, mixed backup strategy in assessing mission success probability, and expected completion time of the considered system. The method is applicable to any distribution of FBs and IBs and has no limitation on system time-to-failure distributions. Second, the optimal backup policy problem is formulated and solved, which finds the distribution of FBs and IBs maximizing the mission success probability. Examples are provided to illustrate the proposed methodology as well as influence of different parameters on the optimal solution. Advantages of adopting nonperiodic backups over periodic ones are also illustrated. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2016 | Optimization of Full versus Incremental Periodic Backup PolicyabstractThis paper models repairable computing systems performing a mission that is successful if the system can accomplish a specified amount of work within the allowed mission time or deadline. During the mission the system is subject to a sequence of full and incremental data backup procedures to facilitate an effective system recovery and avoid repeating the entire mission work from the very beginning when a system failure happens. The repair time is fixed while the system time-to-failure can follow any arbitrary type of distributions. This paper makes novel contributions by first developing a new numerical algorithm to evaluate mission success probability and expected completion time of the considered repairable real-time computing systems subject to mixed full and incremental backups. Correctness of the proposed evaluation algorithm is verified using Monte Carlo simulations. We make another new contribution by formulating and solving the backup schedule optimization problem that finds the full and incremental backup frequencies maximizing the mission success probability. Through illustrative examples, effects of different parameters (including the system time-to-failure distribution parameter, maximum allowed mission time, data backup and retrieval times, storage availability, repair time and efficiency) on the mission success probability and expected completion time as well as on the optimal backup schedule solution are investigated. Gregory Levitin, Liudong Xing, Qingqing Zhai, Yuan-Shun Dai |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2016 | Heterogeneous Non-Repairable Warm Standby Systems With Periodic InspectionsabstractMotivated by practical applications (for example production systems, flow transmission systems, and transportation systems), this paper considers 1-out-of- N: G warm standby systems subject to periodic inspections. Periodic inspections are performed to detect failures of system components. If an online operating component failure is detected during inspection, an available warm standby component is activated to take over the mission task. The mission succeeds if system components can complete a pre-specified amount of work. A state space event transition based numerical algorithm is first suggested for evaluating important mission performance indices including mission success probability, expected mission cost, expected mission time, expected mission completion delay over a desired time, and expected uncompleted work of the considered standby system. Based on the proposed evaluation algorithm, inspection interval optimization problems are formulated and solved, which find the optimal value of the inspection interval to minimize the expected total mission cost subject to providing a desired level of mission success probability, expected completion time, and delay. Further, combined component sequencing and inspection interval optimization problems are solved, for minimizing the expected total mission cost, or maximizing the mission success probability. As illustrated through examples, the proposed methodology can also facilitate solving dynamic optimization problems that update the optimal solution after each component failure, leading to further improvements of various mission indices. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Reliab. | 1 |
| 2016 | Cold Standby Systems With Imperfect BackupabstractCold standby redundancy is a widely-applied design strategy to achieve high system reliability in numerous applications. To facilitate an effective system recovery in case of an online element failure occurring, the backup mechanism is typically implemented which enables a standby element to take over the mission task from a backup point instead of re-executing the entire mission task from the very beginning. However, the backup mechanism is not perfectly reliable in practice, and effect of its failure on the system reliability can be non-monotonic and is correlated to other system and element parameters. In this paper a new numerical approach is first proposed to evaluate reliability and expected mission completion time for 1-out-of- N: G cold standby systems subject to imperfect even backups. The system elements are non-repairable during the mission. Based on the proposed evaluation algorithm, effects of backup system reliability in connection with other system parameters including backup frequency, data backup and retrieval times, replacement failure probability, replacement time, and number of system elements are investigated through examples. Further, the optimal backup frequency and initiation sequencing problem is formulated and solved for heterogeneous cold standby systems, providing solutions that can maximize mission reliability or minimize expected mission completion time depending on design requirements. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Reliab. | 1 |
| 2016 | Heterogeneous Warm Standby Multi-Phase Systems With Variable Mission TimeabstractWhile the majority of existing works on modeling and optimizing standby systems focus on single-phased missions, only a few works consider standby systems with phased-mission requirements, and these works are only applicable to restricted cold-standby configurations with assumptions of negligible replacement times and fixed phase durations. This paper makes new theoretical contributions by suggesting a general model of heterogeneous 1-out-of- N: G warm standby phased-mission systems (PMS) with dynamic phase durations. The model takes into account diverse, phase-dependent performances and time-to-failure distributions of system elements. It also considers non-zero replacement times specific for activated elements, and for the mission phase when the replacement occurs. Both cold and hot standby PMSs are special cases of the proposed model. An algorithm for evaluating the mission reliability and expected completion time is first presented. The algorithm is applicable to any type of time-to-failure distribution for system elements. The heterogeneous standby PMS can demonstrate non-coherent behavior where the reduction of reliability of some elements can cause the increase of the entire mission reliability. Then it is demonstrated that the sequence of the element activation affects the mission reliability and its expected completion time. Hence the element sequence optimization problem is further formulated and solved for the heterogeneous warm standby PMS. Illustrative examples of mission reliability and expected completion time analysis and optimization are presented. Gregory Levitin, Liudong Xing, Amir Ehsani Zonouz, Yuan-Shun Dai |
IEEE Trans. Reliab. | 1 |
| 2015 | Mission Reliability, Cost and Time for Cold Standby Computing Systems with Periodic BackupabstractLife critical applications like space missions and flight controls require their computing systems to be equipped with some fault-tolerance mechanism to meet stringent reliability requirements by performing the intended function even in the case of element failures. Such benefit, however, cannot come without extra time as well as extra overhead and capital costs. This paper for the first time considers the modeling and evaluation of mission reliability, expected mission time and cost simultaneously for 1-out-of-$N$: G non-repairable cold standby computing systems subject to periodic backup actions. Based on the suggested numerical evaluation method, the optimal backup frequency problems are formulated and solved, providing the optimal number of backup operations during the mission to maximize the system reliability or to minimize the mission cost or time. In the case of non-identical system elements, the optimal standby element sequencing problem arises as the order in which the system elements are initiated can impact the system reliability and mission cost and time greatly; such problems are formulated and solved for the 1-out-of-$N$: G cold standby computing systems with periodic backups. Furthermore, a combined optimization problem is considered, where a combination of the element initiation sequence and backup frequency providing the best combination of mission reliability, cost, and time is found. The proposed methodology can facilitate a reliability-cost-time tradeoff study in the practical design of cold standby systems, thus assist in making the optimal decision on the system's standby and backup policy. Examples are provided for illustrating the considered problems and suggested solution methodology. Gregory Levitin, Liudong Xing, Barry W. Johnson, Yuan-Shun Dai |
IEEE Trans. Computers | 1 |
| 2015 | Effect of Failure Propagation on Cold vs. Hot Standby Tradeoff in Heterogeneous 1-Out-of-N: G SystemsabstractThis paper considers 1-out-of- N:G heterogeneous fault-tolerant systems that are designed with a mix of hot and cold standby redundancies to achieve the tradeoff between restoration and operation costs of standby elements. In such systems, the way in which the elements are distributed between hot and cold standby groups and the initiation sequence of all the cold standby elements can greatly affect the system reliability and mission cost. Therefore, it is significant to solve the optimal standby element distributing and sequencing problem (SE-DSP). The failure that occurs in a system element can propagate, causing the outage of other system elements, which complicates the solution to the SE-DSP problem. In this paper, we first propose a numerical method for evaluating the reliability and expected mission cost of 1-out-of- N:G systems with mixed hot and cold redundancy types and propagated failures. Two different failure propagation modes are considered: an element failure causing the outage of all the system elements, and an element failure causing the outage of only working or hot standby elements but not cold standby elements. A genetic algorithm is utilized as an optimization tool for solving the formulated SE-DSP problem, leading to a solution that can minimize the expected mission cost of the system while providing a desired level of the system reliability. Effects of the failure propagation probability on the system reliability, expected mission cost, as well as the optimization results are investigated. The suggested methodology can facilitate a reliability-cost tradeoff study of the considered systems, thus assisting in optimal decision making regarding the system's standby policy. Examples are provided for illustrating the considered problem as well as the proposed solution methodology. Gregory Levitin, Liudong Xing, Hanoch Ben-Haim, Yuan-Shun Dai |
IEEE Trans. Reliab. | 1 |
| 2015 | Reliability of Non-Coherent Warm Standby Systems With ReworkingabstractIn this paper we model and analyze non-repairable 1-out-of- N : G warm standby systems subject to periodic backups and dynamic reworking. Particularly, in such systems, a standby element must redo some portion of already performed work by the failed online element before taking over the mission task, which makes the actual mission time dynamic. The considered systems are widely used in applications such as computing and manufacturing, but have not been well studied in reliability theory. In this work, we make new contributions by suggesting a numerical algorithm to evaluate the reliability of the considered warm standby systems. It is revealed that these systems are non-coherent, where the system reliability has non-monotonic dependence on the reliability of individual elements. Numerical examples further show that the non-coherency phenomenon is more distinguished for elements initiated earlier than those initiated later in the warm standby list. Example results also imply that placing highly unreliable elements at the end of the warm standby waiting list, or even removing them from the system planning, can enhance the reliability of a warm standby system subject to reworking. Findings from this work can guide the reliability design of the considered warm standby systems in practice. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Reliab. | 1 |
| 2015 | Reliability and Mission Cost of 1-Out-of-N: G Systems With State-Dependent Standby Mode TransfersabstractThe paper proposes a new fault tolerant system design model, in particular a 1-out-of- N:G hybrid redundant system with standby elements subject to state-dependent standby mode transfers. Specifically, in such systems, one standby element always resides in the hot standby mode, and thus is ready to replace the failed online element at any time to make the system dependable. If the online operating element or the hot standby element fails, one of the warm standby elements is immediately transferred to the hot standby mode. A numerical algorithm is first suggested for evaluating the reliability and the expected mission cost of the considered system. The algorithm is based on a discrete approximation of element time-to-failure distributions, and can work with any type of distribution. Furthermore, based on the suggested algorithm, the problem of optimal sequencing of standby elements initiation is formulated and solved. The objective of this optimization problem is to minimize the expected mission cost associated with elements' standby and operation expenses, as well as the mode transfer expenses, while meeting a certain system reliability constraint. Illustrative examples are provided. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Reliab. | 1 |
| 2015 | Non-Homogeneous 1-Out-of-N Warm Standby Systems With Random Replacement TimesabstractIn standby systems, when an online working element fails, a replacement procedure is initiated to activate a standby unit which will take over the mission task to sustain the system function. Existing works on standby systems have mostly assumed that such replacement procedure takes a negligible or fixed amount of time. This assumption is not practical in many real-world systems, where the replacement procedure can take times that are random and different for different standby elements. This paper makes novel contributions by considering the effects of the random replacement times in analyzing and optimizing 1-out-of- N: G non-repairable warm standby systems. The system elements are not necessarily identical; different elements can have different time-to-failure distributions, different performance levels, and different replacement time distributions. The system is considered failed if the elements cannot complete the specified amount of work (mission task) within the maximum allowed mission time. A numerical algorithm is first proposed to simultaneously evaluate the mission reliability and expected mission completion time of the considered warm standby system. Influences of different mission and element parameters on the mission reliability and expected completion time are investigated. It is revealed that the considered warm standby systems exhibit non-coherent behavior as mission reliability may increase with the decrease of an element's reliability. Due to heterogeneity of the system elements, the order in which the elements are initiated and replaced can affect the mission reliability and actual completion time significantly. Therefore, based on the suggested numerical evaluation algorithm, we further formulate and solve the optimal element replacement sequencing problem for the considered warm standby system subject to random replacement times. Examples are given to demonstrate the considered problems and proposed methodology. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Reliab. | 1 |
| 2015 | Heterogeneous 1-Out-of-N Warm Standby Systems With Dynamic Uneven BackupsabstractIn this paper, mission reliability, expected mission completion time, and cost of non-repairable 1-out-of- N: G warm standby sparing systems subject to uneven backup actions are modeled and optimized. The backup actions are used to facilitate the data recovery process in the case of an online operating element failure, which enables an activated standby element to take over the mission task through subsequent data retrievals. Both data backup and retrieval times are dynamic, and physically dependent on the amount of work performed. The system elements are not necessarily identical; each element can be characterized by a different time-to-failure distribution, a different performance, and a different level of readiness to take over the system task during the warm standby mode. An iterative numerical method is first proposed to simultaneously evaluate mission reliability, expected mission completion time, and the cost of the considered heterogeneous warm standby systems. Due to the non-monotonic effect of the backup distribution on the mission reliability, time, and cost, we formulate and solve the optimal backup distribution problem considering different combinations of optimization objectives and constraints. In the case of system elements being non-identical, their activation order can influence the mission reliability, expected mission completion time, and mission cost significantly. Therefore, we also formulate and solve the optimal element sequencing problem for the considered system. Furthermore, new integrated optimization problems are formulated and addressed. The integrated optimization aims to identify the optimal combination of backup distribution and element activation order that maximizes the mission reliability, or minimizes the expected mission time or mission cost. As shown through examples, the proposed methodology can implement a tradeoff analysis among the three mission requirements of reliability, cost, and completion time, leading to the optimal decision on both backup and standby policies of warm standby systems. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Reliab. | 1 |
| 2015 | Optimal Backup Distribution in 1-out-of-N Cold Standby SystemsabstractThis paper considers nonrepairable 1-out-of-N: G cold standby (CS) systems subject to uneven backup actions as well as dynamic backup and retrieval times. In such systems, only one element is online and operates with the rest of the elements waiting in the unpowered CS mode. The operating element performs data backup actions when certain fractions of the mission task are accomplished. The backup actions facilitate data recovery in case of an operating element failure, which allows an activated standby element to take over the task through data retrieval. Backup distribution can have a nonmonotonic effect on mission reliability, time, and cost, leading to the optimal backup distribution problem. In this paper, we first suggest a numerical method to model and evaluate mission reliability, expected time, and cost simultaneously for the considered CS systems with uneven backup actions and dynamic backup and retrieval times. Based on the suggested evaluation method and genetic algorithm, the optimal backup distribution problem is then formulated and solved with the objective to minimize the expected mission cost subject to meeting certain levels of mission reliability and expected mission time. Examples show that the proposed methodology can facilitate a tradeoff study between mission reliability, and time and cost, which assists in the optimal decisionmaking for the backup policy used in the practical design of CS systems. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2015 | Optimal Design of Hybrid Redundant Systems With Delayed Failure-Driven Standby Mode TransferabstractStandby redundancy is a design technique that has been widely adopted to enhance system reliability and achieve fault tolerance in many critical applications. In this paper, we consider a 1-out-of-N: G hybrid standby redundant system with unrepairable elements being subject to delayed failure-driven standby mode transfers. Specifically, in the considered system, all the standby elements are initially in a warm standby mode (WSM) but can be transferred to a hot standby mode (HSM) so as to be ready to replace the online operating element when it fails. The WSM to HSM transfer is performed with a fixed time delay after no element resides in HSM either due to the element failure or because the element leaves the HSM to replace the failed online element for operation. A new iterative numerical algorithm is first proposed for evaluating reliability and expected mission cost (relevant to elements' standby expense, operation expense, as well as mode transfer expense) of the considered hybrid standby system. The algorithm has no restriction on element time-to-failure distribution types. Based on the proposed evaluation algorithm, we further formulate and solve a new optimization problem that finds the optimal delay and optimal sequence of standby elements with the objective to minimize expected mission cost while satisfying a certain level of mission reliability constraint. Examples are presented to demonstrate applications of the proposed methodology. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2014 | Minimum Mission Cost Cold-Standby Sequencing in Non-Repairable Multi-Phase SystemsabstractThis paper considers the optimal cold standby element sequencing problem (SESP) for 1-out-of- n: G heterogeneous non-repairable cold-standby systems that accomplish multi-phase missions. Given a fixed set of element choices, the objective of the optimal system design is to select the initiation sequence of the system elements so as to minimize the expected mission cost while providing a desired level of system reliability. It is assumed that during different mission phases the elements are exposed to different stresses, which affects their time-to-failure distributions. The startup and exploitation costs of system elements are also phase dependent. We suggest an algorithm for evaluating the mission reliability and expected mission cost based on a discrete approximation of time-to-failure distributions of the system elements. A genetic algorithm is used as an optimization tool for solving the formulated SESP for multi-phase cold-standby systems. Examples are given to illustrate the considered problem and the proposed solution methodology. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Reliab. | 1 |
| 2014 | Minimizing Bypass Transportation Expenses in Linear Multistate Consecutively-Connected SystemsabstractMany continuous transportation systems can be represented as multi-state linear consecutively connected systems consisting of N+1 linearly ordered nodes. Some of these nodes contain statistically independent multistate elements with different characteristics. Each element j can provide a connection between the node to which it belongs and Xjnext nodes, where Xjis a discrete random variable with known probability mass function. If the system contains nodes not connected with any previous node, then gaps exist that require bypass transportation solutions associated with considerable expenses. An algorithm based on the universal generating function method is suggested for evaluating the expected value of these expenses. A problem of finding the multi-state element allocation that minimizes the expected bypass transportation expenses is formulated and solved. Illustrative examples are presented. Gregory Levitin, Wei-Chang Yeh 0001, Yuan-Shun Dai |
IEEE Trans. Reliab. | 1 |
| 2014 | Structure Optimization of Nonrepairable Phased Mission SystemsabstractSystem structure optimization is a well-studied problem in the field of reliability engineering, which aims to achieve the best possible reliability versus cost solutions for the system design. This problem is usually solved for systems that do not change their task and configuration during the mission. However, many practical systems are phased-mission systems (PMS), where the mission involves multiple, consecutive, and nonoverlapping phases of operation. An accurate analysis of PMS must consider the dynamics in system configuration, success criteria, and element behavior as well as statistical dependence of element states across phases. In this paper, we propose a method for solving the structure optimization problem of multistate PMS consisting of nonidentical nonrepairable binary elements. The system configuration and demand as well as the failure distributions of system elements can change from phase to phase. The proposed approach is based on a recursive algorithm for the reliability evaluation of PMS and a genetic algorithm for the structure optimization of PMS. The method is illustrated using an example of a six-phased mission running on an airborne distributed computing system. Yuan-Shun Dai, Gregory Levitin, Liudong Xing |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2014 | Mission Cost and Reliability of 1-out-of- $N$ Warm Standby Systems With Imperfect Switching MechanismsabstractIn this paper, mission cost and reliability of 1-out-of-N: G nonrepairable warm standby systems with imperfect switching are modeled and analyzed using an iterative method. A general switching structure is considered, which consists of an overall fault detection mechanism (FDM) and a set of individual switches (one for each standby element). The failure of the FDM prevents the replacement of the failed online element by any standby element while the failure of an individual switch only makes the corresponding standby element unavailable. The entire mission fails either when the FDM fails before the failure of the online element during the mission, or when all the system elements have failed or become unavailable before the mission completion. Based on the proposed algorithm for the mission cost and reliability analysis, the optimal element sequencing problem is further formulated and solved for 1-out-of-N: G nonrepairable warm standby systems with nonidentical elements and imperfect switching mechanisms. The objective of the problem is to find the optimal initiation sequence of system elements that can minimize the expected mission cost while providing a certain level of system reliability. Examples are given to illustrate the considered problem and the proposed solution methodology. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2013 | Multi-State Vector-k-Out-of-n SystemsabstractMulti-state generalizations of binary k-out-of-n systems yielded two generic types of models. The first type, component-based systems, accounts for the number of system components being in certain states. The second type, weight-based systems, assigns specific weights (or performance levels) to each system component in each state, and accounts for the cumulative weight of the components. The two types of systems cannot be mapped to each other. This paper introduces a new general model, named the multi-state vector-k-out-of-n system, which is motivated by transmission, transportation, manufacturing, and service applications. It shows that both existing component-based and weight-based k-out-of-n models are special cases of the suggested model. An algorithm for evaluating the state probabilities of the multi-state vector-k-out-of-n system is suggested. A numerical example is presented. Gregory Levitin |
IEEE Trans. Reliab. | 1 |
| 2013 | Defending Threshold Voting Systems With Identical Voting UnitsabstractThe threshold voting system consists ofNunits that each provide a binary decision (0 or 1), or abstain from voting. System output is 1 if the number of 1-opting units is at least a pre-specified fraction τ of the number of all non-abstaining units. Otherwise system output is 0. For a system consisting of voting units with given probabilistic output distribution, one can maximize the entire system reliability by choosing a proper threshold value τ. When a system operates in a hostile environment, some units can be destroyed or compromised by an aggressive media, or by a strategic malicious counterpart. One of the ways to enhance voting system survivability is to protect its units from possible attacks. We consider a situation when an attacker and a defender have fixed resources. The defender can protect, and the attacker can attack, a subset of the units. First, we formulate the problem of maximizing survivability of a threshold voting system by proper choice of system threshold and number of protected units, assuming that all the units are attacked. Then we consider a maxmin game in which the defender chooses an optimal system threshold and number of protected units assuming that the attacker chooses the number of attacked units that minimizes the probability of correct system output. Gregory Levitin, Kjell Hausken |
IEEE Trans. Reliab. | 1 |
| 2013 | Reliability of Series-Parallel Systems With Random Failure Propagation TimeabstractThis paper presents an algorithm for evaluating the reliability and performance distribution of complex non-repairable series-parallel multi-state systems with common cause failures caused by propagation of failures in system elements. The failure propagation can have a selective effect, which means that the failures originating from different elements can cause failures of different subsets of system elements. The failure propagation time is assumed to be a random value with a given distribution. The suggested algorithm is based on the universal generating function approach, and a generalized reliability block diagram method (recursive aggregation of pairs of elements and their replacement by an equivalent one). The evaluation procedure is repeated for each combination of elements affected by the common cause failures. Illustrative examples are provided. Gregory Levitin, Liudong Xing, Hanoch Ben-Haim, Yuan-Shun Dai |
IEEE Trans. Reliab. | 1 |
| 2013 | Optimal Allocation of Connecting Elements in Phase Mission Linear Consecutively-Connected SystemsabstractMany continuous transportation systems and communication networks can be modeled as a linear consecutively connected system (LCCS) consisting of n+1 linearly ordered nodes. Some of these nodes contain connection elements (CEs) with different characteristics. Each CE i in its working state provides a connection between the node to which it belongs and L(i) next nodes, the set of nodes that follow the given node. If the system contains any node not connected with any previous node, then it fails. We consider multi-phase LCCS that should perform a sequence of transmission tasks. In each phase, it should provide continuous connection along a specific path of nodes. Different paths can contain the same nodes, which creates statistical dependence across the phases. Different nodes are characterized by specific conditions expressed by different failure acceleration factors affecting the CEs located at these nodes. Thus, an accurate reliability analysis of a multi-phase LCCS must consider the statistical dependencies of CE states across phases, and dynamics in path structures as well as in the failure acceleration factors. In this paper, we propose a method for optimal allocation of CEs in multi-phase LCCS. The method is based on a recursive algorithm for exact system reliability evaluation, and the genetic algorithm for the optimal allocation of CEs to nodes of LCCS. The proposed approach is illustrated using a practical example of wireless sensor networks. Gregory Levitin, Liudong Xing, Yuan-Shun Dai |
IEEE Trans. Reliab. | 1 |
| 2013 | Optimal Allocation of Multistate Components in Consecutive Sliding Window SystemsabstractThis paper considers a system consisting ofnlinearly ordered multistate components. Each component can have different states: from complete failure, up to perfect functioning. A performance rate is associated with each state. The system fails if in each of at leastmconsecutive overlapping groups ofrconsecutive components (windows) the sum of the performance rates of components belonging to the group is lower than a minimum allowable level. It is shown that, in the case of different components, the system reliability depends on their arrangement. The optimal arrangement problem is formulated, and a numerical tool for solving this problem is suggested. The tool uses an extended universal moment generating function technique for system reliability evaluation, and a genetic algorithm for optimization. Examples of system reliability optimization are presented. Yanping Xiang, Gregory Levitin, Yuan-Shun Dai |
IEEE Trans. Reliab. | 2 |
| 2013 | Algorithm for Reliability Evaluation of Nonrepairable Phased-Mission Systems Consisting of Gradually Deteriorating Multistate ElementsabstractReliability analysis of phased-mission systems (PMS) must consider the statistical dependences of element states across different phases as well as changes in system configuration, success criteria, and component behavior. This paper proposes a recursive method for the exact reliability evaluation of PMS consisting of nonidentical independent nonrepairable multistate elements. The method is based on conditional probabilities and the branch-and-bound principle. It is invariant to changes in system structure, demand, and the elements' state transition rates among phases. The main advantage of this method is that it does not require the composition of decision diagrams and can be fully automated. Both analytical and numerical examples are presented to illustrate the application and advantages of the proposed method. The computational performance of the proposed algorithm is illustrated through comprehensive experimentation on the CPU running time of the algorithm. Gregory Levitin, Suprasad V. Amari, Liudong Xing |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2013 | Reliability of Nonrepairable Phased-Mission Systems With Common Cause FailuresabstractPhased-mission systems (PMSs) are systems supporting missions characterized by multiple, consecutive, and nonoverlapping phases of operation. Examples of PMSs abound in many practical applications such as aerospace, nuclear power, and airborne weapon systems. Reliability analysis of a PMS must consider statistical dependence of component states across different phases, as well as dynamics in system structure functions and component behavior. In this paper, we propose a recursive method for exact reliability evaluation of a binary-state or multistate PMS consisting of nonidentical, binary, and nonrepairable elements. The system elements can fail individually or due to common-cause failures (CCFs) caused by some external factors. The proposed method is based on the branch-and-bound principle, and can be fully automated. The method is applicable to PMSs with nonoverlapping or overlapping sets of elements that can fail as a result of CCFs. The method is illustrated using both analytical and numerical examples. Gregory Levitin, Liudong Xing, Suprasad V. Amari, Yuan-Shun Dai |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2013 | Reliability of Systems Subject to Failures With Dependent Propagation EffectabstractThis paper suggests a method for the reliability analysis of binary-state systems subject to component failures with dependent propagation effect. A propagated failure originating from a system component causes extensive damage to the rest of the system. The level of the damage can be dependent upon the status of other system components and the order of the component failures. Such dependent propagation effect typically takes place in systems with some protection mechanism (e.g., firewalls, filters, and antivirus programs) or systems subject to functional dependence behavior where the failure of a system component, referred to as a trigger, causes other components within the same system to become inaccessible or isolated from the system. A combinatorial and analytical method is proposed for addressing the dependent propagation effect in the system reliability analysis. Basics and application of the proposed method are illustrated through analyses of an example of a network system in two different scenarios for propagation effects and an example of a memory system with multiple trigger events. Liudong Xing, Gregory Levitin, Chaonan Wang, Yuan-Shun Dai |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2012 | Recursive Algorithm for Reliability Evaluation of Non-Repairable Phased Mission Systems With Binary ElementsabstractMany practical systems are phased-mission systems (PMS), where the mission consists of multiple, consecutive, and non-overlapping phases of operation. An accurate reliability analysis of a PMS must consider the statistical dependencies of component states across phases, as well as dynamics in system configurations, success criteria, and component behavior. In this paper, we propose a method for exact reliability evaluation of arbitrary binary or multi-state PMS consisting of non-identical binary non-repairable elements. The method is invariant to changes in system structure and demand among missions, and takes into account the time-varying and phase-dependent failure rates and associated cumulative damage effects. The proposed method is based on conditional probabilities, and an efficient recursive formula to compute these probabilities based on branch and bound. The main advantage of this method is that it does not require composition of decision diagrams, and can be fully automated. The method is illustrated using both an analytical example, and a numerical example. Gregory Levitin, Liudong Xing, Suprasad V. Amari |
IEEE Trans. Reliab. | 1 |
| 2012 | Linear $m$ -Gap Sliding Window SystemsabstractThis paper proposes a new model that generalizes the linear multi-state sliding window system. In this model, the system consists ofnlinearly ordered multi-state elements. Each element can have different states spanning from complete failure up to perfectly functioning. A performance rate is associated with each state. The system fails if the gap between any pair of groups ofrconsecutive elements having the cumulative performance lower than a minimum allowable levelWis less thanmgroups ofrconsecutive elements. An algorithm for system reliability evaluation is suggested which is based on an extended universal moment generating function. Examples of evaluating system reliability and elements' reliability importance indices are presented. Yanping Xiang, Gregory Levitin |
IEEE Trans. Reliab. | 2 |
| 2012 | Linear Multistate Consecutively-Connected Systems With Gap ConstraintsabstractThis paper generalizes the linear multistate consecutively-connected system model by introducing allowable gaps. The new model consists of N +1 linearly ordered nodes. Some of these nodes contain statistically independent multistate elements with different characteristics. Each element j can provide a connection between the node to which it belongs and Xjnext nodes, where Xjis a discrete random variable with known probability mass function. The system fails if it contains at least m consecutive nodes not connected with any previous node (m consecutive gaps). An algorithm based on the universal generating function method is suggested for the system reliability evaluation. Illustrative examples are presented. Yanping Xiang, Gregory Levitin, Yuan-Shun Dai |
IEEE Trans. Reliab. | 2 |
| 2012 | k-out-of- $n$ Sliding Window SystemsabstractThis paper proposes a new model that generalizes the linear multistate sliding window system to the case of multiple failures. In this model, the system consists of independent linearly ordered multistate elements. Each element can have different states: from complete failure up to perfect functioning. A performance rate is associated with each state. The system fails if, at least in groups of consecutive elements (windows), the sum of the performance rates of elements belonging to the group is less than a minimum allowable level. An algorithm for system reliability evaluation is suggested which is based on an extended universal moment generating function. Examples of evaluating system reliability and elements' reliability importance indices are presented. Gregory Levitin, Yuan-Shun Dai |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2011 | Guest Editorial: Computational Intelligence in Reliability EngineeringabstractThe five papers in this special section focus on computational intelligence in reliability engineering. Gregory Levitin, Kai-Yuan Cai |
IEEE Trans. Reliab. | 1 |
| 2011 | Linear m -Consecutive k -out-of-r -From- n: F SystemsabstractThis paper proposes a new model that generalizes the linear consecutivek-out-of-r-from-nsystem to the case ofmconsecutive runs ofrelements. In this model, the system consists ofnlinearly ordered statistically independent elements, and fails iff in each of at leastmconsecutive overlapping groups ofrconsecutive elements at leastkelements fail. This system model is motivated by a heating system for moving parts. An algorithm for system reliability evaluation is suggested which is based on an extended universal moment generating function. Examples of system reliability evaluation are presented. Gregory Levitin, Yuan-Shun Dai |
IEEE Trans. Reliab. | 1 |
| 2011 | Combinatorial Algorithm for Reliability Analysis of Multistate Systems With Propagated Failures and Failure Isolation EffectabstractThis paper considers the reliability analysis of multistate systems (MSSs) subject to propagated failure with global effect (PFGE) and failure isolation effect. The PFGE can be caused by an imperfect fault coverage despite the presence of fault-tolerant mechanism or by a destructive effect of failures that originate from some system components on other components. The failure isolation effect is caused by functional dependence among system components, where the failure of some component can prevent the propagation of failures that originate from other components within the same system. Existing approaches for simultaneously addressing PFGE and failure isolation are limited to binary-state systems in which the system and its components exhibit two and only two states: operation or failure. In practice, however, many systems are MSS in which the system and/or its components may exhibit multiple performance levels corresponding to different states ranging from perfect operation to complete failure. In this paper, a separable and combinatorial methodology is proposed for evaluating the reliability of MSS subject to both PFGE and the failure isolation effect. The proposed method has no limitation on the type of time-to-failure distributions for the system components and is applicable to MSS with any arbitrary system structure. Application and advantages of the proposed method are illustrated through a detailed analysis of an example of a multistate memory system. Liudong Xing, Gregory Levitin |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2009 | Redundancy vs. Protection vs. False Targets for Systems Under AttackabstractA system consists of identical elements. The cumulative performance of these elements should meet a demand. The defender applies different actions to reduce a damage associated with system performance reduction below the demand that is caused by an external attack. There are three types of the defensive actions: providing system redundancy (deploying genuine system elements (GE) with cumulative performance exceeding the demand), deploying false elements (FE), and protecting the GE. The defender tries to allocate its total defense resource optimally among the three defensive actions by choosing the total number of GEN, the number of protected GEK, and the number of FEH. The attacker cannot distinguish GE from FE. It chooses the number of elementsQto attack, and attacksQelements at random distributing its resource evenly among the attacked elements. The model considers a non-cooperative two-period minmax game between the defender and the attacker, and presents an algorithm for determining the optimal agents' strategies. Gregory Levitin, Kjell Hausken |
IEEE Trans. Reliab. | 1 |
| 2009 | Redundancy vs. Protection in Defending Parallel Systems Against Unintentional and Intentional ImpactsabstractThis article considers defense resource allocation in a 1-out-of-N system exposed to external intentional impacts caused by malicious attacks, and unintentional impacts caused by naturally-occurring events or technological accidents. The defender distributes its resource between deploying redundant elements, and their protection. Two cases of the protection are considered: general protections that protect from both intentional, and unintentional impacts; and special protections of two different types that protect either against the intentional impacts, or against the unintentional impacts. Different combinations of intentional and unintentional impact sequences are considered. If the unintentional impact occurs first, the strategic attacker has full information about the elements destroyed, and concentrates all its effort on attacking only the survived elements. The vulnerability of each element is determined by contest success functions, between the defender and the attacker, and between the defender and the unintentional impact. A model, and a methodology for finding the defense resource distribution that minimizes the overall system vulnerability are suggested. Illustrative examples of the optimal defense are presented. Gregory Levitin, Kjell Hausken |
IEEE Trans. Reliab. | 1 |
| 2008 | Optimal Structure of Multi-State Systems With Uncovered FailuresabstractThe paper presents a model of multi-state systems in which some groups of elements can be affected by uncovered failures that cause outages of the entire group. It shows that the maximal reliability of such systems can be achieved through a proper balance between two types of task parallelization: parallel task execution with work sharing, and redundant task execution. The paper also suggests a procedure for finding the optimal balance for series-parallel multi-state systems. This procedure is based on a universal generating function technique, and a genetic algorithm. Illustrative examples are presented. Gregory Levitin |
IEEE Trans. Reliab. | 1 |
| 2007 | Optimal task partition and distribution in grid service system with common cause failures
Yuan-Shun Dai, Gregory Levitin, Xiaolong Wang 0003 |
Future Gener. Comput. Syst. | 2 |
| 2007 | Performance and Reliability of Tree-Structured Grid Services Considering Data Dependence and Failure CorrelationabstractGrid computing is a newly emerging technology aimed at large-scale resource sharing and global-area collaboration. It is the next step in the evolution of parallel and distributed computing. Due to the largeness and complexity of the grid system, its performance and reliability are difficult to model, analyze, and evaluate. This paper presents a model that relaxes some assumptions made in prior research on distributed systems that were inappropriate for grid computing. The paper proposes a virtual tree-structured model of the grid service. This model simplifies the physical structure of a grid service, allows service performance (execution time) to be efficiently evaluated, and takes into account data dependence and failure correlation. Based on the model, an algorithm for evaluating the grid service time distribution and the service reliability indices is suggested. The algorithm is based on Graph theory and probability theory. Illustrative examples and a real case study of the BioGrid are presented. Yuan-Shun Dai, Gregory Levitin, Kishor S. Trivedi |
IEEE Trans. Computers | 2 |
| 2007 | Optimal Resource Allocation for Maximizing Performance and Reliability in Tree-Structured Grid ServicesabstractThe paper considers a grid computing systems in which the resource management systems (RMS) can divide service tasks into execution blocks (EB), and send these blocks to different resources. To provide a desired level of service reliability, the RMS can assign the same EB to several independent resources for parallel (redundant) execution. According to the optimal schedule for service task partition, and distribution among resources, one can achieve the greatest possible expected service performance (i.e. least execution time), or reliability. For solving this optimization problem, the paper suggests an algorithm that is based on graph theory, Bayesian approach, and the evolutionary optimization approach. A virtual tree-structure model is constructed in which failure correlation in common communication channels is taken into account. Illustrative examples are presented. Yuan-Shun Dai, Gregory Levitin |
IEEE Trans. Reliab. | 2 |
| 2007 | Optimal Defense Strategy Against Intentional AttacksabstractThis paper presents a generalized model of damage caused to a complex multi-state series-parallel system by intentional attack. The model takes into account the defense strategy that presumes separation and protection of system elements. The defense strategy optimization methodology is suggested, based on the assumption that the attacker tries to maximize the expected damage of an attack. An optimization algorithm is presented that uses a universal generating function technique for evaluating the losses caused by system performance reduction, and a genetic algorithm for determining the optimal defense strategy. Illustrative examples of defense strategy optimization are presented Gregory Levitin |
IEEE Trans. Reliab. | 1 |
| 2006 | Reliability and performance of tree-structured grid servicesabstractGrid computing is a new emerging technology aiming at large-scale resource sharing, and global-area collaboration. It is a next step in an evolution of parallel and distributed computing. Due to the large scale and complexity of the grid system, its performance and reliability are difficult to model, analyse, and evaluate. This paper presents a model that relaxes some assumptions unsuitable for grid computing systems that have been made in the existed works studying the distributed systems. The paper proposes a virtual tree model of the grid service. This model simplifies the physical structure of a grid service, allows service performance (execution time) to be estimated, and takes into account the common cause failures in communication channels. Based on the model, an algorithm for evaluating the grid service performance distribution and the service reliability indices is suggested. The algorithm is based on graph theory, and Bayesian analysis. Illustrative examples are presented in which the results of the suggested algorithm are compared with simulation results. Yuan-Shun Dai, Gregory Levitin |
IEEE Trans. Reliab. | 2 |
| 2006 | Reliability and Performance of Star Topology Grid Service With Precedence Constraints on Subtask ExecutionabstractThe paper considers grid computing systems with star architectures in which the resource management system (RMS) divides service tasks into subtasks, and sends the subtasks to different specialized resources for execution. To provide the desired level of service reliability, the RMS can assign the same subtasks to several independent resources for parallel execution. Some subtasks cannot be executed until they have received input data, which can be the result of other subtasks. This imposes precedence constraints on the order of subtask execution. The service reliability & performance indices are introduced, and a fast numerical algorithm for their evaluation given any subtask distribution is suggested. Illustrative examples are presented Gregory Levitin, Yuan-Shun Dai, Hanoch Ben-Haim |
IEEE Trans. Reliab. | 1 |
| 2004 | Consecutive k-out-of-r-from-n system with multiple failure criteriaabstractThis paper proposes a new model which generalizes the linear consecutive k-out-of-r-from-n system to the case of multiple failure criteria. In this model, the system consists of n linearly ordered elements, and fails iff in any group of r/sub 1/,...,r/sub H/ consecutive elements less than k/sub 1/,...,k/sub H/ elements respectively are in a working state. An algorithm for system reliability evaluation is suggested which is based on an extended universal moment generating function. Examples of system reliability evaluation are presented. Gregory Levitin |
IEEE Trans. Reliab. | 1 |
| 2003 | Optimal allocation of multistate elements in a linear consecutively-connected systemabstractA linear consecutively-connected system consists of N+1 linear ordered positions. M s-independent multistate elements with different characteristics are to be allocated to the first N positions. Each element can provide a connection between the position to which it is allocated and the next few positions. The reliability of the connection for any given element depends on: (1) the position to which it is allocated; and (2) the number of positions it connects. The system fails if the first position (source) is not connected with the N+1 position (sink). An algorithm based on the universal generating function method is suggested for the linear consecutively-connected system reliability determination. This algorithm can handle cases in which any number of multistate elements are allocated in the same position while some positions remain empty. In many cases, such uneven allocation provides greater system reliability than the even one. A genetic algorithm is used as an optimization tool to solve the optimal element-allocation problem. Gregory Levitin |
IEEE Trans. Reliab. | 1 |
| 2003 | Reliability evaluation for acyclic transmission networks of multi-state elements with delaysabstractAcyclic transmission networks (ATN) consist of a number of positions in which s-independent multi-state elements (ME) capable of receiving and/or sending a signal are allocated. Each network has: a root position where the signal source is located, a number of terminal positions that can only receive a signal, and a number of intermediate positions containing ME that retransmit the received signal to some other positions. The signal propagation is allowed only in direction of the terminal nodes, which avoids cycles in the network. Each ME that is located in a nonterminal node can have different states determined by a set of nodes receiving the signal directly from this ME. The probability of each state is assumed to be known for each ME. The signal retransmission process is associated with delays. The system fails if the signal generated at the first position (source) cannot reach the terminal nodes within a specified time. An algorithm for ATN reliability evaluation is suggested; it is based on an extended universal generating function method that uses a state enumeration approach. A procedure for recurrent derivation of probabilistic distribution of a vector representing delays in ATN is described. A technique for computational complexity reduction is developed. Gregory Levitin |
IEEE Trans. Reliab. | 1 |
| 2003 | Linear multi-state sliding-window systemsabstractThis paper proposes a new model that generalizes the consecutive k-out-of-r-from-n:F system to the multi-state case. In this model (linear multi-state sliding window system), the system consists of n linearly ordered multi-state elements. Each element can have various states: from complete-failure up to perfect-functioning. A performance rate is associated with each state. The sliding window system fails if the sum of the performance rates of any r consecutive multi-state elements is lower than a minimum allowable level. An algorithm for system reliability evaluation is based on an extended universal moment-generating function. Examples of system reliability evaluation are presented. Gregory Levitin |
IEEE Trans. Reliab. | 1 |
| 2003 | Reliability of multi-state systems with two failure-modesabstractSystems with two failure modes (STFM) consist of devices, which can fail in either of two modes. For example, switching systems can not only fail to close when commanded to close but can also fail to open when commanded to open. This paper considers systems consisting of different elements characterized by nominal performance level in each mode. Such systems are multi-state because they have multiple performance levels in both modes, depending on the combination of elements available at the moment. The system availability is defined as the probability of satisfaction of given constraints imposed on system performance in both modes. The paper suggests reliability measures for multi-state systems with 2 failure modes, and presents a procedure for evaluating these measures. The procedure is based on the use of a universal moment generating function (UMGF). It allows one to estimate availability and s-expected performance of complex systems with series-parallel and bridge topology. Basic UMGF technique operators are developed for two types of systems, based on transmitting-capacity and on operation-time. Gregory Levitin |
IEEE Trans. Reliab. | 1 |
| 2001 | Redundancy optimization for multi-state system with fixed resource-requirements and unreliable sourcesabstractThis paper considers a redundancy optimization problem for a multi-state system of: (1) elements that consume a fixed amount of resources to perform their task, and (2) a number of resource generating subsystems. The algorithm finds the optimal system-structure, subject to availability constraints, by choosing system elements from a list of available equipment. Each element is characterized by its productivity, availability, and cost. Elements of the main producing subsystem also have their specific resource consumption limitations. The objective is to minimize the sum of investment costs while satisfying demand, represented by a cumulative demand curve, with given probability. To solve the problem, a genetic algorithm is used for optimization. The procedure, based on the universal generating function, is used to evaluate the system-availability while assuming that the working elements of the main producing subsystem are chosen in such a way that the total system performance rate is maximal under given resource constraints. Examples demonstrate how to obtain the optimal structures of a simple 2-level system for various availability constraints. Gregory Levitin |
IEEE Trans. Reliab. | 1 |
| 2001 | Incorporating common-cause failures into nonrepairable multistate series-parallel system analysisabstractThis paper adapts the universal generating function method of multistate system reliability analysis to incorporate common-cause failures (CCF). An implicit 2-stage approach is used. In stage #1, a polynomial representation of system-output performance distribution is obtained without considering system elements which are subject to CCF. In stage #2, the contribution of CCF is correctly included based on information stored in vector-indicators that describe states of system elements belonging to various common-cause groups. A straightforward procedure is suggested for evaluating reliability functions of nonrepairable series-parallel multistate systems with CCF. This procedure allows the reliability functions to be obtained numerically. Examples are given. Gregory Levitin |
IEEE Trans. Reliab. | 1 |
| 2000 | Multistate series-parallel system expansion-scheduling subject to availability constraintsabstractThis paper addresses the multistage expansion problem for multistate series-parallel systems. The study period is divided into several stages. At each stage the demand distribution is predicted in the form of a cumulative demand curve. The additional elements chosen from a list of available products can be included into any system-component at any stage to increase the total system capacity and/or reliability. Each element is characterized by its capacity (productivity), availability, and cost. The objective is to minimize the sum of costs of the investments over the study period while satisfying reliability constraints at each stage. To solve the problem, a genetic algorithm is used as an optimization tool. The solution encoding technique allows the genetic algorithm to manipulate integer strings representing multistage expansion planes. A solution quality index comprises both reliability and cost estimations. The procedure based on the universal generating function is used for evaluating the availability of multistate series-parallel systems. An example illustrates finding the optimal expansion plan for a coal-transportation system of a power station. Gregory Levitin |
IEEE Trans. Reliab. | 1 |