VLDB 2026 Research / reviewers in the wild / expert
Fumio Machida
dblp:60/6586
· DBLP profile ↗
60ranked-venue papers
18as first author
28since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 22 · 4 first-author · 12 since 2021Security and privacy · 21 · 5 first-author · 6 since 2021Systems, architecture and hardware · 13 · 6 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Dynamic adaptive task offloading for UAV-based road traffic monitoringabstractUnmanned Aerial Vehicles (UAVs) are increasingly used for road traffic monitoring due to their mobility and wide-area coverage. However, their limited onboard resources make real-time video analysis challenging under dynamic traffic conditions. To overcome this, computational task offloading to nearby fog nodes is often employed. The main challenge lies in deciding when to process locally or offload, as both traffic and computational load vary continuously. Existing heuristic-based approaches are lightweight but rely on fixed thresholds, leading to unstable switching and degraded performance under fluctuating conditions. Meanwhile, Deep Reinforcement Learning (DRL)–based methods can adaptively optimize offloading but require extensive training and high computational costs, limiting their practicality on UAVs. To address this challenge, we propose Dynamic Vehicle Density-aware Offloading (DVDOffload), an adaptive task offloading technique designed to maximize performance and resource efficiency by adapting the offloading decision to road traffic conditions. The proposed method uses vehicle density as the primary workload indicator and dynamically adjusts offloading thresholds using an Exponential Moving Average (EMA) to ensure adaptive and stable decisions. Experimental results in multiple realistic traffic scenarios show that DVDOffload achieves higher accuracy, faster processing, and lower resource consumption compared to several baseline heuristic and DRL-based approaches in the evaluated UAV–fog traffic monitoring system. Mohammad Dwipa Furqan, Fumio Machida, Ermeson Carneiro de Andrade |
Future Gener. Comput. Syst. | 2 |
| 2026 | Distributed Performability Optimization for Multi-UAV Road Traffic Monitoring
Qingyang Zhang 0007, Koji Noshiro, Mohammad Dwipa Furqan, Koji Hasebe, Fumio Machida |
ICAART (1) | 5 |
| 2026 | SPADE: Simulator-assisted Performability Design for UAV-based monitoring systemsabstractAs Uncrewed Aerial Vehicles (UAV) have been used widely in a variety of real-world monitoring applications, quality design of UAV-based monitoring systems becomes an emergent challenge as it involves complex trade-offs among several performance criteria. While analytical models have been used for performance analysis of UAV systems, they often rely on hypothetical parameter values due to difficulty in accessing real-world systems, resulting in a gap between theory and practice. To fill this gap, this paper proposes SPADE (Simulator-assisted PerformAbility Design methodology for UAV-based Systems), an approach that integrates performance profiling with a realistic flight scenario generated by a UAV simulator and model-based performance analysis. We demonstrate the application of SPADE through a case study that focuses on designing a UAV-based ecological monitoring system using an object detection algorithm (YOLOv5). Our analysis explores key trade-offs among several quality metrics, including detection accuracy, performance, energy consumption, and service availability. Using Stochastic Petri Nets, we conduct numerical evaluations, with baseline parameter values estimated from performance profiling on an emulated computing device. Experimental results using YOLOv5 provide valuable insights into how image resolution and computation modes impact UAV-based system performance and availability. These findings offer practical guidance for improving UAV system design. Qingyang Zhang 0007, Fumio Machida, Ermeson Carneiro de Andrade |
Future Gener. Comput. Syst. | 2 |
| 2026 | Machine learning for software aging detection: A systematic mapping study
Rafael José Moura, Maria Gizele Nascimento, Fumio Machida, Domenico Cotroneo, Ermeson Carneiro de Andrade |
J. Syst. Softw. | 3 |
| 2026 | Experimental investigation of memory-related software aging in LLM systemsabstractLarge Language Models (LLMs) have been increasingly adopted in a wide range of applications, many of which require long-running inference processes. However, these systems may be subject to software aging phenomena, leading to progressive performance degradation and potential failures. In this work, we experimentally investigate memory-related software aging in LLM inference. We performed 48-hour experiments with three open-source models (Pythia, OPT, and GPT-Neo) under low, medium, and high workloads, monitoring memory consumption at both system and process levels. Using the Mann–Kendall test and Sen’s slope estimator, we observed monotonic growth in RAM usage across all models on Central Processing Units (CPUs), with OPT presenting the steepest slopes. Process-level analysis further revealed that LLM processes were the primary contributors to memory growth, along with background services. Additionally, we conducted identical experiments on Graphics Processing Units (GPUs). Unlike the experiments without a GPU, GPU-based experiments revealed bounded oscillations and abrupt resets likely due to driver-level memory management, while host RAM and process-level monitoring still revealed clear symptoms of aging. These findings demonstrate that software aging manifests differently across execution environments, reinforcing the need for environment-specific monitoring approaches. César Augusto Ribeiro dos Santos, Fumio Machida, Ermeson Carneiro de Andrade |
J. Syst. Softw. | 2 |
| 2026 | On Metaverse Application Dependability AnalysisabstractMetaverse as-a-Service (MaaS) enablesMetaverse tenants to execute theirAPPlications (MetaAPP) by allocating Metaverse resources in the form of Metaverse service functions (MSF). Usually, each MSF is deployed in a virtual machine (VM) for better resiliency and security. However, these MSFs along with VMs and virtual machine monitors (VMM) running them will encounter software aging after prolonged continuous operation. Then, there is a decrease in MetaAPP dependability, namely, the dependability of the MSF chain (MSFC), consisting of MSFs allocated to MetaAPP. This paper aims to investigate the impact of both software aging and rejuvenation techniques on MetaAPP dependability in the scenarios, where both active components (MSF, VM and VMM) and their backup components are subject to software aging. We develop a hierarchical model to capture behaviors of aging, failure, and recovery by applying Semi-Markov process and reliability block diagram. Numerical analysis and simulation experiments are conducted to evaluate the approximation accuracy of the proposed model and dependability metrics. We then identify the key parameters for improving the MetaAPP/MSFC dependability through sensitivity analysis. The investigation is also made about the influence of various parameters on MetaAPP/MSFC dependability. Yingfan Zong, Jing Bai 0009, Xiaolin Chang, Fumio Machida, Yingsi Zhao |
IEEE Trans. Cloud Comput. | 4 |
| 2025 | Multi-version Machine Learning and Rejuvenation for Resilient Perception in Safety-critical SystemsabstractMachine learning (ML) has become a crucial component in safety-critical systems, such as those used in autonomous vehicle perception. However, the correctness and, therefore, the safety of these systems can be compromised by out-of-distribution data, accidental faults, and security breaches. This paper investigates using a replicated ML architecture to mitigate the risks associated with complex single-points-of-failure. Additionally, it explores the application of rejuvenation to sustain healthy majorities when facing persistent threats. We evaluate the output reliability of the proposed architecture in two case studies: traffic sign detection and perception for autonomous driving. We adopt models and reliability functions, validating our findings using realistic data sets and fault injection experiments. We also evaluate driving safety using the proposed architecture in the CARLA simulator. Our results show that our models can present a good generalization and multi-version ML with proactive rejuvenation can improve correctness and, thus, safety despite faults and cyberattacks. Julio Mendonca 0001, Fumio Machida, Marcus Völp |
DSN | 3 |
| 2025 | Vehicle Density-Aware Adaptive Offloading for UAV-Based Road Traffic MonitoringabstractUnmanned Aerial Vehicles (UAVs) have emerged as a transformative technology for real-time road traffic monitoring, offering enhanced efficiency and responsiveness to modern traffic management systems. However, the resource limitations of UAVs and the dynamic nature of traffic densities present significant challenges for continuous operation. To address these constraints, this study proposes a vehicle-density-aware adaptive offloading mechanism that dynamically alternates between local processing and task offloading to fog nodes, based on real-time traffic conditions. The mechanism operates in three distinct modes: Low-CPU Mode for low vehicle density, Full Offloading Mode for moderate density, and Local Processing Mode for high-density scenarios. Preliminary results reveal that the proposed VD-aware adaptive offloading mechanism effectively balances performance, resource efficiency, and communication costs. It maintains competitive accuracy, optimizes throughput, and dynamically manages CPU utilization and communication overhead. These findings highlight the adaptability and efficiency of the proposed mechanism, making it an ideal solution for UAV-based road traffic monitoring in dynamic and resource-constrained environments. Mohammad Dwipa Furqan, Fumio Machida, Ermeson Carneiro de Andrade |
ICFEC | 2 |
| 2025 | Exploiting the Availability-Continuity Trade-off in Imperfect Retraining of Machine Learning SystemsabstractMachine Learning Systems (MLSs) often combine diverse models to achieve complex objectives but face performance degradation due to dataset shifts. Regular performance monitoring and model retraining are essential to mitigate this risk. However, model retraining may not always fully restore the system’s performance, which is known as the imperfect retraining problem. This study examines model retraining policies to maintain MLS performance in the face of imperfect retraining. First, we demonstrate real-world applications that encounter imperfect retraining in computer vision and natural language processing tasks. Next, we theoretically analyze two retraining policies, progressive and conservative, to counteract performance degradation. We formulate the dynamics of model degradation and retraining using semi-Markov processes and quantitatively evaluate service availability and continuity, which measures how long the service can maintain its performance. The numerical analysis results demystify the notable trade-off between service availability and continuity, guiding a proposed retraining strategy to better sustain MLS performance. Zhengji Wang, Fumio Machida |
ISSRE | 2 |
| 2025 | Understanding Container-Based Services Under Software Aging: Dependability and Performance ViewsabstractContainer technology, as the key enabler behind microservice architectures, is widely applied in Cloud and Edge Computing. A long and continuous running of operating system (OS) hosting container-based services can encounter software aging that leads to performance deterioration and even causes system failures. OS rejuvenation techniques can mitigate the impact of software aging but the rejuvenation trigger interval needs to be carefully determined to reduce the downtime cost due to rejuvenation. This paper proposes a comprehensive semi-Markov-based approach to quantitatively evaluate the effect of OS rejuvenation on the dependability and the performance of a container-based service. In contrast to the existing studies, we neither restrict the distributions of time intervals of events to be exponential nor assume that backup resources are always available. Through the numerical study, we show the optimal container-migration trigger intervals that can maximize the dependability or minimize the performance of a container-based service. Jing Bai 0009, Xiaolin Chang, Fumio Machida, Kishor S. Trivedi |
IEEE Trans. Sustain. Comput. | 3 |
| 2024 | Exploiting Transformer Models in Three-Version Image Classification SystemsabstractMachine learning (ML) models are extensively employed in a wide range of real-world applications, including safety-critical ones. The reliability of ML application systems is a critical concern, particularly in situations where incorrect system outputs lead to severe consequences. This paper aims to enhance the reliability of ML-based image classification systems by exploiting a transformer-based model with non-transformer models in two distinct architectural frameworks, specifically the Majority Voting (MV) and Recovery Block (RB). Instead of relying on a single ML prediction, the proposed architectures leverage multiple ML models, executed simultaneously for the same input, to improve system output reliability. Our exper-imental results on image classification tasks show that three-version systems employing transformer models exhibit reliability enhancements in both MV and RB architectures. Moreover, considering performance overhead imposed by transformer models, we evaluated the response times as a performance metric of ML systems. The evaluation results show that RB architecture proves to have shorter response times than MV architecture. Shamima Afrin, Fumio Machida |
COMPSAC | 2 |
| 2024 | Maintaining Performance of a Machine Learning System Against Imperfect RetrainingabstractMachine learning systems (MLS) often consist of diverse machine learning models to attain demanding and complex objectives. Despite their advanced functionalities, MLSs encounter the inevitable challenge of performance deterioration caused by distribution changes often referred to as dataset shift. When models encountering dataset shift, retraining machine learning model with new datasets is essential to restore the performance. However, model retraining is not always perfect due to component entanglement, resulting in failures to maintain the required level of performance. To address this issue, this paper investigates the impact of imperfect retraining on the overall performance of MLSs, and accordingly propose two maintenance policies, progressive and conservative retraining policies. We consider an MLS consisting of two sequentially-dependent machine learning models and develop continuous-time Markov chains capturing the dynamics of performance degradation and retraining of machine learning models. The results of parametric sensitivity analysis demonstrate that the progressive retraining policy and conservative retraining policy could provide higher service availability in different conditions. Zhengji Wang, Fumio Machida |
COMPSAC | 2 |
| 2024 | Performability Modeling and Analysis for Real-Time Object Detection on UAV SystemsabstractWith the widespread application of Uncrewed Aerial Vehicles (UAVs) in various real-world surveillance scenarios, the quality analysis of UAV-based monitoring systems has become an emergent challenge. Previous model-based studies often made theoretical assumptions to estimate the performance and availability of UAV computing systems, without detailed considerations of the interaction between UAVs and fog computing nodes, as well as computation steps of object detection algorithms. In this paper, we propose Stochastic Reward Nets (SRNs) to capture computational behavior and analyze performance, availability, and performability metrics of a UAV system that utilizes computation offloading. In order to obtain more realistic parameters for model-based analysis, we conduct empirical experiments using an edge computing device to emulate real-time object detection on a UAV. We measure the throughput of real-time object detection process in three stages by experiments that are fed into parameters for numerical analysis on the proposed model. Through the sensitivity analysis, we demonstrate the impact of different computation modes and video resolutions on performance and availability metrics, providing insights for improving UAV system design and operation. Qingyang Zhang 0007, Fumio Machida, Ermeson Carneiro de Andrade |
COMPSAC | 2 |
| 2024 | Empirical Architecture Comparison of Two-input Machine Learning Systems for Vision TasksabstractAs machine learning models have been deployed in many vision systems, including autonomous vehicles and robots, designing architectures for machine learning systems (MLSs) has emerged as a critical concern. Previous studies have shown that enhancing the reliability of MLS outputs can be achieved by comparing multiple inference results on distinct inputs. Nevertheless, the architectures facilitating multiple inferences incur non-negligible performance overhead and energy consumption that have been less investigated. This article delves into the trade-offs among reliability, performance, and energy efficiency of architectures for two-input MLSs through real experiments conducted on image classification and object detection tasks. Specifically, we scrutinize the comparison between parallel- and shared-type architectures of two-input MLSs for vision tasks. The experiments confirm that the shared-type architecture can achieve a shorter response time and smaller energy consumption by using a shared machine learning module for both image classification and object detection tasks. However, the parallel-type architecture can benefit the redundant machine learning modules for improving throughput and fault tolerance. Our empirical results also show the service time distributions of image classification and object detection tasks fit well with a log-normal distribution and a mixture of the Gaussian model, respectively. Kazuya Wakigami, Fumio Machida, Tuan Phung-Duc |
Formal Aspects Comput. | 2 |
| 2024 | Assuring Autonomy of UAVs in Mission-critical Scenarios by Performability Modeling and AnalysisabstractUncrewed Aerial Vehicles (UAVs) have been used in mission-critical scenarios such as Search and Rescue (SAR) missions. In such a mission-critical scenario, flight autonomy is a key performance metric that quantifies how long the UAV can continue the flight with a given battery charge. In a UAV running multiple software applications, flight autonomy can also be impacted by faulty application processes that excessively consume energy. In this article, we propose Flight Autonomy Assurance as a framework to assure the autonomy of a UAV considering faulty application processes through performability modeling and analysis. The framework employs hierarchically configured stochastic Petri nets, evaluates the performability-related metrics, and guides the design of mitigation strategies to improve autonomy. We consider a SAR mission as a case study and evaluate the feasibility of the framework through extensive numerical experiments. The numerical results quantitatively show how autonomy is enhanced by offloading and restarting faulty application processes. Ermeson Carneiro de Andrade, Fumio Machida |
ACM Trans. Cyber Phys. Syst. | 2 |
| 2023 | Characterizing Reliability of Three-version Traffic Sign Classifier System through Diversity MetricsabstractThe N-version machine learning (ML) system is an architecture approach to enhance the reliability of ML system outputs by exploiting ML model diversity and input data diversity. While existing studies theoretically show the relation between diversity metrics and system reliability, there is a shortage of empirical studies validating reliability models with diversity parameters in real datasets. In this paper, focusing on traffic sign recognition tasks, we empirically analyze the impact of diversity parameter estimations for predicting the reliability of three-version traffic sign classifier systems. Using five real-world traffic sign datasets, we confirm that the three-version architecture effectively enhances system reliability by applying diverse models and diversified input images. Then, we estimate the diversity parameters and apply them to variants of reliability prediction models. The prediction residuals between the observed reliability and the predicted reliability are mostly less than 0.017 across all data sets, which is half of the residual achieved by the conventional prediction model, except for the architecture of a single model with triple input. As the estimated values of diversity parameters tend to be stable with a relatively small number of samples, we consider that the reliability prediction models using diversity parameters are useful in the early-stage design of ML systems. Fumio Machida |
ISSRE | 2 |
| 2023 | Reliability and Performance Evaluation of Two-input Machine Learning SystemsabstractThe multiple-input machine learning system (MLS) is a system architecture exploiting data diversity to improve the output reliability of the system by comparing prediction results on multiple input data. While the output reliability is enhanced by redundancy, the architecture imposes additional costs and non-negligible processing overheads. The performance of multiple-input MLSs has been theoretically investigated in the previous study using queueing analysis. However, it is little known how real MLSs are impacted by the multiple predictions and comparison processes needed in the architecture. In this paper, we implement two-input MLSs in two different configurations, a parallel type architecture and a shared type architecture, and evaluate the reliability, performance, and energy consumption of the system by experiments. Our empirical results unveil several advantages of the shared type architecture that can suppress the increases in response time and energy consumption by using a shared machine learning module for predictions of two inputs. We also compare the results of the performance simulation of two-input MLS with the empirical results. While we confirm the effectiveness of the simulation, we also find some gaps in the real observations. For example, we observe that the inference time distribution fits well in the log-normal distribution rather than the exponential distribution assumed in the simulation. Our findings could be useful for developing performance models for multiple-input MLSs. Kazuya Wakigami, Fumio Machida, Tuan Phung-Duc |
PRDC | 2 |
| 2023 | Performability analysis of adaptive drone computation offloading with fog computing
Fumio Machida, Qingyang Zhang 0007, Ermeson Carneiro de Andrade |
Future Gener. Comput. Syst. | 1 |
| 2023 | Understanding NFV-Enabled Vehicle Platooning Application: A Dependability ViewabstractThis paper aims to use analytical modeling technique to quantitatively study the dependability of Vehicle Platooning Application, which consists of Multiple Sub-Services (VPP-MSS) to achieve its functionality. Each sub-service (SS), based on network function virtualization technology, is executed in a container. Both SSes and OSes which SSes run on can suffer from software aging after a long and continuous running, reducing VPP-MSS dependability. Rejuvenation techniques are usually used to combat software aging, but they require the support of backup components. Quantitative study of VPP-MSS dependability enables in-depth understanding of the effectiveness of rejuvenation techniques based on analytical models. In contrast to the existing studies, we develop a semi-Markov process (SMP) model to jointly analyze the impact of rejuvenation technique trigger intervals (RTTIs), backup components’ behaviors, time-dependent interactions between various behaviors and the number of active SSes deployed on an OS on the effectiveness of rejuvenation technique. Sensitivity analysis helps identify key parameters for improving the dependability of VPP-MSS. Extensive numerical experiments demonstrate the necessity of considering backup components’ behaviors and investigating non-exponentially distributed failure times. We also determine both the optimal RTTI combination and the optimal combination of SSes and OSes, which can maximize VPP-MSS dependability. Jing Bai 0009, Xiaolin Chang, Fumio Machida, Kishor S. Trivedi |
IEEE Trans. Cloud Comput. | 4 |
| 2023 | A Comparative Analysis of Software Aging in Image Classifiers on Cloud and EdgeabstractImage classifiers for recognizing real-world objects are widely used in the Internet of Things (IoT) and Cyber-Physical Systems(CPSs). A classifier is trained offline by machine learning algorithms with training data sets, and then it is deployed on a cloud or an edge computing system for online label predictions. As the classifier's performance depends on the underlying software infrastructure, it may degrade over time due to software faults causing software aging. In this paper, we address this issue and experimentally investigate software aging observed in an image classification system that continuously runs on cloud and edge computing environments. We apply several statistical techniques to analyze degradation trends in the systems under stress tests. Our statistical trend analysis confirms the degradation trends in the throughput as well as the available memory resources both in the cloud and the edge environments. Contrary to our expectation, the edge computing environment under test had much less impact on the performance degradation than our cloud environment when the workload is high, although the latter one has four times larger allocated memory resources. We also show that the observed performance degradation trends are associated with the memory usage of specific processes by performing correlation analysis. Ermeson Carneiro de Andrade, Roberto Pietrantuono, Fumio Machida, Domenico Cotroneo |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2023 | Impact of Service Function Aging on the Dependability for MEC Service Function ChainabstractThe Multi-access Edge Computing (MEC) and Network Function Virtualization (NFV) integrated architecture is a key enabling platform for 5G to run multiple customized services in the form of service function chain (SFC) configured as an ordered set of service functions (SFs). However, memory-related software aging in the SF that can be exploited by attackers becomes a new threat to the dependability of MEC-SFC services. To provide dependable MEC-SFC services, proactive rejuvenation techniques to counteract the SF aging problem are essential. In this paper, we develop a semi-Markov model to quantitatively investigate the transient availability and steady-state dependability (availability and reliability) of MEC-SFC services. Our model enables the analysis of a MEC-SFC with any number of SFs, and can capture complex time-dependent behaviors of aging, failure, and recovery. The approximate accuracies of the presented model on dependability measures are comprehensively evaluated through comparative studies with simulation experiments. We then detect potential bottlenecks for a MEC-SFC system through sensitivity analysis and further analyze the impact of event-time interval distributions on steady-state dependability. Finally, we investigate the transient behaviors of a MEC-SFC service when varying system parameters during MEC-SFC operation. Jing Bai 0009, Xiaolin Chang, Fumio Machida, Lili Jiang 0004, Zhen Han 0001, Kishor S. Trivedi |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2023 | Model-Driven Dependability Assessment of Microservice Chains in MEC-Enabled IoTabstractMulti-accessedgecomputing (MEC)-enabledInternetofThings (IoT) is considered as a promising paradigm to deliver computation-intensive and delay-sensitive services to users. IoT service requests can be served by multiplemicroservices (MSs) that form a chain, called amicroservicechain (MSC). However, the high complexity of MSs and security threats in MEC-enabled IoT pose new challenges to MSC dependability. Proactive rejuvenation techniques can mitigate the impact of resource degradation of MSs and hostoperatingsystems (OSes) executing them. In this article, we develop a multi-dimensional semi-Markov model to investigate the effectiveness of proactive rejuvenation techniques in improving the dependability (availability and reliability) of a dynamic and heterogeneous MSC. The results of numerical experiments firstly reveal how MSs can be effectively combined, in different deployment configurations, with host OSes to improve MSC dependability, secondly jointly optimize the rejuvenation trigger intervals of host OS and MSs running on it, and finally show the impact of time-varying parameters. We also identify the bottlenecks for MSC dependability improvement by sensitivity analysis, and give the ranges of important parameter values guaranteeing five-nines availability. In addition, the superiority of our model is demonstrated by comparison with the continuous-time Markov chain model. Jing Bai 0009, Xiaolin Chang, Fumio Machida, Kishor S. Trivedi |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | How Data Diversification Benefits the Reliability of Three-version Image Classification SystemsabstractRecently, we have witnessed an increased use of systems employing Machine learning (ML) models. The dependability of such systems is significantly impacted by the inference results of ML models that may not always be correct. Nversion ML systems can be adopted for improving output reliability by detecting or correcting errors by combining multiple inference results. In this study, we focus on a three-version ML system that combines the inference results of an ML model on diversified input data to make the system reliable for ML image classification tasks. The three-version image classification system uses only one deep-learning classifier while generating three inference results in response to diversified inputs. We use several image transformation methods to generate diversified input data for inferences. The system reliability is evaluated by the coverage of errors and the certainty of accurate predictions, as those metrics are decision method agnostic. We confirm that appropriately combined inference results from diversified data can increase the coverage of errors and improve reliability while maintaining the certainty of accurate predictions. In order to search for such effective combinations of diversification methods, we propose the neuron coverage improvement rate (NCIR) as an indicator of data diversity. Through the experiments, we show that the NCIR tends to have correlations with the coverage of errors and the certainty of accurate predictions, indicating the usefulness of the indicator. Mitsuho Takahashi, Fumio Machida |
PRDC | 2 |
| 2022 | An Empirical Study on Software Aging of Long-Running Object Detection AlgorithmsabstractEfficient and effective object detection is a key problem in Computer Vision. Numerous object detection algorithms have been developed, whose aim is to achieve two conflicting goals, namely accuracy and efficiency, while being executed in real-time with high robustness. Many of these algorithms must run for an extended period of time, i.e., in video surveillance or in self-driving cars – a working condition that make them subject to the risk of software aging.In this work, we focus on evaluating several object detection algorithms to understand if and to what extent they are affected by software aging. A measurement-based aging approach was adopted, with a series of long-running tests and subsequent data analysis. The results report significant trends of performance degradation, sometimes leading to aging-related failures, as well as memory consumption trends, which turned out to be the main issue across all the experiments. Roberto Pietrantuono, Domenico Cotroneo, Ermeson Carneiro de Andrade, Fumio Machida |
QRS | 4 |
| 2022 | Quantitative understanding serial-parallel hybrid sfc services: a dependability perspective
Jing Bai 0009, Xiaolin Chang, Fumio Machida, Zhen Han 0001, Yang Xu 0013, Kishor S. Trivedi |
Peer-to-Peer Netw. Appl. | 3 |
| 2022 | Analysis of Optimal File Placement for Energy-Efficient File-Sharing Cloud Storage SystemabstractPopular data concentration is a widely accepted storage energy-saving technique which places frequently-accessed data on a small subset of hard disks and spins-down other infrequently-accessed disks. Many previous studies use intuitive heuristic algorithms for data placement that promote the imbalance in the access frequencies across hard disks. However, the relevance and the optimality of such file placements have not been rigorously investigated. In this paper, we formally define the energy-saving file placement problem under the capacity and performance constraints as a combinatorial optimization problem and show the theory of the optimal file placement where the file access rates in the next period are given. Our analysis based on a stochastic process of disk state transitions gives the theoretical support for the common heuristic placement method. To examine the effectiveness of the optimal file placement, we experimentally evaluate the energy-efficiency of a test storage system using the file access rates generated from the real access traces from Flickr. The experimental results show that the energy consumption can be reduced by 31.8 percent with the optimal file placement compared to the evenly distributed file placement. We also conduct simulation experiments to confirm the energy-saving impacts in larger-scale storage systems. Fumio Machida, Koji Hasebe, Hirotake Abe, Kazuhiko Kato |
IEEE Trans. Sustain. Comput. | 1 |
| 2021 | PA-Offload: Performability-Aware Adaptive Fog Offloading for Drone Image ProcessingabstractSmart drone systems have built-in computing resources for processing real-world images captured by cameras to recognize their surroundings. Due to limited resources and battery constraints, resource-intensive image processing tasks cannot always run on drones. Thus, offloading computation tasks to any available node in a fog computing infrastructure can be considered as a promising solution. An important challenge when applying fog offloading is deciding when to start or stop offloading, taking into account performance and availability impacts under varying workloads and communication link states. In this paper, we present a performability-aware adaptive offloading scheme called PA-Offload that controls the offloading of image processing tasks from a drone to a fog node. To incorporate uncertainty factors, we introduce Stochastic Reward Nets (SRNs) to model the entire system behavior and compute a performability metric that is a composite measure of service throughput and system availability. The estimated performability value is then used to determine when to start or stop the offloading in order to make a better trade-off between performance and availability. Our numerical experiments show the effectiveness of PA-offload in terms of performability compared to non-adaptive fog offloading schemes. Fumio Machida, Ermeson Carneiro de Andrade |
ICFEC | 1 |
| 2021 | Availability Modeling for Drone Image Processing Systems with Adaptive OffloadingabstractAvailability of a computing process running on a flying drone is an essential quality aspect for mission-critical drone systems. Computing tasks such as image processing tasks can be lost when the process encounters a failure. Since the failure probability of the process depends on workload intensities, reducing drone workloads by computation offloading or load-balancing must have impacts on the system availability. While many existing studies discuss the performance-cost tradeoff associated with computation offloading, potential impacts on the system availability have not been deeply investigated yet. In this paper, we propose stochastic models for estimating the availability, the performance, and the energy consumption of a drone system with image processing tasks that can be either offloaded to a fog node or distributed to a collaborative drone. Our comprehensive numerical analysis with the proposed model clarifies the trade-offs among the availability, the throughputs, and the energy consumption under different computation modes. Furthermore, we propose an adaptive offloading scheme that can change the computation modes dynamically according to workload intensities and network conditions. A simulation study with a phased mission scenario shows that the proposed adaptive scheme can achieve high availability with 26% of energy reduction and less than 4% of throughput losses. Fumio Machida, Ermeson Carneiro de Andrade |
PRDC | 1 |
| 2019 | On the Diversity of Machine Learning Models for System ReliabilityabstractThe diversity of system components is one of the important contributing factors of reliable and secure software systems. In a software fault-tolerant system using diverse versions of software components, a component failure caused by defects or malicious attacks can be covered by other versions. Machine learning systems can also benefit from such a multi-version approach to improve the system reliability. Nevertheless, there are few studies addressing this issue. In this paper, we experimentally analyze how outputs of machine learning modules can be diversified by using different versions of machine learning algorithms, neural network architectures and perturbated input data. The experiments are conducted on image classification tasks of MNIST data set and Belgian Traffic Sign data set. Different neural network architectures, support vector machines and random forests are used for constructing diverse machine learning models. The diversity is characterized by the coverage of errors over the test samples. We observe that the different machine learning models have quite different error coverages that can be leveraged for system reliability design. Based on the experimental results, we construct the reliability model for three-version machine learning architecture with a diversity measure defined as the intersection of error spaces in the sample space. From the presented reliability model, we derive a necessary condition under which three-version architecture achieves a higher system reliability than a single machine learning module. Fumio Machida |
PRDC | 1 |
| 2018 | Performability Modeling for RAID Storage Systems by Markov Regenerative ProcessabstractThis paper presents a performability model for RAID storage systems using Markov regenerative process to compare different RAID architectures. While homogeneous Markov models are extensively used for reliability analysis of RAID storage systems, the memory-less property of the sojourn time assumed in such models is not satisfied in reality, especially in disk rebuild process whose progress is not interrupted even at an event of another disk failure. In this paper, we use Markov regenerative process which allows us to model the generally distributed rebuild times providing a needed extension of the traditional Markov models. The Markov regenerative process is then used to assess the performability of the storage system by assigning reward rates to each state based on the real storage benchmark results. Our numerical study characterizes the performability advantage of RAID6 architecture over RAID10 architecture in terms of sequential read access. Our findings include that the effect of exponential assumption for the rebuild times has practically negligible effect when we focus on data availability. However, the effect this approximation on performability prediction may not be negligible especially when the performance level drastically changes in degraded states. Our MRGP model provides more accurate prediction of performability in such cases. Fumio Machida, Ruofan Xia, Kishor S. Trivedi |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2017 | Lifetime Extension of Software Execution Subject to AgingabstractSoftware aging is a phenomenon of progressive degradation of software execution environment caused by software faults. In this paper, we propose software life-extension as an operational countermeasure against software aging and present the mathematical foundations of software life-extension by means of stochastic modeling. A semi-Markov process is used to capture the behavior of a system with software life-extension and to analyze the system's availability and completion times of jobs running on it. The semi-Markov process can correctly model the time-based life-extension and allows us to derive the optimal trigger for starting life-extension in terms of system availability and mean job completion time. We also present an effective combination of software life-extension and software rejuvenation that can maximize the system availability compared with a system using either rejuvenation or software life-extension. Fumio Machida, Jianwen Xiang, Kumiko Tadano, Yoshiharu Maeno |
IEEE Trans. Reliab. | 1 |
| 2016 | A Test Design Method for Resilient System on Cloud InfrastructureabstractIT systems for social infrastructure need to be resilient against a variety of disturbance events including hardware failures, software malfunctions, and application load surges. Techniques of dynamic system reconfiguration such as auto-scaling and auto-healing, which are typically provided in cloud services, can help maintain the system availability and good performance even after disturbance events. Such reconfiguration operations should be carefully designed and tested well on the production system (in real use) since the offline test is not enough to guarantee the functionality and performance of the production system after reconfiguration. In this paper, we propose a new test design method, called Live System Evaluation, to test an expected system configuration after reconfiguration operation triggered by disturbance events (e.g., a server failure). The test system is designed in a way to be embedded on the production system so that it can validate the expected system with high confidence with the minimum additional resources. We formulate the problem as an optimization problem and develop a prototype solver. An example case study on Network Function Virtualization application shows that our design method reduces the total resource cost by half compared with a conventional approach which requires a separated test system. Masaya Fujiwaka, Fumio Machida, Seiichi Koizumi |
CLOUD | 2 |
| 2015 | Just-in-time server procurement to private cloud for mobile thin-client serviceabstractMobile thin-client services are gaining lots of attention from companies who concern about the security yet recognize the benefits of mobile computing in their businesses. The service is based on a private cloud system hosting virtual machines that can execute mobile OS instances. The owner of the private cloud needs to prepare sufficient server resources for hosting those virtual machines. In this paper, we propose a framework to guide server procurement decisions in a private cloud for mobile thin-client service, which aims to minimize the cost of unused servers while avoiding service level violation due to the lack of resources. To make a timely server procurement decision, the framework combines the techniques for workload estimation to individual VMs, demand estimation of newly created VMs and repetitive simulation of VM replacement algorithm. Through a simulation study, we show that the proposed framework can reduce the cost of unused servers by 25% while satisfying a service level, compared with a time-based heuristic decision method. Fumio Machida, Shunsuke Kohno, Kosuke Maebara, Masayuki Nakagawa |
CNSM | 1 |
| 2015 | A Scalable Optimization Framework for Storage Backup Operations Using Markov Decision ProcessesabstractExplosive growth of data generation and increasing reliance of business analysis on massive data make data loss more damaging than ever before. Thus it has also become a critical issue for businesses to protect important data effectively. In a system with multiple data sets, complex system configurations and data protection requirements, backup planning plays an important role for maintaining the desired level of data protection while minimizing the impact on system operation. In this paper we investigate the use of Markov Decision Process (MDP) to guide the planning of data backup operations. To improve the applicability of the MDP framework to large systems, we present a novel approximation method to enhance its scalability. The benefit of the framework is demonstrated through numerical examples, where our MDP method reduces the storage system downtime by over 50% compared to the best heuristic approach. Ruofan Xia, Fumio Machida, Kishor S. Trivedi |
PRDC | 2 |
| 2015 | An Imperfect Fault Coverage Model With Coverage of Irrelevant ComponentsabstractThis paper addresses the coverage (including identification and isolation) of irrelevant components in systems with imperfect fault coverage (IFC). In fault-tolerant systems, a single not-covered component fault may thwart the automatic recovery mechanisms, and lead to a system or subsystem failure. The models that consider the effects of IFC are known as coverage models (CMs). In traditional CMs, except those considering functional dependency (a similar concept to relevancy but with different assumptions and semantics), coverage is typically limited to faulty components regardless of their relevancies. Consequently, an operational but irrelevant component will not be isolated, and may threaten the system by its future uncovered (not-covered) failures. Although the system is generally assumed to be coherent, which implies the relevancy of each component in the initial system state, the traditional CMs do not consider the fact that an initially relevant component could become irrelevant after the failures of other components. We propose the irrelevancy coverage model (ICM) to cover the irrelevant components in addition to the faulty components. In the ICM, a component will be isolated from the system whenever it becomes irrelevant (even it is not failed), such that its future not-covered failures will not affect the system anymore. By incorporating the coverage of irrelevant components, the ICM opens up a new cost-effective approach to improve system reliability without additional redundancy. Jianwen Xiang, Fumio Machida, Kumiko Tadano, Yoshiharu Maeno |
IEEE Trans. Reliab. | 2 |
| 2014 | A Markov Decision Process Approach for Optimal Data Backup SchedulingabstractThe explosive growth of data generation and increasing reliance of business analysis on massive data make data loss more damaging than ever before. Nowadays many organizations start relying on cloud services for keeping their valuable data. It is a critical issue for cloud service provider to protect the data for individual users securely and effectively. To protect the data in a system with multiple data sources, backup schedule plays an important role for achieving the desired level of data protection while minimizing the impact on system operation. In this paper we investigate the use of Markov Decision Process (MDP) to guide the scheduling of data backup operation and propose a framework to automatically generate an MDP instance from system specifications and data protection requirements. We then demonstrate the benefits of the MDP approach. Ruofan Xia, Fumio Machida, Kishor S. Trivedi |
DSN | 2 |
| 2014 | Virtualized server infrastructure for resilient voice communication serviceabstractThe huge earthquake struck Japan on 11th March, 2011, caused a massive congestion of call attempts from mobile phones that resulted in only 5% of accepted connections due to the congestion control by telephone companies. It is an emergent issue to improve resiliency of communication service in anticipation of future disasters and thus communication service infrastructure necessitates the flexibility of its capacity. In this paper, we introduce server virtualization technology to provide a flexible communication service infrastructure and design a communication service controller that provisions additional capacity by virtual machines in response to increased call attempts from mobile phones after a disaster. The communication service controller is designed with constraint programming and performance/availability estimations for deciding the optimum virtual machine placement according to network operator's instructions. Through an experimental disaster test, we confirm the capacity of communication service is increased five-fold by virtual machine provisioning with half an hour latency. Fumio Machida, Ryota Mibu, Junichi Gokurakuji, Kazuo Yanoo, Kumiko Tadano, Yoshiharu Maeno, Tomoyoshi Sugawara |
NOMS | 1 |
| 2014 | Computing Defects per Million in Cloud Caused by Virtual Machine Failures with ReplicationabstractVirtual machines (VM) are used in cloud computing systems to handle user requests for service. A typical user request goes through several cloud service provider specific processing steps from the instant it is submitted until the service is completed. In the process of providing the service, VM failures cause the user's request to be dropped. To mitigate the adverse impact of VM failure, replication mechanisms, either using cold, warm or hot replication, can be used. In this paper, we model the system behavior with a structure-state process to characterize the failure-recovery behavior of a VM in a cloud that uses one of the aforementioned replication schemes. We use a service-oriented dependability metric called Defects Per Million (DPM), defined as the number of user requests dropped out of a million. The structure-state process approach is used to analyze the job completion time distribution and subsequently we compute the DPM by counting the number of requests exceed the specified deadline. The effectiveness of replication schemes are demonstrated through numerical results. Subrota K. Mondal, Jogesh K. Muppala, Fumio Machida, Kishor S. Trivedi |
PRDC | 3 |
| 2014 | Analysis of Persistence of Relevance in Systems with Imperfect Fault Coverage
Jianwen Xiang, Fumio Machida, Kumiko Tadano, Yoshiharu Maeno |
SAFECOMP | 2 |
| 2014 | A Systematic Differential Analysis for Fast and Robust Detection of Software AgingabstractSoftware systems running continuously for a long time often confront software aging, which is the phenomenon of progressive degradation of execution environment caused by latent software faults. Removal of such faults in software development process is a crucial issue for system reliability. A known major obstacle is typically the large latency to discover the existence of software aging. We propose a systematic approach to detect software aging which has in a shorter test time and higher accuracy compared to traditional aging detection via stress testing and trend detection with high confidence. The approach is based on a comparative differential analysis where a software version under test is compared with against a previous robust version by observing in terms of behavioral (signal) changes during system tests of resource metrics. A key instrument adopted is a divergence chart, which expresses time-dependent differences between two signals, allowing us to detect changes in the system metrics' values which indicate the existence of software aging. In our experimental study, we focuses on memory-leak detection and the and evaluates divergence charts are computed using various multiple statistical techniques combined paired with different application-level memory related metrics (RSS and Heap Usage). The experimental results show that the statistical process control techniques used in our approach proposed method achieves good performance for memory-leak detection, when compared with other in comparison to techniques widely adopted in previous works (e.g., linear regression, moving average and median). Rivalino Matias, Artur Andrzejak 0001, Fumio Machida, Diego Costa 0001, Kishor S. Trivedi |
SRDS | 3 |
| 2014 | Job completion time on a virtualized server with software rejuvenationabstractThis article analyzes the completion time of a job running on a virtualized server subject to software aging and rejuvenation in a virtual machine monitor (VMM). A job running on the server may be interrupted by virtual machine (VM) failure, VMM failure or VMM rejuvenation. The job interruption is categorized as either preemptive-repeat ( prt ), in which case the interrupted job needs to restart from the beginning, or preemptive-resume ( prs ), in which case the job resumes execution from the point of interruption. Using a semi-Markov process (SMP) to model the server behavior, the steady-state server availability is computed and the theory developed in Kulkarni et al. [1987] is used to obtain the Laplace-Stieltjes transform (LST) of the job completion time. In the numerical experiments, we introduce four types of aging behavior of VMM. The effectiveness of VMM rejuvenation on job completion time is discussed in association with the type of interruption it causes and the VMM aging type. With our parameter settings, VMM rejuvenation with prs job interruption improves the performance of job execution regardless of the aging type, with performance degradation is taken into account. Fumio Machida, Victor F. Nicola, Kishor S. Trivedi |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2014 | Performance and Availability Modeling of ITSystems with Data Backup and RestoreabstractIn modern IT systems, data backup and restore operations are essential for providing protection against data loss from both natural and man-made incidents. On the other hand, data backup and restore operations can be resource-intensive and lead to performance degradation, or may require the system to be offline entirely. Therefore, it is important to properly choose backup and restore techniques and policies to ensure adequate data protection while minimizing the impact on system availability and performance. In this paper, we present an analytical modeling approach for such a purpose. We study a file service system that undergoes periodic data backups, and investigate metrics concerning system availability, data loss and rejection of user requests. To obtain the metrics, we combine a variety of model types, including Markov chains, queuing networks and Stochastic Reward Nets, to construct a set of analytical models that capture the operational details of the system. We then compute the metrics of interest under different backup/restore techniques, policies, and workload scenarios. The numerical results allow us to compare the effects of different backup/restore techniques and policies in terms of the tradeoff between protective power and impact on system performance and availability. Ruofan Xia, Xiaoyan Yin 0002, Javier Alonso 0001, Fumio Machida, Kishor S. Trivedi |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2013 | Performability analysis of RAID10 versus RAID6abstractDesign of storage system configuration is one of the key issues for providing dependable IT systems. An appropriate RAID storage configuration should consider both performance and availability. To assist the design, this paper presents the performability models for RAID10 and RAID6 that can be used to compare the configuration quantitatively. A performability advantage of RAID6 over RAID10 in sequential read access is discovered by the numerical study in conjunction with performance benchmark results. Fumio Machida, Jianwen Xiang, Kumiko Tadano, Yoshiharu Maeno, Takashi Horikawa |
DSN | 1 |
| 2013 | Composing hierarchical stochastic model from SysML for system availability analysisabstractComprehensive analytic model for system availability analysis often confronts the largeness issue where a system designer cannot easily handle the model and the solution is not given in a feasible solution time. Hierarchical decomposition of a large state-space model gives a promising solution to the largeness issue when the model is decomposable. However, the decomposability of analytic model is not always manually tractable especially when the model is generated in an automated manner. In this paper, we propose an automated model composition technique from a system design to a hierarchical stochastic model which is the judicious combination of combinatorial and state-space models. In particular, from SysML-based system specifications, a top-level fault tree and associated stochastic reward nets are automatically generated in hierarchical manner. The obtained hierarchical stochastic model can be solved analytically considerably faster than monolithic state-space models. Through an illustrative example of three-tier web application system on a virtualized infrastructure, the accuracy and efficiency of the solution are evaluated in comparison to a monolithic state space model and a static fault tree. Fumio Machida, Jianwen Xiang, Kumiko Tadano, Yoshiharu Maeno |
ISSRE | 1 |
| 2013 | Performability Modeling of Manual Resolution of Data Inconsistencies for Optimization of Data Synchronization Interval
Kumiko Tadano, Jianwen Xiang, Fumio Machida, Yoshiharu Maeno |
MODELSWARD | 3 |
| 2013 | Modeling and analysis of software rejuvenation in a server virtualized system with live VM migration
Fumio Machida, Dong Seong Kim 0001, Kishor S. Trivedi |
Perform. Evaluation | 1 |
| 2012 | Software Life-Extension: A New Countermeasure to Software AgingabstractThis paper presents software life-extension, a new technique for counteracting software aging by preventive operation to extend the lifetime of software execution. Software aging is a phenomenon of progressive degradation of execution environment due to aging-related software faults and it might cause resource depletion resulting in system failures. To extend the lifetime of the software affected by aging, we use a virtual machine to execute the software and allocate additional memory to the virtual machine upon software aging detection. Although software life-extension is a temporal solution as it only postpones the occurrence of a failure, it provides a simple, cost-effective, and non-intrusive countermeasure to software aging. The feasibility and effectiveness of software life-extension are studied by the experiments on memcached, a widely adopted general-purpose in-memory cache server. From the experimental results, we present a Semi-Markov process (SMP) describing the general behavior of software life-extension and analyze the model which gives the prediction of the system availability as well as the user-perceived availability. Fumio Machida, Jianwen Xiang, Kumiko Tadano, Yoshiharu Maeno |
ISSRE | 1 |
| 2012 | Identification of Minimal Unacceptable Combinations of Simultaneous Component Failures in Information SystemsabstractLarge-scale disasters may cause simultaneous failures of many components in information systems. In the design for disaster recovery, operational procedures to recover from simultaneous component failures need to be determined so as to satisfy the time-to-recovery objective within the limited budget. For this purpose, it is beneficial to identify the minimal unacceptable combination of component failures which violates the requirements for time-to-recovery or the required cost. The identified combination allows us to know the limitation of the recovery capability of the designed recovery operation procedure. In this paper, we propose a technique to identify the minimal unacceptable combination of component failures by predicting the required time and cost for recovery from each combination of component failures. We synthesize analytic models from the description of recovery operation procedure in the form of SysML Activity Diagram, and solve the models to predict the time-to-recovery and the cost. The feasibility of the proposed technique is evaluated in an example of recovery operation procedures for a commercial database management system. Kumiko Tadano, Fumio Machida, Jianwen Xiang, Yoshiharu Maeno |
PRDC | 2 |
| 2012 | Availability Modeling and Analysis for Data Backup and Restore OperationsabstractData backup operation is an essential part of common IT system administration to protect against data loss caused by any storage failures, human errors, or disasters. Lost data can be recovered from the backed up data if it exists. Since the backup and restore operations accrue downtime overhead or performance degradation, they have to be designed to ensure the data reliability while minimizing the performance and availability overhead. In this paper, we study the impacts of different backup policies on availability measures such as storage availability, system availability, and user-perceived availability. Backup and restore operations are designed using SysML Activity diagrams that are automatically translated into Stochastic Reward Net (SRN) to compute the availability measures. Our numerical results show the effectiveness of the combination of full backup and partial backup in terms of user-perceived data availability and data loss rate. Furthermore, the sensitivity ranking can help improve the availability measures. Xiaoyan Yin 0002, Javier Alonso 0001, Fumio Machida, Ermeson Carneiro de Andrade, Kishor S. Trivedi |
SRDS | 3 |
| 2012 | Sensitivity Analysis of Server Virtualized System AvailabilityabstractServer virtualization is a technology used in many enterprise systems to reduce operation and acquisition costs, and increase the availability of their critical services. Virtualized systems may be even more complex than traditional nonvirtualized systems; thus, the quantitative assessment of system availability is even more difficult. In this paper, we propose a sensitivity analysis approach to find the parameters that deserve more attention for improving the availability of systems. Our analysis is based on Markov reward models, and suggests that host failure rate is the most important parameter when the measure of interest is the system mean time to failure. For capacity oriented availability, the failure rate of applications was found to be another major concern. The results of both analyses were cross-validated by varying each parameter in isolation, and checking the corresponding change in the measure of interest. A cost-based optimization method helps to highlight the parameter that should have higher priority in system enhancement. Rúbens de Souza Matos Júnior, Paulo Romero Martins Maciel, Fumio Machida, Dong Seong Kim 0001, Kishor S. Trivedi |
IEEE Trans. Reliab. | 3 |
| 2011 | Modeling and Analyzing Server System with Rejuvenation through SysML and Stochastic Reward NetsabstractHigh-availability assurance of server systems is becoming an important issue, since many mission-critical applications are implemented on server systems. To achieve high-availability, software rejuvenation is a practical technique to reduce unexpected downtime caused by software aging in software applications running on server systems. Although analytic models of software rejuvenation are well-studied, such analysis is not used in server system administration due to the complexity of modeling. In this paper, we present an availability modeling method for server system with software rejuvenation based on SysML that is used to describe system configurations and maintenance operations semi-formally. The proposed approach allows system administrators, who do not have expertise in availability modeling, to design and study the effects of different rejuvenation policies deployed in server systems. To show the applicability of the proposed modeling and evaluation process, a case study of a web application server is presented. We show the correctness of our modeling method by comparing the conventional models for condition-based and time-based software rejuvenation. Ermeson Carneiro de Andrade, Fumio Machida, Dong Seong Kim 0001, Kishor S. Trivedi |
ARES | 2 |
| 2011 | Efficient Analysis of Fault Trees with Voting GatesabstractThe voting gate, or k-out-of-n (k/n) gate, is a standard logic gate used in fault trees modelling fault-tolerant systems. It is traditionally expanded into a combination of AND and OR gates, and this expansion may result in combinatorial explosion problem in the calculation of minimal cut sets (MCSs) of the fault tree for even a not very big n, especially when the voting gate inputs are intermediate rather than basic events. In this paper we propose a set of reduction rules to simplify the voting gates without direct expanding, and also propose a concept of minimal cut vote (MCV) denoting a k/n gate whose inputs are all basic events and whose k-combinations are all MCSs of the fault tree. With the proposed reduction rules and MCV concept, the MCSs of fault trees can be evaluated and weeded more efficiently and the result can be represented in a more compact form. The results of experiments on practical fault trees with voting gates show that our method not only outperforms conventional MCS evaluation methods by several orders of magnitude but also provides performance comparably to that provided by binary decision tree (BDD) based algorithms. Jianwen Xiang, Kazuo Yanoo, Yoshiharu Maeno, Kumiko Tadano, Fumio Machida, Atsushi Kobayashi, Takao Osaki |
ISSRE | 5 |
| 2011 | Candy: Component-based Availability Modeling Framework for Cloud Service Management Using SysMLabstractHigh-availability assurance of cloud service is a critical and challenging issue for cloud service providers. To quantify the availability of cloud services from both architectural and operational points of views, availability modeling and evaluation are essential. This paper presents a component-based availability modeling framework, named Candy, which constructs a comprehensive availability model semi-automatically from system specifications described by Systems Modeling Language (SysML). SysML diagrams are translated into components of availability model and the components are assembled together to form the entire availability model in Stochastic Reward Nets (SRNs). In order to incorporate the maintenance operations of cloud services in availability models, Candy defines the translation rules from Activity diagram to SRN and synchronizes the related SRNs according to SysML allocation notations. The feasibility of the proposed modeling and availability evaluation process is studied by an illustrative example of a web application service hosted on a cloud infrastructure having multiple failure isolation zones and automatic scale-up function. Fumio Machida, Ermeson Carneiro de Andrade, Dong Seong Kim 0001, Kishor S. Trivedi |
SRDS | 1 |
| 2010 | Resource Information Cache Update Control for Scalable Access Control Management SystemsabstractIn private clouds that host many enterprise applications, scalable security management has become an important issue. In our previous work, we had developed integrated access control manager that manages access permissions to a large number of various resources using resource information provided by a resource information service. To improve performance of the resource information service, we introduced a resource information cache and a proactive cache update control method. To avoid overload of the management server due to updating cached information, the proposed method selects a part of cached information by content priority as an update target. In this work, we evaluated the query response time of the resource information service in a third-party enterprise system using the search queries issued during system operations by an administrator. The proposed method reduced average query response time by 35% compared to a conventional reactive update control method. Kumiko Tadano, Masahiro Kawato, Fumio Machida, Yoshiharu Maeno |
IEEE CLOUD | 3 |
| 2010 | Renovating legacy distributed systems using virtual appliance with dependency graphabstractLegacy distributed systems hosted on old infrastructures can be renovated using virtual appliance that is a package of virtual machines, installed applications and their configurations. By converting a legacy distributed system to a virtual appliance, we can conserve the entire system and restart the application on the specific virtualization platforms. However, in order to execute the application service properly on the new hosting environment, some additional network configurations are required to resolve the dependencies inherited in the original system. In this paper, we propose a framework named DS Renovator that converts a legacy distributed system to a virtual appliance and renovates the system on a new hosting environment with optimum deployment for resolving the dependencies. In the virtual appliance generation process, DS Renovator analyzes server dependencies inherent in the legacy distributed system and generates a dependency graph that consists of nodes and edges representing servers and their dependencies respectively. In the virtual appliance deployment process, DS Renovator applies graph partitioning algorithm to the dependency graph to determine the optimum virtual machine placement which minimizes the dependencies between the hosting servers under the capacity limitations. As a graph partitioning algorithm, we propose an edge-contraction based efficient algorithm. The performance of the proposed algorithm is evaluated with case studies in comparison to other approximation algorithms. Fumio Machida, Masahiro Kawato, Yoshiharu Maeno |
CNSM | 1 |
| 2010 | Digital Watermarking of Virtual Machine Images
Kumiko Tadano, Masahiro Kawato, Ryo Furukawa 0003, Fumio Machida, Yoshiharu Maeno |
IFIP Int. Conf. Digital Forensics | 4 |
| 2010 | Redundant virtual machine placement for fault-tolerant consolidated server clustersabstractConsolidated server systems using server virtualization involves serious risks of host server failures that induce unexpected downs of all hosted virtual machines and applications. To protect applications requiring high-availability from unpredictable host server failures, redundant configuration using virtual machines can be an effective countermeasure. This paper presents a virtual machine placement method for establishing a redundant configuration against host server failures with less host servers. The proposed method estimates the requisite minimum number of virtual machines according to the performance requirements of application services and decides an optimum virtual machine placement so that minimum configurations survive at any k host server failures. The evaluation results clarify that the proposed method achieves requested fault-tolerance level with less number of hosting servers compared to the conventional N+M redundant configuration approach. Fumio Machida, Masahiro Kawato, Yoshiharu Maeno |
NOMS | 1 |
| 2009 | Availability Modeling and Analysis of a Virtualized SystemabstractThis paper develops an availability model of a virtualized system. We construct non-virtualized and virtualized two hosts system models using a two-level hierarchical approach in which fault trees are used in the upper level and homogeneous continuous time Markov chains (CTMC) are used to represent sub-models in lower level. In the models, we incorporate not only hardware failures (e.g., CPU, memory, power, etc) but also software failures including Virtual Machine Monitor (VMM), Virtual Machine (VM), and application failures. We also incorporate high availability (HA) service and VM live migration in the virtualized system. Metrics we use are system steady state availability, downtime in minutes per year and capacity oriented availability. Dong Seong Kim 0001, Fumio Machida, Kishor S. Trivedi |
PRDC | 2 |
| 2006 | Guarantee of Freshness in Resource Information Cache on WSPE: Web Service Polling EngineabstractManaging resource information using web services is an important issue in grid computing. Although some standards for grid computing based on web services have been published, the overhead of SOAP/HTTP and XML processing may cause performance problems in the resource information service for the grid. This paper proposes a scheduling update mechanism for resource information cache to meet the performance requirements of the resource information service. We have designed and implemented this mechanism on an information service named Web Service Polling Engine (WSPE) that collects resource information in accordance with WSDM/MUWS standards. From the results of evaluations, we conclude that WSPE can guarantee the freshness of resource information with a stable load on the WSPE server. Fumio Machida, Masahiro Kawato, Yoshiharu Maeno |
CCGRID | 1 |
| 2004 | Polimatica: Abstraction for Customizable Private Virtual Organizations in Global GridsabstractWe have developed an application service hosting middleware for enterprise information systems, "Polimatica", upon which a policy-customizable private virtual organization (pVO) is built. The pVO is a collection of customizable application services in the global grid computing environment. "Polimatica" is distinctive in terms of abstraction capabilities in (1) policy instance description, (2) view synthesis of an extensive resource data model, (3) operation aggregation on the OGSI extension interfaces. The performance to enforce the policy instances is measured in a laboratory testbed for an enterprise ASP, providing B2B e-commerce (BizEngine/WS) and interactive media communication service (LiveComm). We have verified that, in 60 seconds, the feedback loop in a pVO adjusts the application services as customized given the changing environment. Yoshiharu Maeno, Masahiro Kawato, Shoji Nishimura, Fumio Machida, Tsunehiko Kamachi |
ICWS | 4 |