Kenji Yoshihira

dblp:38/467 · DBLP profile ↗
← Back
26ranked-venue papers
0as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6Security and privacy · 6Databases, data management, data science and information retrieval · 6Artificial intelligence and machine learning · 2Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Distributed systems · 70% Performance modeling and evaluation · 24% Cloud and datacenter computing · 6%
Software engineering, system software, and programming languages
1 paper
Program analysis · 50% Runtime systems and virtual machines · 50%
Computer networks
2 papers
Network measurement and analytics · 88% Content delivery and video streaming · 12%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
fault tolerance
0.342008
Monitoring High-Dimensional Data for Failure Detection and Localization in Large-Scale Computing Systems · IEEE Trans. Knowl. Data Eng. 2008
Failure Detection in Large-Scale Internet Services by Principal Subspace Mapping · IEEE Trans. Knowl. Data Eng. 2007
Modeling and Tracking of Transaction Flow Dynamics for Fault Detection in Complex Systems · IEEE Trans. Dependable Secur. Comput. 2006
Performance modeling and evaluation
workload characterization
0.232008
Understanding internet video sharing site workload: a view from data center design · WWW 2008
Modeling and Tracking of Transaction Flow Dynamics for Fault Detection in Complex Systems · IEEE Trans. Dependable Secur. Comput. 2006
Efficient and Scalable Algorithms for Inferring Likely Invariants in Distributed Systems · IEEE Trans. Knowl. Data Eng. 2007
Program analysis › dynamic analysis
dynamic instrumentation
0.212013
iProbe: A lightweight user-level dynamic instrumentation tool · ASE 2013
Distributed systems
anomaly detection
0.222008
Monitoring High-Dimensional Data for Failure Detection and Localization in Large-Scale Computing Systems · IEEE Trans. Knowl. Data Eng. 2008
Failure Detection in Large-Scale Internet Services by Principal Subspace Mapping · IEEE Trans. Knowl. Data Eng. 2007
Distributed systems › fault tolerance
failure detection
0.222008
Monitoring High-Dimensional Data for Failure Detection and Localization in Large-Scale Computing Systems · IEEE Trans. Knowl. Data Eng. 2008
Failure Detection in Large-Scale Internet Services by Principal Subspace Mapping · IEEE Trans. Knowl. Data Eng. 2007
Network measurement and analytics
anomaly detection
0.112008
Automatic Profiling of Network Event Sequences: Algorithm and Applications · INFOCOM 2008
Network measurement and analytics
traffic characterization
0.112008
Automatic Profiling of Network Event Sequences: Algorithm and Applications · INFOCOM 2008
Cloud and datacenter computing
datacenter architecture
0.112008
Understanding internet video sharing site workload: a view from data center design · WWW 2008
Distributed systems › fault tolerance › failure diagnosis
failure localization
0.112008
Monitoring High-Dimensional Data for Failure Detection and Localization in Large-Scale Computing Systems · IEEE Trans. Knowl. Data Eng. 2008
Performance modeling and evaluation
system modeling
0.112008
Exploiting Local and Global Invariants for the Management of Large Scale Information Systems · ICDM 2008
Performance modeling and evaluation › performance monitoring
monitoring data analysis
0.112007
Efficient and Scalable Algorithms for Inferring Likely Invariants in Distributed Systems · IEEE Trans. Knowl. Data Eng. 2007
Distributed systems › fault tolerance
fault detection and diagnosis
0.112006
Modeling and Tracking of Transaction Flow Dynamics for Fault Detection in Complex Systems · IEEE Trans. Dependable Secur. Comput. 2006
Network measurement and analytics
traffic analysis
0.012008
Automatic Profiling of Network Event Sequences: Algorithm and Applications · INFOCOM 2008
Distributed systems › observability
distributed monitoring
0.012007
Efficient and Scalable Algorithms for Inferring Likely Invariants in Distributed Systems · IEEE Trans. Knowl. Data Eng. 2007

Methods — techniques the papers use, named apart from their topics

offline compilation · 0.2hot patching · 0.2workload measurement · 0.2queueing theory · 0.2principal component analysis · 0.2squared prediction error · 0.1hotelling t2 · 0.1mixture model · 0.1markov chain · 0.1laplacian prior · 0.1expectation-maximization · 0.1bayesian regression · 0.1flow intensity · 0.1constant relationship search · 0.1canonical correlation analysis · 0.1
YearPublicationVenuePosition
2014 Software system performance debugging with kernel events feature guidance
abstract
To diagnose performance problems in production systems, many OS kernel-level monitoring and analysis tools have been proposed. Using low level kernel events provides benefits in efficiency and transparency to monitor application software. On the other hand, such approaches miss application-specific semantic information which can be effective to differentiate the trace patterns from distinct application logic. This paper introduces new trace analysis techniques based on event features to improve kernel event based performance diagnosis tools. Our prototype, AppDiff, is based on two analysis features: system resource features convert kernel events to resource usage metrics, thereby enabling the detection of various performance anomalies in a unified way; program behavior features infer the application logic behind the low level events. By using these features and conditional probability, AppDiff can detect outliers and improve the diagnosis of application performance.
Junghwan Rhee, Hui Zhang 0002, Nipun Arora, Guofei Jiang, Kenji Yoshihira
NOMS5
2014 Uscope: A scalable unified tracer from kernel to user space
abstract
Unified tracing is the process of collecting trace logs across the boundary of kernel and user spaces, and has been used to understand the in-depth correspondence between low level events and application program context for diagnosing system failures and performance problems. Crossing the boundary from the kernel space to a user space to collect trace events from dual spaces imposes challenges compared to crossing the boundary in the other way from a user space to the kernel space due to multiple scheduled programs and diverse code layouts in the user space regarding the tracing target. In this paper, we propose a novel unified tracing system called Uscope to systematically trace kernel and unprecedented user code with low overhead. The key idea is to use an efficient variant of stack walking. Uscope lowers stack walking overhead by adjusting the scope of walking in two ways: (1) a highly configurable focus within the call stack, and (2) a per-application tracing that systematically tracks a dynamic set of new, exiting, or transforming processes and threads of an application software. This system is realized by using a flexible stack walking algorithm and a runtime kernel structure, Trace Map. These key features lead to low run-time overhead under 6% relative to native execution on a set of widely used benchmarks.
Junghwan Rhee, Hui Zhang 0002, Nipun Arora, Guofei Jiang, Kenji Yoshihira
NOMS5
2014 CLUE: System trace analytics for cloud service performance diagnosis
abstract
In this paper, we present CLUE, a system event analytics tool for black-box performance diagnosis in production Cloud Computing systems. CLUE provides an unified and extensible means of profiling service transactional behaviors, and builds structured data called event sketches. CLUE further offers a set of analytic tools for summarizing and analyzing event sketches by integrating data mining and statistical analysis. CLUE has been developed in NEC as an internal tool and applied in diagnosing a diverse set of real performance problems for multi-tiered IT applications running on multi-core servers of major platforms including Linux (Redhat, Fedora), Unix (HP-UX), and Windows (Windows Server 2008). We demonstrated the evaluation of our framework on real-world IT systems, and showed how it can enable visibility and effective diagnosis of service system performance problems.
Hui Zhang 0002, Junghwan Rhee, Nipun Arora, Sahan Gamage, Guofei Jiang, Kenji Yoshihira, Dongyan Xu
NOMS6
2014 Proactive Workload Management in Hybrid Cloud Computing
abstract
The hindrances to the adoption of public cloud computing services include service reliability, data security and privacy, regulation compliant requirements, and so on. To address those concerns, we propose a hybrid cloud computing model which users may adopt as a viable and cost-saving methodology to make the best use of public cloud services along with their privately-owned (legacy) data centers. As the core of this hybrid cloud computing model, an intelligent workload factoring service is designed for proactive workload management. It enables federation between on- and off-premise infrastructures for hosting Internet-based applications, and the intelligence lies in the explicit segregation of base workload and flash crowd workload, the two naturally different components composing the application workload. The core technology of the intelligent workload factoring service is a fast frequent data item detection algorithm, which enables factoring incoming requests not only on volume but also on data content, upon a changing application data popularity. Through analysis and extensive evaluation with real-trace driven simulations and experiments on a hybrid testbed consisting of local computing platform and Amazon Cloud service platform, we showed that the proactive workload management technology can enable reliable workload prediction in the base workload zone (with simple statistical methods), achieve resource efficiency (e.g., 78% higher server capacity than that in base workload zone) and reduce data cache/replication overhead (up to two orders of magnitude) in the flash crowd workload zone, and react fast (with an X^2 speed-up factor) to the changing application data popularity upon the arrival of load spikes.
Hui Zhang 0002, Guofei Jiang, Kenji Yoshihira
IEEE Trans. Netw. Serv. Manag.3
2013 Fault detection and localization in distributed systems using invariant relationships
abstract
Recent advances in sensing and communication technologies enable us to collect round-the-clock monitoring data from a wide-array of distributed systems including data centers, manufacturing plants, transportation networks, automobiles, etc. Often this data is in the form of time series collected from multiple sensors (hardware as well as software based). Previously, we developed a time-invariant relationships based approach that uses Auto-Regressive models with eXogenous input (ARX) to model this data. A tool based on our approach has been effective for fault detection and capacity planning in distributed systems. In this paper, we first describe our experience in applying this tool in real-world settings. We also discuss the challenges in fault localization that we face when using our tool, and present two approaches - a spatial approach based on invariant graphs and a temporal approach based on expected broken invariant patterns - that we developed to address this problem.
Abhishek B. Sharma, Kenji Yoshihira, Guofei Jiang
DSN4
2013 Predictive VM consolidation on multiple resources: Beyond load balancing
abstract
Effective consolidation of different applications on common resources is often akin to black art as application performance interference may result in unpredictable system and workload delays. In this paper we consider the problem of fair load balancing on multiple servers within a virtualized data center setting. We especially focus on multi-tiered applications with different resource demands per tier and address the problem on how to best match each application tier on each resource, such that performance interference is minimized. To address this problem, we propose a two-step approach. First, a fair load balancing scheme assigns different virtual machines (VMs) across different servers; this process is formulated as a multi-dimensional vector scheduling problem that uses a new polynomial-time approximation scheme (PTAS) to minimize the maximum utilization across all server resources and results in multiple load balancing solutions. Second, a queueing network analytic model is applied on the proposed min-max solutions in order to select the optimal one. We experimentally evaluate the proposed two-stage mechanism using a Xen virtualization testbed that hosts multiple RUBiS multi-tier applications. Experimental results show that the proposed mechanism is robust as it always predicts the optimal consolidation strategy.
Hui Zhang 0002, Evgenia Smirni, Guofei Jiang, Kenji Yoshihira
IWQoS5
2013 iProbe: A lightweight user-level dynamic instrumentation tool
abstract
We introduce a new hybrid instrumentation tool for dynamic application instrumentation called iProbe, which is flexible and has low overhead. iProbe takes a novel 2-stage design, and offloads much of the dynamic instrumentation complexity to an offline compilation stage. It leverages standard compiler flags to introduce “place-holders” for hooks in the program executable. Then it utilizes an efficient user-space “HotPatching” mechanism which modifies the functions to be traced and enables execution of instrumented code in a safe and secure manner. In its evaluation on a micro-benchmark and SPEC CPU2006 benchmark applications, the iProbe prototype achieved the instrumentation overhead an order of magnitude lower than existing state-of-the-art dynamic instrumentation tools like SystemTap and DynInst.
Nipun Arora, Hui Zhang 0002, Junghwan Rhee, Kenji Yoshihira, Guofei Jiang
ASE4
2011 Application Behavior Mapping across Heterogeneous Hardware Platforms
abstract
Predicting the application behavior such as its resource utilization in a new hardware machine is becoming an urgent issue as the increasing number of servers with various configurations show up in data centers and clouds. Current two categories of approaches, the test bed evaluation based and the software simulation based methods, both have certain shortcomings. While the test bed evaluation based approaches suffer from the lack of measurement data to build the prediction model, the simulation based methods intrinsically introduce uncertainties and errors in the data. In order to overcome those issues, this paper proposes a new solution that combines the current two separate processes. We develop a generalized regression model with L1 penalty to predict the application behavior from software simulation. Meanwhile we also use evaluations on real hardware instances to improve the model obtained from simulation. Our model improvement is grounded on the Bayesian learning theory, which elegantly embeds outcomes from both simulation and real evaluation stages into the final prediction. Experimental results show the higher prediction accuracy of our method compared with current techniques.
Guofei Jiang, Kenji Yoshihira
DASC4
2011 Effective VM sizing in virtualized data centers
abstract
In this paper, we undertake the problem of server consolidation in virtualized data centers from the perspective of approximation algorithms. We formulate server consolidation as a stochastic bin packing problem, where the server capacity and an allowed server overflow probability p are given, and the objective is to assign VMs to as few physical servers as possible, and the probability that the aggregated load of a physical server exceeds the server capacity is at most p.
Ming Chen 0002, Hui Zhang 0002, Ya-Yunn Su, Guofei Jiang, Kenji Yoshihira
Integrated Network Management6
2010 Invariants Based Failure Diagnosis in Distributed Computing Systems
abstract
This paper presents an instance based approach to diagnosing failures in computing systems. Owing to the fact that a large portion of occurred failures are repeated ones, our method takes advantage of past experiences by storing historical failures in a database and retrieving similar instances in the occurrence of failure. We extract the system `invariants' by modeling consistent dependencies between system attributes during the operation, and construct a network graph based on the learned invariants. When a failure happens, the status of invariants network, i.e., whether each invariant link is broken or not, provides a view of failure characteristics. We use a high dimensional binary vector to store those failure evidences, and develop a novel algorithm to efficiently retrieve failure signatures from the database. Experimental results in a web based system have demonstrated the effectiveness of our method in diagnosing the injected failures.
Guofei Jiang, Kenji Yoshihira, Akhilesh Saxena
SRDS3
2010 A Cooperative Sampling Approach to Discovering Optimal Configurations in Large Scale Computing Systems
abstract
With the growing scale of current computing systems, traditional configuration tuning methods become less effective because they usually assume a small number of parameters in the system. In order to handle the scalability issue of configuration tuning, this paper proposes a cooperative optimization framework, which mimics the behavior of team playing to discover the optimal configuration setting in computing systems. We follow a `best of the best' rule to decompose the tuning task into a number of small subtasks with manageable size and complexity. While each decomposed module is responsible for the optimization of its own configuration parameters, all the modules share the performance evaluations of new samples as common feedbacks to enhance their optimization objectives. As a result, the qualities of generated samples become improved during the search, and the cooperative sampling will eventually discover the optimal configurations in the system. Experimental results demonstrate that our proposed cooperative optimization can identify better solutions within limited time periods compared with other state of the art configuration search methods. Such advantage becomes more significant when the number of configuration parameters increases.
Guofei Jiang, Hui Zhang 0002, Kenji Yoshihira
SRDS4
2010 Understanding Internet Video sharing site workload: A view from data center design
Xiaozhu Kang, Hui Zhang 0002, Guofei Jiang, Xiaoqiao Meng, Kenji Yoshihira
J. Vis. Commun. Image Represent.6
2008 Enabling Information Confidentiality in Publish/Subscribe Overlay Services
abstract
"Alice has a piece of valuable information which she is willing to sell to anyone who is interested in; she is too busy and wants to ask Bob, a professional broker, to sell that information for her; but Alice is in a dilemma where she cannot trust Bob with that information but Bob cannot help her find her customers without knowing that information." In this paper, we propose a security mechanism called information foiling to address new confidentiality problems arising in pub/sub overlay services [1]. Information foiling extends Rivest's "Chaffing and Winnowing" [2], and its basic idea is to carefully generate a set of fake messages to hide an authentic message. Information foiling requires no modification inside the broker network so that the routing/filtering capabilities of broker nodes remains intact. We formally present the information foiling mechanism in the context of publish/subscribe overlay services, and discuss its applicability in other Internet applications. For publish/subscribe applications, we propose a suite of optimal schemes for fake message generation in different scenarios. Real-world data are used in our evaluation to demonstrate the effectiveness of the proposed schemes.
Hui Zhang 0002, Abhishek B. Sharma, Guofei Jiang, Xiaoqiao Meng, Kenji Yoshihira
ICC6
2008 Exploiting Local and Global Invariants for the Management of Large Scale Information Systems
abstract
This paper presents a data oriented approach to modeling the complex computing systems, in which an ensemble of correlation models are discovered to represent the system status. If the discovered correlations can continually hold under different user scenarios and workloads, they are regarded as invariants of the information system. In our previous work, we have developed an algorithm to automatically search the invariants between any pair of system attributes, which we call local invariants. However that method is unable to deal with the high order dependency models due to the combinatorial explosion of search space. In this paper we use Bayesian regression technique to discover those high order correlation models, called global invariants. We treat each attribute as a response variable in turn and express its dependency with the other attributes in a regression model. By adding the prior constraint of Laplacian distribution to the regression coefficients, we can find the solution in which only the correlated attributes with respect to the response have nonzero regression coefficients. After that we further consider the temporal dependencies of those extracted attributes by incorporating their past observations. We also provide a confidence metric and a validation procedure to measure the reliability of learned models. If the model does not break down in the validation, it is regarded as a true invariant of the system. Experimental results on a real wireless networking system show that the discovered invariants can be used to effectively detect system failures as well as provide valuable information about the failure source.
Haibin Cheng, Guofei Jiang, Kenji Yoshihira
ICDM4
2008 Measurement, Modeling, and Analysis of Internet Video Sharing Site Workload: A Case Study
abstract
In this paper we measured and analyzed the workload on Yahoo! Video, the 2nd largest U.S. video sharing site, to understand its nature and the impact on online video data center design. We discovered interesting statistical properties on both static and temporal dimensions of the workload; they include file duration and popularity distributions, arrival rate dynamics and predictability, and workload stationarity and burstiness. Complemented with queueing-theoretic techniques, we extended our understanding on the measurement data with a virtual data center design assuming the same workload as measured, which reveals results regarding the impact of workload arrival distribution, service level agreements (SLAs) and workload scheduling schemes on the design and operations of such large-scale video distribution systems.
Xiaozhu Kang, Hui Zhang 0002, Guofei Jiang, Xiaoqiao Meng, Kenji Yoshihira
ICWS6
2008 Automatic Profiling of Network Event Sequences: Algorithm and Applications
abstract
The behavior of network entities, such as flows, sessions, hosts, and users, can often be described by communication event sequences in the time domain. For the purpose of many network measurement and monitoring tasks, it is desirable to have an accurate yet information-compact profiling of the behavior of massive event sequences. This paper proposes a new method to achieve this goal. On a given set of event sequences, the proposed method automatically learns a mixture model which fully captures the sequence behavior including both event pattern and duration between events. The learned mixture model is information-compact as it classifies sequences into a set of behavior templates, each of which is described by a Markov Chain. The model parameters are estimated in an iterative procedure which is developed from the Expectation Maximization algorithm. Two network management applications are proposed based on the method: a visualization tool for network administrators to conduct exploratory traffic analysis, and an efficient anomaly detection mechanism. In the evaluation, we validate the method accuracy as well as the usefulness of the two applications by using three networking datasets with different types: TCP packet traces, VoIP calls, and syslog traces in wireless networks.
Xiaoqiao Meng, Guofei Jiang, Hui Zhang 0002, Kenji Yoshihira
INFOCOM5
2008 Correlating real-time monitoring data for mobile network management
abstract
With a proliferation of new mobile data services, the complexity of wireless mobile networks is rapidly growing. While large amount of operational monitoring data such as performance measurement statistics is available, it is a great challenge to correlate such data effectively for real time performance analysis. Meantime, the dynamics of mobile applications and environments introduce another dimension of complexity for us to track the evolving system status. In this paper, we analyze the spatial and temporal correlations of Key Performance Indicators (KPIs) to track and interpret the operational status of wide-area cellular systems. We first correlate large number of raw measurements into limited number of KPIs. Further we exploit spatial and temporal correlations of these KPIs for cellular network management. We use large volume of field data collected from real cellular systems in our analysis. Experimental results demonstrate that it is promising to build a real-time data management and support system by effectively correlating KPIs.
Nanyan Jiang, Guofei Jiang, Kenji Yoshihira
WOWMOM4
2008 Understanding internet video sharing site workload: a view from data center design
abstract
In this paper we measured and analyzed the workload on Yahoo! Video, the 2nd largest U.S. video sharing site, to understand its nature and the impact on online video data center design. We discovered interesting statistical properties on both static and temporal dimensions of the workload including file duration and popularity distributions, arrival rate dynamics and predictability, and workload stationarity and burstiness. Complemented with queueing-theoretic techniques, we further extended our understanding on the measurement data with a virtual design on the workload and capacity management components of a data center assuming the same workload as measured, which reveals key results regarding the impact of Service Level Agreements (SLAs) and workload scheduling schemes on the design and operations of such large-scale video distribution systems.
Xiaozhu Kang, Hui Zhang 0002, Guofei Jiang, Xiaoqiao Meng, Kenji Yoshihira
WWW6
2008 Monitoring High-Dimensional Data for Failure Detection and Localization in Large-Scale Computing Systems
abstract
It is a major challenge to process high-dimensional measurements for failure detection and localization in large-scale computing systems. However, it is observed that in information systems, those measurements are usually located in a low-dimensional structure that is embedded in the high-dimensional space. From this perspective, a novel approach is proposed to model the geometry of underlying data generation and detect anomalies based on that model. We consider both linear and nonlinear data generation models. Two statistics, that is, the Hotelling T2and the squared prediction error (SPE), are used to reflect data variations within and outside the model. We track the probabilistic density of extracted statistics to monitor the system's health. After a failure has been detected, a localization process is also proposed to find the most suspicious attributes related to the failure. Experimental results on both synthetic data and a real e-commerce application demonstrate the effectiveness of our approach in detecting and localizing failures in computing systems.
Guofei Jiang, Kenji Yoshihira
IEEE Trans. Knowl. Data Eng.3
2007 Failure Detection in Large-Scale Internet Services by Principal Subspace Mapping
abstract
Fast and accurate failure detection is becoming essential in managing large scale Internet services. This paper proposes a novel detection approach based on the subspace mapping between system inputs and internal measurements. By exploring these contextual dependencies, our detector can initiate repair actions accurately, increasing the availability of system. While a classical statistical method, the canonical correlation analysis (CCA), is presented in the paper to achieve subspace mapping, we also propose a more advanced technique, the principal canonical correlation analysis (PCCA), to improve the performance of CCA based detector. PCCA extracts a principal subspace from internal measurements that is not only highly correlated with the inputs, but also a significant representative of original measurements. Experimental results on a J2EE based web application demonstrate that such property of PCCA is especially beneficial to failure detection tasks.
Guofei Jiang, Kenji Yoshihira
IEEE Trans. Knowl. Data Eng.3
2007 Efficient and Scalable Algorithms for Inferring Likely Invariants in Distributed Systems
abstract
Distributed systems generate a large amount of monitoring data such as log files to track their operational status. However, it is hard to correlate such monitoring data effectively across distributed systems and along observation time for system management. In previous work, we proposed a concept named flow intensity to measure the intensity with which internal monitoring data reacts to the volume of user requests. We calculated flow intensity measurements from monitoring data and proposed an algorithm to automatically search constant relationships between flow intensities measured at various points across distributed systems. If such relationships hold all the time, we regard them as invariants of the underlying systems. Invariants can be used to characterize complex systems and support various system management tasks. However, the computational complexity of the previous invariant search algorithm is high so that it may not scale well in large systems with thousands of measurements. In this paper, we propose two efficient but approximate algorithms for inferring invariants in large-scale systems. The computational complexity of new randomized algorithms is significantly reduced, and experimental results from a real system are also included to demonstrate the accuracy and efficiency of our new algorithms.
Guofei Jiang, Kenji Yoshihira
IEEE Trans. Knowl. Data Eng.3
2007 Online Tracking of Component Interactions for Failure Detection and Localization in Distributed Systems
abstract
This paper proposes a novel failure-detection approach that can handle high-dimensional observation and frequent system changes. We extract two statistics from the subspace decomposition of observations, and use the mixture of Gaussians to model their probability density. Instead of monitoring the original data, the density model of extracted statistics is adaptively updated and examined regularly to detect failures. We also present a localization method to identify the faulty components once the failure happens. Applying our technique to monitor the component interactions in an e-commerce application shows satisfactory results in detecting a variety of injected failures.
Guofei Jiang, Cristian Ungureanu, Kenji Yoshihira
IEEE Trans. Syst. Man Cybern. Part C4
2007 Multiresolution Abnormal Trace Detection Using Varied-Length n-Grams and Automata
abstract
Detection and diagnosis of faults in a large-scale distributed system is a formidable task. Interest in monitoring and using traces of user requests for fault detection has been on the rise recently. In this paper we propose novel fault detection methods based on abnormal trace detection. One essential problem is how to represent the large amount of training trace data compactly as an oracle. Our key contribution is the novel use of varied-length n-grams and automata to characterize normal traces. A new trace is compared against the learned automata to determine whether it is abnormal. We develop algorithms to automatically extract n-grams and construct multiresolution automata from training data. Further, both deterministic and multihypothesis algorithms are proposed for detection. We inspect the trace constraints of real application software and verify the existence of long n-grams. Our approach is tested in a real system with injected faults and achieves good results in experiments
Guofei Jiang, Cristian Ungureanu, Kenji Yoshihira
IEEE Trans. Syst. Man Cybern. Part C4
2006 Tracking Probabilistic Correlation of Monitoring Data for Fault Detection in Complex Systems
abstract
Due to their growing complexity, it becomes extremely difficult to detect and isolate faults in complex systems. While large amount of monitoring data can be collected from such systems for fault analysis, one challenge is how to correlate the data effectively across distributed systems and observation time. Much of the internal monitoring data reacts to the volume of user requests accordingly when user requests flow through distributed systems. In this paper, we use Gaussian mixture models to characterize probabilistic correlation between flow-intensities measured at multiple points. A novel algorithm derived from Expectation-Maximization (EM) algorithm is proposed to learn the "likely" boundary of normal data relationship, which is further used as an oracle in anomaly detection. Our recursive algorithm can adaptively estimate the boundary of dynamic data relationship and detect faults in real time. Our approach is tested in a real system with injected faults and the results demonstrate its feasibility.
Guofei Jiang, Kenji Yoshihira
DSN4
2006 Modeling and Tracking of Transaction Flow Dynamics for Fault Detection in Complex Systems
abstract
With the prevalence of Internet services and the increase of their complexity, there is a growing need to improve their operational reliability and availability. While a large amount of monitoring data can be collected from systems for fault analysis, it is hard to correlate this data effectively across distributed systems and observation time. In this paper, we analyze the mass characteristics of user requests and propose a novel approach to model and track transaction flow dynamics for fault detection in complex information systems. We measure the flow intensity at multiple checkpoints inside the system and apply system identification methods to model transaction flow dynamics between these measurements. With the learned analytical models, a model-based fault detection and isolation method is applied to track the flow dynamics in real time for fault detection. We also propose an algorithm to automatically search and validate the dynamic relationship between randomly selected monitoring points. Our algorithm enables systems to have self-cognition capability for system management. Our approach is tested in a real system with a list of injected faults. Experimental results demonstrate the effectiveness of our approach and algorithms
Guofei Jiang, Kenji Yoshihira
IEEE Trans. Dependable Secur. Comput.3
2005 Failure detection and localization in component based systems by online tracking
abstract
The increasing complexity of today's systems makes fast and accurate failure detection essential for their use in mission-critical applications. Various monitoring methods provide a large amount of data about system's behavior. Analyzing this data with advanced statistical methods holds the promise of not only detecting the errors faster, but also detecting errors which are difficult to catch with current monitoring tools. Two challenges to building such detection tools are: the high dimensionality of observation data, which makes the models expensive to apply, and frequent system changes, which make the models expensive to update. In this paper, we present algorithms to reduce the dimensionality of data in a way that makes it easy to adapt to system changes. We decompose the observation data into signal and noise subspaces. Two statistics, the Hotelling T2 score and squared prediction error (SPE) are calculated to represent the data characteristics in signal and noise subspaces respectively. Instead of tracking the original data, we use a sequentially discounting expectation maximization (SDEM) algorithm to learn the distribution of the two extracted statistics. A failure event can then be detected based on the abnormal change of the distribution. Applying our technique to component interaction data in a simple e-commerce application shows better accuracy than building independent profiles for each component. Additionally, experiments on synthetic data show that the detection accuracy is high even for changing systems.
Guofei Jiang, Cristian Ungureanu, Kenji Yoshihira
KDD4