Guofei Jiang

dblp:47/4422 · DBLP profile ↗
← Back
87ranked-venue papers
6as first author
0since 2021 · last 2020
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 25 · 2 first-authorComputer networks · 20Security and privacy · 17 · 1 first-authorArtificial intelligence and machine learning · 15 · 1 first-authorSystems, architecture and hardware · 10Software engineering, systems software and programming languages · 6Human-computer interaction and ubiquitous computing · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
13 papers
Distributed systems · 44% Cloud and datacenter computing · 29% Performance modeling and evaluation · 18%
Network and information security
6 papers
Blockchain and cryptocurrency security · 19% Hardware security and side channels · 19% Systems and software security · 19%
Databases, data mining, and information retrieval
5 papers
Data mining · 80% Graph data management · 20%
Artificial intelligence
3 papers
Deep learning architectures and training · 28% Optimization for machine learning · 24% Learning paradigms · 21%
Software engineering, system software, and programming languages
4 papers
Debugging and program repair · 26% Program analysis · 22% Operating systems · 20%
Computer networks
3 papers
Network measurement and analytics · 88% Physical-layer communications · 7% Content delivery and video streaming · 5%

Topics — the 30 heaviest of 70, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
fault tolerance
0.552016
CloudSeer: Workflow Monitoring of Cloud Infrastructures via Interleaved Logs · ASPLOS 2016
Monitoring High-Dimensional Data for Failure Detection and Localization in Large-Scale Computing Systems · IEEE Trans. Knowl. Data Eng. 2008
Failure Detection in Large-Scale Internet Services by Principal Subspace Mapping · IEEE Trans. Knowl. Data Eng. 2007
Data mining
pattern mining
0.522016
Temporal Skeletonization on Sequential Data: Patterns, Categorization, and Visualization · IEEE Trans. Knowl. Data Eng. 2016
Behavior Query Discovery in System-Generated Temporal Graphs · Proc. VLDB Endow. 2015
Blockchain and cryptocurrency security › permissioned blockchain
consortium blockchain
0.412020
Confidentiality Support over Financial Grade Consortium Blockchain · SIGMOD Conference 2020
Hardware security and side channels
trusted execution environments
0.412020
Confidentiality Support over Financial Grade Consortium Blockchain · SIGMOD Conference 2020
Machine learning › Deep learning architectures and training
attention mechanism
0.312017
A Dual-Stage Attention-Based Recurrent Neural Network for Time Series Prediction · IJCAI 2017
Machine learning › Deep learning architectures and training
recurrent neural network
0.312017
A Dual-Stage Attention-Based Recurrent Neural Network for Time Series Prediction · IJCAI 2017
Machine learning › Time series and sequential data › time series analysis
time series forecasting
0.312017
A Dual-Stage Attention-Based Recurrent Neural Network for Time Series Prediction · IJCAI 2017
Machine learning › Optimization for machine learning › sparse learning
group sparsity
0.212016
Annealed Sparsity via Adaptive and Dynamic Shrinking · KDD 2016
Machine learning › Learning paradigms
multi-task learning
0.212016
Annealed Sparsity via Adaptive and Dynamic Shrinking · KDD 2016
Machine learning › Optimization for machine learning
sparse learning
0.212016
Annealed Sparsity via Adaptive and Dynamic Shrinking · KDD 2016
Data mining
anomaly detection
0.212016
Ranking Causal Anomalies via Temporal and Dynamical Analysis on Vanishing Correlations · KDD 2016
Data mining
clustering
0.212016
Temporal Skeletonization on Sequential Data: Patterns, Categorization, and Visualization · IEEE Trans. Knowl. Data Eng. 2016
Data mining › pattern mining
sequential pattern mining
0.212016
Temporal Skeletonization on Sequential Data: Patterns, Categorization, and Visualization · IEEE Trans. Knowl. Data Eng. 2016
Digital forensics and information hiding › digital forensics
forensic analysis
0.212016
High Fidelity Data Reduction for Big Data Security Dependency Analyses · CCS 2016
Cloud and datacenter computing › cloud service management
cloud infrastructure management
0.212016
CloudSeer: Workflow Monitoring of Cloud Infrastructures via Interleaved Logs · ASPLOS 2016
Distributed systems
fault management
0.212016
Ranking Causal Anomalies via Temporal and Dynamical Analysis on Vanishing Correlations · KDD 2016
Data mining › pattern mining
discriminative pattern mining
0.212015
Behavior Query Discovery in System-Generated Temporal Graphs · Proc. VLDB Endow. 2015
Graph data management
temporal graph
0.212015
Behavior Query Discovery in System-Generated Temporal Graphs · Proc. VLDB Endow. 2015
Graph data management › temporal graph mining
temporal subgraph pattern mining
0.212015
Behavior Query Discovery in System-Generated Temporal Graphs · Proc. VLDB Endow. 2015
Data mining
time series analysis
0.212015
Efficient Long-Term Degradation Profiling in Time Series for Complex Physical Systems · KDD 2015
Systems and software security › information flow control
information flow analysis
0.212015
Checking More and Alerting Less: Detecting Privacy Leakages via Enhanced Data-flow Analysis and Peer Voting · NDSS 2015
Web and mobile security
mobile security
0.212015
SUPOR: Precise and Scalable Sensitive User Input Detection for Android Apps · USENIX Security Symposium 2015
Privacy and data protection › differential privacy › privacy auditing
privacy leak detection
0.212015
Checking More and Alerting Less: Detecting Privacy Leakages via Enhanced Data-flow Analysis and Peer Voting · NDSS 2015
Mathematical optimization › continuous optimization › nonlinear optimization
quadratic programming
0.212015
Efficient Long-Term Degradation Profiling in Time Series for Complex Physical Systems · KDD 2015
Web and mobile security › mobile security
android security
0.222015
CHEX: statically vetting Android apps for component hijacking vulnerabilities · CCS 2012
SUPOR: Precise and Scalable Sensitive User Input Detection for Android Apps · USENIX Security Symposium 2015
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.212014
Improving Semi-Supervised Target Alignment via Label-Aware Base Kernels · AAAI 2014
Machine learning › Learning paradigms
semi-supervised learning
0.212014
Improving Semi-Supervised Target Alignment via Label-Aware Base Kernels · AAAI 2014
Operating systems › kernel instrumentation
kernel tracing
0.212014
IntroPerf: transparent context-sensitive multi-layer performance inference using system stack traces · SIGMETRICS 2014
Debugging and program repair › crash report analysis
stack trace analysis
0.212014
IntroPerf: transparent context-sensitive multi-layer performance inference using system stack traces · SIGMETRICS 2014
Performance modeling and evaluation
performance diagnosis
0.212014
IntroPerf: transparent context-sensitive multi-layer performance inference using system stack traces · SIGMETRICS 2014

Methods — techniques the papers use, named apart from their topics

quadratic programming · 0.7non-negative QP · 0.7temporal analysis · 0.5network diffusion · 0.5dependency analysis · 0.5aggregation algorithm · 0.5data encryption · 0.4TEE · 0.4system stack traces · 0.4OS kernel-level tracers · 0.4swarm intelligence optimization · 0.3static analysis · 0.3encoder-decoder · 0.3dual-stage attention · 0.3NARX · 0.3streaming log checking · 0.2metric space embedding · 0.2l1 regularization · 0.2
YearPublicationVenuePosition
2020 Confidentiality Support over Financial Grade Consortium Blockchain
abstract
Confidentiality is an indispensable requirement in financial applications of blockchain technology, and supporting it along with high performance and friendly programmability is technically challenging. In this paper, we present a system design called CONFIDE to support on-chain confidentiality by leveraging Trust Execution Environment (TEE). CONFIDE's secure data transmission protocol and data encryption protocol, together with a highly efficient virtual machine run in TEE, guarantee the confidentiality in the life cycle of a transaction from end to end. CONFIDE proposes a secure data model along with an application-driven secure protocol to guarantee data confidentiality and integrity. Its smart contract language extension offers users the flexibility to define complex confidentiality models. CONFIDE is implemented as a plugin module to Antfin Blockchain's proprietary platform, and can be plugged into other blockchain platforms as well with its universal interface design. Nowadays, CONFIDE is supporting millions of commercial transactions daily on consortium blockchain running financial applications including supply chain finance, ABS, commodity provenance, and cold-chain logistics.
Ying Yan 0002, Changzheng Wei, Xuepeng Guo, Xuming Lu, Xiaofu Zheng, Chenhui Zhou, Xuyang Song, Boran Zhao, Hui Zhang 0002, Guofei Jiang
SIGMOD Conference11
2018 TGNet: Learning to Rank Nodes in Temporal Graphs
abstract
Node ranking in temporal networks are often impacted by heterogeneous context from node content, temporal, and structural dimensions. This paper introduces TGNet , a deep learning framework for node ranking in heterogeneous temporal graphs. TGNet utilizes a variant of Recurrent Neural Network to adapt context evolution and extract context features for nodes. It incorporates a novel influence network to dynamically estimate temporal and structural influence among nodes over time. To cope with label sparsity, it integrates graph smoothness constraints as a weak form of supervision. We show that the application of TGNet is feasible for large-scale networks by developing efficient learning and inference algorithms with optimization techniques. Using real-life data, we experimentally verify the effectiveness and efficiency of TGNet techniques. We also show that TGNet yields intuitive explanations for applications such as alert detection and academic impact ranking, as verified by our case study.
Qi Song 0004, Bo Zong, Yinghui Wu 0001, Lu-An Tang, Hui Zhang 0002, Guofei Jiang
CIKM6
2018 LogLens: A Real-Time Log Analysis System
abstract
Administrators of most user-facing systems depend on periodic log data to get an idea of the health and status of production applications. Logs report information, which is crucial to diagnose the root cause of complex problems. In this paper, we present a real-time log analysis system called LogLens that automates the process of anomaly detection from logs with no (or minimal) target system knowledge and user specification. In LogLens, we employ unsupervised machine learning based techniques to discover patterns in application logs, and then leverage these patterns along with the real-time log parsing for designing advanced log analytics applications. Compared to the existing systems which are primarily limited to log indexing and search capabilities, LogLens presents an extensible system for supporting both stateless and stateful log analysis applications. Currently, LogLens is running at the core of a commercial log analysis solution handling millions of logs generated from the large-scale industrial environments and reported up to 12096x man-hours reduction in troubleshooting operational problems compared to the manual approach.
Biplob Debnath, Mohiuddin Solaimani, Muhammad Ali Gulzar, Nipun Arora, Cristian Lumezanu, Jianwu Xu, Bo Zong, Hui Zhang 0002, Guofei Jiang, Latifur Khan
ICDCS9
2017 Polygravity: traffic usage accountability via coarse-grained measurements in multi-tenant data centers
abstract
Network usage accountability is critical in helping operators and customers of multi-tenant data centers deal with concerns such as capacity planning, resource allocation, hotspot detection, link failure detection, and troubleshooting. However, the cost of measurements and instrumentation to achieve flow-level accountability is non-trivial. We propose Polygravity to determine tenant traffic usage via lightweight measurements in multi-tenant data centers. We adopt a tomogravity model widely used in ISP networks, and adapt it to a multi-tenant data center environment. By integrating datacenter-specific domain knowledge, sampling-based partial estimation and gravity-based internal sinks/sources estimation, Polygravity addresses two key challenges for adapting tomogravity to a data center environment: sparse traffic matrices and internal traffic sinks/sources. We conducted extensive evaluation of our approach using realistic data center workloads. Our results show that Polygravity can determine tenant IP flow usage with less than 1% average relative error for tenants with fine-grained domain knowledge. In addition, for tenants with coarse-grained domain knowledge and with partial host-based sampling, Polygravity reduces the relative error of sampling-based estimation by 1/3.
Hyun Wook Baek, Cheng Jin 0008, Guofei Jiang, Cristian Lumezanu, Jacobus E. van der Merwe, Ning Xia
SoCC3
2017 A Dual-Stage Attention-Based Recurrent Neural Network for Time Series Prediction
abstract
The Nonlinear autoregressive exogenous (NARX) model, which predicts the current value of a time series based upon its previous values as well as the current and past values of multiple driving (exogenous) series, has been studied for decades. Despite the fact that various NARX models have been developed, few of them can capture the long-term temporal dependencies appropriately and select the relevant driving series to make predictions. In this paper, we propose a dual-stage attention-based recurrent neural network (DA-RNN) to address these two issues. In the first stage, we introduce an input attention mechanism to adaptively extract relevant driving series (a.k.a., input features) at each time step by referring to the previous encoder hidden state. In the second stage, we use a temporal attention mechanism to select relevant encoder hidden states across all time steps. With this dual-stage attention scheme, our model can not only make predictions effectively, but can also be easily interpreted. Thorough empirical studies based upon the SML 2010 dataset and the NASDAQ 100 Stock dataset demonstrate that the DA-RNN can outperform state-of-the-art methods for time series prediction.
Yao Qin 0001, Dongjin Song, Wei Cheng 0002, Guofei Jiang, Garrison W. Cottrell
IJCAI5
2017 Ranking Causal Anomalies for System Fault Diagnosis via Temporal and Dynamical Analysis on Vanishing Correlations
abstract
Detecting system anomalies is an important problem in many fields such as security, fault management, and industrial optimization. Recently, invariant network has shown to be powerful in characterizing complex system behaviours. In the invariant network, a node represents a system component and an edge indicates a stable, significant interaction between two components. Structures and evolutions of the invariance network, in particular the vanishing correlations, can shed important light on locating causal anomalies and performing diagnosis. However, existing approaches to detect causal anomalies with the invariant network often use the percentage of vanishing correlations to rank possible casual components, which have several limitations: (1) fault propagation in the network is ignored, (2) the root casual anomalies may not always be the nodes with a high percentage of vanishing correlations, (3) temporal patterns of vanishing correlations are not exploited for robust detection, and (4) prior knowledge on anomalous nodes are not exploited for (semi-)supervised detection. To address these limitations, in this article we propose a network diffusion based framework to identify significant causal anomalies and rank them. Our approach can effectively model fault propagation over the entire invariant network and can perform joint inference on both the structural and the time-evolving broken invariance patterns. As a result, it can locate high-confidence anomalies that are truly responsible for the vanishing correlations and can compensate for unstructured measurement noise in the system. Moreover, when the prior knowledge on the anomalous status of some nodes are available at certain time points, our approach is able to leverage them to further enhance the anomaly inference accuracy. When the prior knowledge is noisy, our approach also automatically learns reliable information and reduces impacts from noises. By performing extensive experiments on synthetic datasets, bank information system datasets, and coal plant cyber-physical system datasets, we demonstrate the effectiveness of our approach.
Wei Cheng 0002, Jingchao Ni, Kai Zhang 0001, Guofei Jiang, Yu Shi 0002, Xiang Zhang 0001, Wei Wang 0010
ACM Trans. Knowl. Discov. Data5
2016 CloudSeer: Workflow Monitoring of Cloud Infrastructures via Interleaved Logs
abstract
Cloud infrastructures provide a rich set of management tasks that operate computing, storage, and networking resources in the cloud. Monitoring the executions of these tasks is crucial for cloud providers to promptly find and understand problems that compromise cloud availability. However, such monitoring is challenging because there are multiple distributed service components involved in the executions. CloudSeer enables effective workflow monitoring. It takes a lightweight non-intrusive approach that purely works on interleaved logs widely existing in cloud infrastructures. CloudSeer first builds an automaton for the workflow of each management task based on normal executions, and then it checks log messages against a set of automata for workflow divergences in a streaming manner. Divergences found during the checking process indicate potential execution problems, which may or may not be accompanied by error log messages. For each potential problem, CloudSeer outputs necessary context information including the affected task automaton and related log messages hinting where the problem occurs to help further diagnosis. Our experiments on OpenStack, a popular open-source cloud infrastructure, show that CloudSeer's efficiency and problem-detection capability are suitable for online monitoring.
Pallavi Joshi, Jianwu Xu, Guoliang Jin, Hui Zhang 0002, Guofei Jiang
ASPLOS6
2016 Automated IT system failure prediction: A deep learning approach
abstract
In mission critical IT services, system failure prediction becomes increasingly important; it prevents unexpected system downtime, and assures service reliability for end users. While operational console logs record rich and descriptive information on the health status of those IT systems, existing system management technologies mostly use them in a labor-intensive forensics approach, i.e., identifying what went wrong after the fact. Recent efforts on log-based system management take an automation approach with text mining techniques, such as term frequency - inverse document frequency (TF-IDF). However, those techniques lead to a high-dimensional feature space, and are not easily generalizable to heterogeneous log formats. In this paper, we present a novel system that automatically parses streamed console logs and detects early warning signals for IT system failure prediction. In particular, our solution includes a log pattern extraction method by clustering together logs with similar format and content. We then resemble the TF-IDF idea by considering each pattern as a word and the set of patterns in each discretized epoch as a document. This leads to a feature space with significantly lower dimensionality that can provide robust signals for the status of the system. As system failures tend to occur very rare, we apply a recurrent neural network, namely, Long Short-Term Memory (LSTM), to deal with the “rarity” of labeled data in the training process. LSTM is able to capture the long-range dependency across sequences, therefore outperforms traditional supervised learning methods in our application domain. We evaluated and compared our proposed technology with state-of-the-art machine learning approaches using real log traces from two large enterprise systems. The results showed the advantage and potentials of our system in prediction of complex IT failures. To our knowledge, our work is the first that employs LSTM for log-based system failure prediction.
Ke Zhang 0013, Jianwu Xu, Martin Renqiang Min, Guofei Jiang, Konstantinos Pelechrinis, Hui Zhang 0002
IEEE BigData4
2016 High Fidelity Data Reduction for Big Data Security Dependency Analyses
abstract
Intrusive multi-step attacks, such as Advanced Persistent Threat (APT) attacks, have plagued enterprises with significant financial losses and are the top reason for enterprises to increase their security budgets. Since these attacks are sophisticated and stealthy, they can remain undetected for years if individual steps are buried in background "noise." Thus, enterprises are seeking solutions to "connect the suspicious dots" across multiple activities. This requires ubiquitous system auditing for long periods of time, which in turn causes overwhelmingly large amount of system audit events. Given a limited system budget, how to efficiently handle ever-increasing system audit logs is a great challenge. This paper proposes a new approach that exploits the dependency among system events to reduce the number of log entries while still supporting high-quality forensic analysis. In particular, we first propose an aggregation algorithm that preserves the dependency of events during data reduction to ensure the high quality of forensic analysis. Then we propose an aggressive reduction algorithm and exploit domain knowledge for further data reduction. To validate the efficacy of our proposed approach, we conduct a comprehensive evaluation on real-world auditing systems using log traces of more than one month. Our evaluation results demonstrate that our approach can significantly reduce the size of system logs and improve the efficiency of forensic analysis without losing accuracy.
Zhang Xu, Zhenyu Wu 0003, Zhichun Li, Kangkook Jee, Junghwan Rhee, Xusheng Xiao, Fengyuan Xu, Haining Wang 0001, Guofei Jiang
CCS9
2016 LogMine: Fast Pattern Recognition for Log Analytics
abstract
Modern engineering incorporates smart technologies in all aspects of our lives. Smart technologies are generating terabytes of log messages every day to report their status. It is crucial to analyze these log messages and present usable information (e.g. patterns) to administrators, so that they can manage and monitor these technologies. Patterns minimally represent large groups of log messages and enable the administrators to do further analysis, such as anomaly detection and event prediction. Although patterns exist commonly in automated log messages, recognizing them in massive set of log messages from heterogeneous sources without any prior information is a significant undertaking. We propose a method, named LogMine, that extracts high quality patterns for a given set of log messages. Our method is fast, memory efficient, accurate, and scalable. LogMine is implemented in map-reduce framework for distributed platforms to process millions of log messages in seconds. LogMine is a robust method that works for heterogeneous log messages generated in a wide variety of systems. Our method exploits algorithmic techniques to minimize the computational overhead based on the fact that log messages are always automatically generated. We evaluate the performance of LogMine on massive sets of log messages generated in industrial applications. LogMine has successfully generated patterns which are as good as the patterns generated by exact and unscalable method, while achieving a 500× speedup. Finally, we describe three applications of the patterns generated by LogMine in monitoring large scale industrial systems.
Hossein Hamooni, Biplob Debnath, Jianwu Xu, Hui Zhang 0002, Guofei Jiang, Abdullah Mueen
CIKM5
2016 Ranking Causal Anomalies via Temporal and Dynamical Analysis on Vanishing Correlations
abstract
Modern world has witnessed a dramatic increase in our ability to collect, transmit and distribute real-time monitoring and surveillance data from large-scale information systems and cyber-physical systems. Detecting system anomalies thus attracts significant amount of interest in many fields such as security, fault management, and industrial optimization. Recently, invariant network has shown to be a powerful way in characterizing complex system behaviours. In the invariant network, a node represents a system component and an edge indicates a stable, significant interaction between two components. Structures and evolutions of the invariance network, in particular the vanishing correlations, can shed important light on locating causal anomalies and performing diagnosis. However, existing approaches to detect causal anomalies with the invariant network often use the percentage of vanishing correlations to rank possible casual components, which have several limitations: 1) fault propagation in the network is ignored; 2) the root casual anomalies may not always be the nodes with a high-percentage of vanishing correlations; 3) temporal patterns of vanishing correlations are not exploited for robust detection. To address these limitations, in this paper we propose a network diffusion based framework to identify significant causal anomalies and rank them. Our approach can effectively model fault propagation over the entire invariant network, and can perform joint inference on both the structural, and the time-evolving broken invariance patterns. As a result, it can locate high-confidence anomalies that are truly responsible for the vanishing correlations, and can compensate for unstructured measurement noise in the system. Extensive experiments on synthetic datasets, bank information system datasets, and coal plant cyber-physical system datasets demonstrate the effectiveness of our approach.
Wei Cheng 0002, Kai Zhang 0001, Guofei Jiang, Zhengzhang Chen, Wei Wang 0010
KDD4
2016 Annealed Sparsity via Adaptive and Dynamic Shrinking
abstract
Sparse learning has received tremendous amount of interest in high-dimensional data analysis due to its model interpretability and the low-computational cost. Among the various techniques, adaptive l1-regularization is an effective framework to improve the convergence behaviour of the LASSO, by using varying strength of regularization across different features. In the meantime, the adaptive structure makes it very powerful in modelling grouped sparsity patterns as well, being particularly useful in high-dimensional multi-task problems. However, choosing an appropriate, global regularization weight is still an open problem. In this paper, inspired by the annealing technique in material science, we propose to achieve "annealed sparsity" by designing a dynamic shrinking scheme that simultaneously optimizes the regularization weights and model coefficients in sparse (multi-task) learning. The dynamic structures of our algorithm are twofold. Feature-wise (spatially), the regularization weights are updated interactively with model coefficients, allowing us to improve the global regularization structure. Iteration-wise (temporally), such interaction is coupled with gradually boosted l1-regularization by adjusting an equality norm-constraint, achieving an annealing effect to further improve model selection. This renders interesting shrinking behaviour in the whole solution path. Our method competes favorably with state-of-the-art methods in sparse (multi-task) learning. We also apply it in expression quantitative trait loci analysis (eQTL), which gives useful biological insights in human cancer (melanoma) study.
Kai Zhang 0001, Shandian Zhe, Chaoran Cheng, Zhi Wei 0001, Zhengzhang Chen, Guofei Jiang, Yuan Qi 0001, Jieping Ye
KDD7
2016 Characterizing Rule Compression Mechanisms in Software-Defined Networks
Curtis Yu, Cristian Lumezanu, Harsha V. Madhyastha, Guofei Jiang
PAM4
2016 Detecting Stack Layout Corruptions with Robust Stack Unwinding
Yangchun Fu, Junghwan Rhee, Zhiqiang Lin 0001, Zhichun Li, Hui Zhang 0002, Guofei Jiang
RAID6
2016 Integrating Community and Role Detection in Information Networks
abstract
Community detection and role detection in information networks have received wide attention recently, where the former aims to detect the groups of nodes that are closely connected to each other and the latter aims to discover the underlying roles of nodes in the network. Traditional studies treat these two problems as orthogonal issues and propose algorithms for these two tasks separately. In this paper, we propose to integrate communities and roles in a unified model and detect both of them simultaneously for information networks. Intuitively, (1) correctly detecting the communities in a network will lead to the success of detecting roles of nodes, such as opinion leaders and followers in social networks; and (2) correctly identifying the roles of the nodes will lead to a better network modeling and thus a better detection of communities. A novel probabilistic network model, the Mixed Membership Community and Role model (MMCR), is then proposed, which models the latent community and role of each node at the same time, and the probability of links are defined accordingly. By testing our model on synthetic networks and two real-world networks, we demonstrate that our approach leads to better performance for both community detection and role detection. Moreover, our model has a better interpretation for link generation in networks according to the link prediction task.
Ting Chen 0007, Lu-An Tang, Yizhou Sun, Zhengzhang Chen, Guofei Jiang
SDM6
2016 Enhancing semi-supervised learning through label-aware base kernels
Qiaojun Wang, Kai Zhang 0001, Zhengzhang Chen, Dequan Wang, Guofei Jiang, Ivan Marsic
Neurocomputing5
2016 Temporal Skeletonization on Sequential Data: Patterns, Categorization, and Visualization
abstract
Sequential pattern analysis aims at finding statistically relevant temporal structures where the values are delivered in a sequence. With the growing complexity of real-world dynamic scenarios, more and more symbols are often needed to encode the sequential values. This is so-called “curse of cardinality”, which can impose significant challenges to the design of sequential analysis methods in terms of computational efficiency and practical use. Indeed, given the overwhelming scale and the heterogeneous nature of the sequential data, new visions and strategies are needed to face the challenges. To this end, in this paper, we propose a “temporal skeletonization” approach to proactively reduce the cardinality of the representation for sequences by uncovering significant, hidden temporal structures. The key idea is to summarize the temporal correlations in an undirected graph, and use the “skeleton” of the graph as a higher granularity on which hidden temporal patterns are more likely to be identified. As a consequence, the embedding topology of the graph allows us to translate the rich temporal content into a metric space. This opens up new possibilities to explore, quantify, and visualize sequential data. Our approach has shown to greatly alleviate the curse of cardinality in challenging tasks of sequential pattern mining and clustering. Evaluation on a business-to-business (B2B) marketing application demonstrates that our approach can effectively discover critical buying paths from noisy customer event data.
Chuanren Liu, Kai Zhang 0001, Hui Xiong 0001, Guofei Jiang, Qiang Yang 0001
IEEE Trans. Knowl. Data Eng.4
2015 Discover and Tame Long-running Idling Processes in Enterprise Systems
abstract
Reducing attack surface is an effective preventive measure to strengthen security in large systems. However, it is challenging to apply this idea in an enterprise environment where systems are complex and evolving over time. In this paper, we empirically analyze and measure a real enterprise to identify unused services that expose attack surface. Interestingly, such unused services are known to exist and summarized by security best practices, yet such solutions require significant manual effort.
Jun Wang 0141, Zhiyun Qian, Zhichun Li, Zhenyu Wu 0003, Junghwan Rhee, Xia Ning, Peng Liu 0005, Guofei Jiang
AsiaCCS8
2015 Efficient Long-Term Degradation Profiling in Time Series for Complex Physical Systems
abstract
The long term operation of physical systems inevitably leads to their wearing out, and may cause degradations in performance or the unexpected failure of the entire system. To reduce the possibility of such unanticipated failures, the system must be monitored for tell-tale symptoms of degradation that are suggestive of imminent failure. In this work, we introduce a novel time series analysis technique that allows the decomposition of the time series into trend and fluctuation components, providing the monitoring software with actionable information about the changes of the system's behavior over time. We analyze the underlying problem and formulate it to a Quadratic Programming (QP) problem that can be solved with existing QP-solvers. However, when the profiling resolution is high, as generally required by real-world applications, such a decomposition becomes intractable to general QP-solvers. To speed up the problem solving, we further transform the problem and present a novel QP formulation, Non-negative QP, for the problem and demonstrate a tractable solution that bypasses the use of slow general QP-solvers. We demonstrate our ideas on both synthetic and real datasets, showing that our method allows us to accurately extract the degradation phenomenon of time series. We further demonstrate the generality of our ideas by applying them beyond classic machine prognostics to problems in identifying the influence of news events on currency exchange rates and stock prices. We fully implement our profiling system and deploy it into several physical systems, such as chemical plants and nuclear power plants, and it greatly helps detect the degradation phenomenon, and diagnose the corresponding components.
Liudmila Ulanova, Tan Yan, Guofei Jiang, Eamonn J. Keogh, Kai Zhang 0001
KDD4
2015 Checking More and Alerting Less: Detecting Privacy Leakages via Enhanced Data-flow Analysis and Peer Voting
Kangjie Lu, Zhichun Li, Vasileios P. Kemerlis, Zhenyu Wu 0003, Long Lu, Cong Zheng, Zhiyun Qian, Wenke Lee, Guofei Jiang
NDSS9
2015 Software-Defined Latency Monitoring in Data Center Networks
Curtis Yu, Cristian Lumezanu, Abhishek B. Sharma, Guofei Jiang, Harsha V. Madhyastha
PAM5
2015 From Categorical to Numerical: Multiple Transitive Distance Learning and Embedding
abstract
Categorical data are ubiquitous in real-world databases. However, due to the lack of an intrinsic proximity measure, many powerful algorithms for numerical data analysis may not work well on their categorical counterparts, making it a bottleneck in practical applications. In this paper, we propose a novel method to transform categorical data to numerical representations, so that abundant numerical learning methods can be exploited in categorical data mining. Our key idea is to learn a pairwise dissimilarity among categorical symbols, henceforth a continuous embedding, which can then be used for subsequent numerical treatment. There are two important criteria for learning the dissimilarities. First, it should capture the important “transitivity” which has shown to be particularly useful in measuring the proximity relation in categorical data. Second, the pairwise sample geometry arising from the learned symbol distances should be maximally consistent with prior knowledge (e.g., class labels) to obtain a good generalization performance. We achieve them through multiple transitive distance learning and embedding. Encouraging results are observed on a number of benchmark classification tasks against state-of-the-art.
Kai Zhang 0001, Qiaojun Wang, Zhengzhang Chen, Ivan Marsic, Vipin Kumar 0001, Guofei Jiang, Jie Zhang 0012
SDM6
2015 SUPOR: Precise and Scalable Sensitive User Input Detection for Android Apps
Jianjun Huang 0001, Zhichun Li, Xusheng Xiao, Zhenyu Wu 0003, Kangjie Lu, Xiangyu Zhang 0001, Guofei Jiang
USENIX Security Symposium7
2015 Behavior Query Discovery in System-Generated Temporal Graphs
abstract
Computer system monitoring generates huge amounts of logs that record the interaction of system entities. How to query such data to better understand system behaviors and identify potential system risks and malicious behaviors becomes a challenging task for system administrators due to the dynamics and heterogeneity of the data. System monitoring data are essentially heterogeneous temporal graphs with nodes being system entities and edges being their interactions over time. Given the complexity of such graphs, it becomes time-consuming for system administrators to manually formulate useful queries in order to examine abnormal activities, attacks, and vulnerabilities in computer systems. In this work, we investigate how to query temporal graphs and treat query formulation as a discriminative temporal graph pattern mining problem. We introduce TGMiner to mine discriminative patterns from system logs, and these patterns can be taken as templates for building more complex queries. TGMiner leverages temporal information in graphs to prune graph patterns that share similar growth trend without compromising pattern quality. Experimental results on real system data show that TGMiner is 6-32 times faster than baseline methods. The discovered patterns were verified by system experts; they achieved high precision (97%) and recall (91%).
Bo Zong, Xusheng Xiao, Zhichun Li, Zhenyu Wu 0003, Zhiyun Qian, Xifeng Yan, Ambuj K. Singh, Guofei Jiang
Proc. VLDB Endow.8
2015 A Framework of Mining Trajectories from Untrustworthy Data in Cyber-Physical System
abstract
A cyber-physical system (CPS) integrates physical (i.e., sensor) devices with cyber (i.e., informational) components to form a context-sensitive system that responds intelligently to dynamic changes in real-world situations. The CPS has wide applications in scenarios such as environment monitoring, battlefield surveillance, and traffic control. One key research problem of CPS is called mining lines in the sand . With a large number of sensors (sand) deployed in a designated area, the CPS is required to discover all trajectories (lines) of passing intruders in real time. There are two crucial challenges that need to be addressed: (1) the collected sensor data are not trustworthy, and (2) the intruders do not send out any identification information. The system needs to distinguish multiple intruders and track their movements. This study proposes a method called LiSM (Line-in-the-Sand Miner) to discover trajectories from untrustworthy sensor data. LiSM constructs a watching network from sensor data and computes the locations of intruder appearances based on the link information of the network. The system retrieves a cone model from the historical trajectories to track multiple intruders. Finally, the system validates the mining results and updates sensors’ reliability scores in a feedback process. In addition, LoRM (Line-on-the-Road Miner) is proposed for trajectory discovery on road networks— mining lines on the roads . LoRM employs a filtering-and-refinement framework to reduce the distance computational overhead on road networks and uses a shortest-path-measure to track intruders. The proposed methods are evaluated with extensive experiments on big datasets. The experimental results show that the proposed methods achieve higher accuracy and efficiency in trajectory mining tasks.
Lu-An Tang, Xiao Yu 0007, Quanquan Gu, Jiawei Han 0001, Guofei Jiang, Alice Leung, Thomas La Porta
ACM Trans. Knowl. Discov. Data5
2014 Improving Semi-Supervised Target Alignment via Label-Aware Base Kernels
Qiaojun Wang, Kai Zhang 0001, Guofei Jiang, Ivan Marsic
AAAI3
2014 Extracting discriminative shapelets from heterogeneous sensor data
abstract
We study the problem of identifying discriminative features in Big Data arising from heterogeneous sensors. We highlight the heterogeneity in sensor data from engineering applications and the challenges involved in automatically extracting only the most interesting features from large datasets. We formulate this problem as that of classification of multivariate time series and design shapelet-based algorithms for this task. We design a novel approach, called Shapelet Forests (SF), which combines shapelet extraction with feature selection. We evaluate our proposed method with other approaches for mining shapelets from multivariate time series using data from real-world engineering applications. Quantitative analysis of the experiments shows that SF performs better than the baseline approaches and achieves high classification accuracy. In addition, the method enables identification of noisy sensors from multivariate data and discounts their use for classification.
Om Prasad Patri, Abhishek B. Sharma, Guofei Jiang, Anand V. Panangadan, Viktor Prasanna 0001
IEEE BigData4
2014 DeltaPath: Precise and Scalable Calling Context Encoding
Qiang Zeng 0001, Junghwan Rhee, Hui Zhang 0002, Nipun Arora, Guofei Jiang, Peng Liu 0005
CGO5
2014 Software system performance debugging with kernel events feature guidance
abstract
To diagnose performance problems in production systems, many OS kernel-level monitoring and analysis tools have been proposed. Using low level kernel events provides benefits in efficiency and transparency to monitor application software. On the other hand, such approaches miss application-specific semantic information which can be effective to differentiate the trace patterns from distinct application logic. This paper introduces new trace analysis techniques based on event features to improve kernel event based performance diagnosis tools. Our prototype, AppDiff, is based on two analysis features: system resource features convert kernel events to resource usage metrics, thereby enabling the detection of various performance anomalies in a unified way; program behavior features infer the application logic behind the low level events. By using these features and conditional probability, AppDiff can detect outliers and improve the diagnosis of application performance.
Junghwan Rhee, Hui Zhang 0002, Nipun Arora, Guofei Jiang, Kenji Yoshihira
NOMS4
2014 Uscope: A scalable unified tracer from kernel to user space
abstract
Unified tracing is the process of collecting trace logs across the boundary of kernel and user spaces, and has been used to understand the in-depth correspondence between low level events and application program context for diagnosing system failures and performance problems. Crossing the boundary from the kernel space to a user space to collect trace events from dual spaces imposes challenges compared to crossing the boundary in the other way from a user space to the kernel space due to multiple scheduled programs and diverse code layouts in the user space regarding the tracing target. In this paper, we propose a novel unified tracing system called Uscope to systematically trace kernel and unprecedented user code with low overhead. The key idea is to use an efficient variant of stack walking. Uscope lowers stack walking overhead by adjusting the scope of walking in two ways: (1) a highly configurable focus within the call stack, and (2) a per-application tracing that systematically tracks a dynamic set of new, exiting, or transforming processes and threads of an application software. This system is realized by using a flexible stack walking algorithm and a runtime kernel structure, Trace Map. These key features lead to low run-time overhead under 6% relative to native execution on a set of widely used benchmarks.
Junghwan Rhee, Hui Zhang 0002, Nipun Arora, Guofei Jiang, Kenji Yoshihira
NOMS4
2014 CLUE: System trace analytics for cloud service performance diagnosis
abstract
In this paper, we present CLUE, a system event analytics tool for black-box performance diagnosis in production Cloud Computing systems. CLUE provides an unified and extensible means of profiling service transactional behaviors, and builds structured data called event sketches. CLUE further offers a set of analytic tools for summarizing and analyzing event sketches by integrating data mining and statistical analysis. CLUE has been developed in NEC as an internal tool and applied in diagnosing a diverse set of real performance problems for multi-tiered IT applications running on multi-core servers of major platforms including Linux (Redhat, Fedora), Unix (HP-UX), and Windows (Windows Server 2008). We demonstrated the evaluation of our framework on real-world IT systems, and showed how it can enable visibility and effective diagnosis of service system performance problems.
Hui Zhang 0002, Junghwan Rhee, Nipun Arora, Sahan Gamage, Guofei Jiang, Kenji Yoshihira, Dongyan Xu
NOMS5
2014 IntroPerf: transparent context-sensitive multi-layer performance inference using system stack traces
abstract
Performance bugs are frequently observed in commodity software. While profilers or source code-based tools can be used at development stage where a program is diagnosed in a well-defined environment, many performance bugs survive such a stage and affect production runs. OS kernel-level tracers are commonly used in post-development diagnosis due to their independence from programs and libraries; however, they lack detailed program-specific metrics to reason about performance problems such as function latencies and program contexts. In this paper, we propose a novel performance inference system, called IntroPerf, that generates fine-grained performance information -- like that from application profiling tools -- transparently by leveraging OS tracers that are widely available in most commodity operating systems. With system stack traces as input, IntroPerf enables transparent context-sensitive performance inference, and diagnoses application performance in a multi-layered scope ranging from user functions to the kernel. Evaluated with various performance bugs in multiple open source software projects, IntroPerf automatically ranks potential internal and external root causes of performance bugs with high accuracy without any prior knowledge about or instrumentation on the subject software. Our results show IntroPerf's effectiveness as a lightweight performance introspection tool for post-development diagnosis.
Junghwan Rhee, Hui Zhang 0002, Nipun Arora, Guofei Jiang, Xiangyu Zhang 0001, Dongyan Xu
SIGMETRICS5
2014 Proactive Workload Management in Hybrid Cloud Computing
abstract
The hindrances to the adoption of public cloud computing services include service reliability, data security and privacy, regulation compliant requirements, and so on. To address those concerns, we propose a hybrid cloud computing model which users may adopt as a viable and cost-saving methodology to make the best use of public cloud services along with their privately-owned (legacy) data centers. As the core of this hybrid cloud computing model, an intelligent workload factoring service is designed for proactive workload management. It enables federation between on- and off-premise infrastructures for hosting Internet-based applications, and the intelligence lies in the explicit segregation of base workload and flash crowd workload, the two naturally different components composing the application workload. The core technology of the intelligent workload factoring service is a fast frequent data item detection algorithm, which enables factoring incoming requests not only on volume but also on data content, upon a changing application data popularity. Through analysis and extensive evaluation with real-trace driven simulations and experiments on a hybrid testbed consisting of local computing platform and Amazon Cloud service platform, we showed that the proactive workload management technology can enable reliable workload prediction in the base workload zone (with simple statistical methods), achieve resource efficiency (e.g., 78% higher server capacity than that in base workload zone) and reduce data cache/replication overhead (up to two orders of magnitude) in the flash crowd workload zone, and react fast (with an X^2 speed-up factor) to the changing application data popularity upon the arrival of load spikes.
Hui Zhang 0002, Guofei Jiang, Kenji Yoshihira
IEEE Trans. Netw. Serv. Manag.2
2013 Modeling heterogeneous time series dynamics to profile big sensor data in complex physical systems
abstract
While a massive amount of time series can now be collected in many physical systems, it is a challenge to build an analytic model that can correctly profile the data because those time series usually exhibit various behaviors. In this paper we propose an integrated method to address the heterogeneity issue in modeling big time series data. We first extracts relevant features to summarize the underlying dynamics of those series. We present both linear and nonlinear feature extraction techniques, as well as a procedure to determine the right extraction method for individual time series. Given extracted features, our method further models the trajectory pattern of time series in the feature space. Both a regression based and a density based method are presented to profile different types of feature trajectories. Experimental results in a real power plant illustrate that our feature extraction and trajectory model are effective to profile various time series. Our method has been used to successfully detect anomalies in the system.
Bin Liu 0045, Abhishek B. Sharma, Guofei Jiang, Hui Xiong 0001
IEEE BigData4
2013 Fault detection and localization in distributed systems using invariant relationships
abstract
Recent advances in sensing and communication technologies enable us to collect round-the-clock monitoring data from a wide-array of distributed systems including data centers, manufacturing plants, transportation networks, automobiles, etc. Often this data is in the form of time series collected from multiple sensors (hardware as well as software based). Previously, we developed a time-invariant relationships based approach that uses Auto-Regressive models with eXogenous input (ARX) to model this data. A tool based on our approach has been effective for fault detection and capacity planning in distributed systems. In this paper, we first describe our experience in applying this tool in real-world settings. We also discuss the challenges in fault localization that we face when using our tool, and present two approaches - a spatial approach based on invariant graphs and a temporal approach based on expected broken invariant patterns - that we developed to address this problem.
Abhishek B. Sharma, Kenji Yoshihira, Guofei Jiang
DSN5
2013 NetFuse: Short-circuiting traffic surges in the cloud
abstract
Modern cloud and data center platforms suffer failures and performance degradation from large traffic surges caused by both external (e.g., DDoS attacks) or internal (e.g., workload changes, operator errors, routing misconfigurations) factors. If not mitigated, traffic overload could have significant financial and availability implications for cloud providers. In this paper, we propose NetFuse, a mechanism to protect against traffic overload in OpenFlow-based data center networks. NetFuse is (1) scalable because it uses passively-collected OpenFlow control messages to detect active network flows; (2) accurate because it uses multi-dimensional flow aggregation to determine the right criteria to combine network flows that lead to overloading behavior; and (3) effective in limiting the damage of surges while not affecting the normal traffic because it uses a toxin-antitoxinlike mechanism to adaptively shape the rate of the flow based on application feedback. Experimental results on a real OpenFlow testbed show that NetFuse is effective in identifying and isolating misbehaving traffic with a small false positive rate (<; 9%).
Yueping Zhang, Vishal K. Singh, Cristian Lumezanu, Guofei Jiang
ICC5
2013 Diagnosing Data Center Behavior Flow by Flow
abstract
Multi-tenant data centers are complex environments, running thousands of applications that compete for the same infrastructure resources and whose behavior is guided by (sometimes) divergent configurations. Small workload changes or simple operator tasks may yield unpredictable results and lead to expensive failures and performance degradation. In this paper, we propose a holistic approach for detecting operational problems in data centers. Our framework, FlowDiff, collects information from all entities involved in the operation of a data center -- applications, operators, and infrastructure -- and continually builds behavioral models for the operation. By comparing current models with pre-computed, known-to-be-stable models, FlowDiff is able to detect many operational problems, ranging from host and network failures to unauthorized access. FlowDiff also identifies common system operations (e.g., VM migration, software upgrades) to validate the behavior changes against planned operator tasks. We show that using passive measurements on control traffic from programmable switches to a centralized controller is sufficient to build strong behavior models; FlowDiff does not require active measurements or expensive server instrumentation. Our experimental results using NEC data center testbed, Amazon EC2, and simulations demonstrate that FlowDiff is effective and robust in detecting anomalous behavior. FlowDiff scales well with the number of applications running in the data center and their traffic volume.
Ahsan Arefin, Vishal K. Singh, Guofei Jiang, Yueping Zhang, Cristian Lumezanu
ICDCS3
2013 Efficient Invariant Search for Distributed Information Systems
abstract
In today's distributed information systems, a large amount of monitoring data such as log files have been collected. These monitoring data at various points of a distributed information system provide unparallel opportunities for us to characterize and track the information system via effectively correlating all monitoring data across the distributed system. Jiang1 proposed a concept named flow intensity to measure the intensity with which the monitoring data reacts to the volume of different user requests. The Autoregressive model with exogenous inputs (ARX) was used to quantify the relationship between each pair of flow intensity measured at various points across distributed systems. If such relationships hold all the time, they are considered as invariants of the underlying systems. Such invariants have been successfully used to characterize complex systems and support various system management tasks, such as system fault detection and localization. However, it is very time-consuming to search the complete set of invariants of large scale systems and existing algorithms are not scalable for thousands of flow intensity measurements. To this end, in this paper, we develop effective pruning techniques based on the identified upper bounds. Accordingly, two efficient algorithms are proposed to search the complete set of invariants based on the pruning techniques. Finally we demonstrate the efficiency and effectiveness of our algorithms with both real-world and synthetic data sets.
Guofei Jiang
ICDM2
2013 Network-aware coordination of virtual machine migrations in enterprise data centers and clouds
Guofei Jiang, Yueping Zhang
IM3
2013 Predictive VM consolidation on multiple resources: Beyond load balancing
abstract
Effective consolidation of different applications on common resources is often akin to black art as application performance interference may result in unpredictable system and workload delays. In this paper we consider the problem of fair load balancing on multiple servers within a virtualized data center setting. We especially focus on multi-tiered applications with different resource demands per tier and address the problem on how to best match each application tier on each resource, such that performance interference is minimized. To address this problem, we propose a two-step approach. First, a fair load balancing scheme assigns different virtual machines (VMs) across different servers; this process is formulated as a multi-dimensional vector scheduling problem that uses a new polynomial-time approximation scheme (PTAS) to minimize the maximum utilization across all server resources and results in multiple load balancing solutions. Second, a queueing network analytic model is applied on the proposed min-max solutions in order to select the optimal one. We experimentally evaluate the proposed two-stage mechanism using a Xen virtualization testbed that hosts multiple RUBiS multi-tier applications. Experimental results show that the proposed mechanism is robust as it always predicts the optimal consolidation strategy.
Hui Zhang 0002, Evgenia Smirni, Guofei Jiang, Kenji Yoshihira
IWQoS4
2013 iProbe: A lightweight user-level dynamic instrumentation tool
abstract
We introduce a new hybrid instrumentation tool for dynamic application instrumentation called iProbe, which is flexible and has low overhead. iProbe takes a novel 2-stage design, and offloads much of the dynamic instrumentation complexity to an offline compilation stage. It leverages standard compiler flags to introduce “place-holders” for hooks in the program executable. Then it utilizes an efficient user-space “HotPatching” mechanism which modifies the functions to be traced and enables execution of instrumented code in a safe and secure manner. In its evaluation on a micro-benchmark and SPEC CPU2006 benchmark applications, the iProbe prototype achieved the instrumentation overhead an order of magnitude lower than existing state-of-the-art dynamic instrumentation tools like SystemTap and DynInst.
Nipun Arora, Hui Zhang 0002, Junghwan Rhee, Kenji Yoshihira, Guofei Jiang
ASE5
2013 FlowSense: Monitoring Network Utilization with Zero Measurement Cost
Curtis Yu, Cristian Lumezanu, Yueping Zhang, Vishal K. Singh, Guofei Jiang, Harsha V. Madhyastha
PAM5
2013 Automating Cloud Network Optimization and Evolution
abstract
With the ever-increasing number and complexity of applications deployed in data centers, the underlying network infrastructure can no longer sustain such a trend and exhibits several problems, such as resource fragmentation and low bisection bandwidth. In pursuit of a real-world applicable cloud network (CN) optimization approach that continuously maintains balanced network performance with high cost effectiveness, we design a topology independent resource allocation and optimization approach, NetDEO. Based on a swarm intelligence optimization model, NetDEO improves the scalability of the CN by relocating virtual machines (VMs) and matching resource demand and availability. NetDEO is capable of (1) incrementally optimizing an existing VM placement in a data center; (2) deriving optimal deployment plans for newly added VMs; and (3) providing hardware upgrade suggestions, and allowing the CN to evolve as the workload changes over time. We evaluate the performance of NetDEO using realistic workload traces and simulated large-scale CN under various topologies.
Zhenyu Wu 0003, Yueping Zhang, Vishal K. Singh, Guofei Jiang, Haining Wang 0001
IEEE J. Sel. Areas Commun.4
2013 Ranking Metric Anomaly in Invariant Networks
abstract
The management of large-scale distributed information systems relies on the effective use and modeling of monitoring data collected at various points in the distributed information systems. A traditional approach to model monitoring data is to discover invariant relationships among the monitoring data. Indeed, we can discover all invariant relationships among all pairs of monitoring data and generate invariant networks, where a node is a monitoring data source (metric) and a link indicates an invariant relationship between two monitoring data. Such an invariant network representation can help system experts to localize and diagnose the system faults by examining those broken invariant relationships and their related metrics, since system faults usually propagate among the monitoring data and eventually lead to some broken invariant relationships. However, at one time, there are usually a lot of broken links (invariant relationships) within an invariant network. Without proper guidance, it is difficult for system experts to manually inspect this large number of broken links. To this end, in this article, we propose the problem of ranking metrics according to the anomaly levels for a given invariant network, while this is a nontrivial task due to the uncertainties and the complex nature of invariant networks. Specifically, we propose two types of algorithms for ranking metric anomaly by link analysis in invariant networks. Along this line, we first define two measurements to quantify the anomaly level of each metric, and introduce the m R ank algorithm. Also, we provide a weighted score mechanism and develop the g R ank algorithm, which involves an iterative process to obtain a score to measure the anomaly levels. In addition, some extended algorithms based on m R ank and g R ank algorithms are developed by taking into account the probability of being broken as well as noisy links. Finally, we validate all the proposed algorithms on a large number of real-world and synthetic data sets to illustrate the effectiveness and efficiency of different algorithms.
Yong Ge 0001, Guofei Jiang, Hui Xiong 0001
ACM Trans. Knowl. Discov. Data2
2012 CHEX: statically vetting Android apps for component hijacking vulnerabilities
abstract
An enormous number of apps have been developed for Android in recent years, making it one of the most popular mobile operating systems. However, the quality of the booming apps can be a concern [4]. Poorly engineered apps may contain security vulnerabilities that can severally undermine users' security and privacy. In this paper, we study a general category of vulnerabilities found in Android apps, namely the component hijacking vulnerabilities. Several types of previously reported app vulnerabilities, such as permission leakage, unauthorized data access, intent spoofing, and etc., belong to this category.
Long Lu, Zhichun Li, Zhenyu Wu 0003, Wenke Lee, Guofei Jiang
CCS5
2012 Application dependency discovery using matrix factorization
abstract
Driven by the large-scale growth of applications deployment in data centers and complicated interactions between service components, automated application dependency discovery becomes essential to daily system management and operation. In this paper, we present ADD, which extracts dependency paths for each application by decomposing the application-layer connectivity graph inferred from passive network monitoring data. ADD utilizes a series of statistical techniques and is based on the combination of global observation of application traffic matrix in the data center and local observation of traffic volumes at small time scales on each server. Compared to existing approaches, ADD is especially effective in the presence of overlapping and multi-hop applications and resilient to data loss and estimation errors.
Vishal K. Singh, Yueping Zhang, Guofei Jiang
IWQoS4
2012 NetDEO: Automating network design, evolution, and optimization
abstract
With the ever-increasing number and complexity of applications deployed in data centers, the underlying network infrastructure can no longer sustain such a trend and exhibits several problems, such as resource fragmentation and low bisection bandwidth. In pursuit of a real-world applicable data center network (DCN) optimization approach that continuously maintains balanced network performance with high cost effectiveness, we design a topology independent resource allocation and optimization approach, NetDEO. Based on a swarm intelligence optimization model, NetDEO improves the scalability of the DCN by relocating virtual machines (VMs) and matching resource demand and availability. NetDEO is capable of (1) incrementally optimizing an existing VM placement in a data center; (2) deriving optimal deployment plans for newly added VMs; and (3) providing hardware upgrade suggestions and allowing the DCN to evolve as the workload changes over time. We evaluate the performance of NetDEO using realistic workload traces and simulated large-scale DCN under various topologies.
Zhenyu Wu 0003, Yueping Zhang, Vishal K. Singh, Guofei Jiang, Haining Wang 0001
IWQoS4
2011 Application Behavior Mapping across Heterogeneous Hardware Platforms
abstract
Predicting the application behavior such as its resource utilization in a new hardware machine is becoming an urgent issue as the increasing number of servers with various configurations show up in data centers and clouds. Current two categories of approaches, the test bed evaluation based and the software simulation based methods, both have certain shortcomings. While the test bed evaluation based approaches suffer from the lack of measurement data to build the prediction model, the simulation based methods intrinsically introduce uncertainties and errors in the data. In order to overcome those issues, this paper proposes a new solution that combines the current two separate processes. We develop a generalized regression model with L1 penalty to predict the application behavior from software simulation. Meanwhile we also use evaluations on real hardware instances to improve the model obtained from simulation. Our model improvement is grounded on the Bayesian learning theory, which elegantly embeds outcomes from both simulation and real evaluation stages into the final prediction. Experimental results show the higher prediction accuracy of our method compared with current techniques.
Guofei Jiang, Kenji Yoshihira
DASC3
2011 Effective VM sizing in virtualized data centers
abstract
In this paper, we undertake the problem of server consolidation in virtualized data centers from the perspective of approximation algorithms. We formulate server consolidation as a stochastic bin packing problem, where the server capacity and an allowed server overflow probability p are given, and the objective is to assign VMs to as few physical servers as possible, and the probability that the aggregated load of a physical server exceeds the server capacity is at most p.
Ming Chen 0002, Hui Zhang 0002, Ya-Yunn Su, Guofei Jiang, Kenji Yoshihira
Integrated Network Management5
2011 CloudInsight: Shedding Light on the Cloud
abstract
Cloud computing provides a revolutionary new computing paradigm for deploying enterprise applications and Internet services. Rather than operating their own data centers, today cloud users run their applications on the remote cloud infrastructures that are owned and managed by cloud providers. However, the cloud computing paradigm also introduces some new challenges in system management. Cloud users create virtual machine instances to run their specific application logic without knowing the underlying physical infrastructure. On the other side, cloud providers manage and operate their cloud infrastructures without knowing their customers' applications. Due to the decoupled ownership of applications and infrastructures, if a problem occurs, there is no visibility for either cloud users or providers to understand the whole context of the incident and solve it quickly. To this end, we propose a software solution, Cloud Insight, to provide some visibility through the middle virtualization layer for both cloud users and providers to address their problems quickly. Cloud Insight automatically tracks each VM instance's configuration status and maintains their life-cycle configuration records in a configuration management database (CMDB). When a user reports a problem, our algorithms automatically analyze CMDB to probabilistically determine the root cause and invoke a recovery process by interacting with the cloud user. Experimental results over data from Amazon EC2 online support forum and NEC Labs' research cloud infrastructures demonstrate that our approach can effectively automate the problem troubleshooting process in cloud environments.
Ahsan Arefin, Guofei Jiang
SRDS2
2011 Understanding data center network architectures in virtualized environments: A view from multi-tier applications
Yueping Zhang, Ao-Jan Su, Guofei Jiang
Comput. Networks3
2011 Experience Transfer for the Configuration Tuning in Large-Scale Computing Systems
abstract
This paper proposes a new strategy, the experience transfer, to facilitate the management of large-scale computing systems. It deals with the utilization of management experiences in one system (or previous systems) to benefit the same management task in other systems (or current systems). We use the system configuration tuning as a case application to demonstrate all procedures involved in the experience transfer including the experience representation, experience extraction, and experience embedding. The dependencies between system configuration parameters are treated as transferable experiences in the configuration tuning for two reasons: 1) because such knowledge is helpful to the efficiency of the optimal configuration search, and 2) because the parameter dependencies are typically unchanged between two similar systems. We use the Bayesian network to model configuration dependencies and present a configuration tuning algorithm based on the Bayesian network construction and sampling. As a result, after the configuration tuning is completed in the original system, we can obtain a Bayesian network as the by-product which records the dependencies between system configuration parameters. Such a network is then embedded into the tuning process in other similar systems as transferred experiences to improve the configuration search efficiency. Experimental results in a web-based system show that with the help of transferred experiences, the configuration tuning process can be significantly accelerated.
Guofei Jiang
IEEE Trans. Knowl. Data Eng.3
2011 Combined Power and Performance Management of Virtualized Computing Environments Serving Session-Based Workloads
abstract
This paper develops an online resource provisioning framework for combined power and performance management in a virtualized computing environment serving session-based workloads. We pose this management problem as one of sequential optimization under uncertainty and solve it using limited lookahead control (LLC), a form of model-predictive control. The approach accounts for the switching costs incurred when provisioning virtual machines and explicitly encodes the risk of provisioning resources in an uncertain and dynamic operating environment. We experimentally validate the control framework on a server cluster supporting three online services. When managed using LLC, our cluster setup saves, on average, 41% in power-consumption costs over a twenty-four hour period when compared to a system operating without dynamic control. Finally, we use trace-based simulations to analyze LLC performance on server clusters larger than our testbed and show how concepts from approximation theory can be used to further reduce the computational burden of controlling large systems.
Dara Kusic, Nagarajan Kandasamy, Guofei Jiang
IEEE Trans. Netw. Serv. Manag.3
2010 Optimal Probing for Unicast Network Delay Tomography
abstract
Network tomography has been proposed to ascertain internal network performances from end-to-end measurements. In this work, we present priority probing, an optimal probing scheme for unicast network delay tomography that is proven to provide the most accurate estimation. We first demonstrate that the Fisher information matrix in unicast network delay tomography can be decomposed into an additive form where each term can be obtained numerically. This establishes the space over which we can design the optimal probing scheme. Then, we formulate the optimal probing problem into a semi-definite programming (SDP) problem. High computation complexity constrains the SDP solution to only small scale scenarios. In response, we propose a greedy algorithm that approximates the optimal solution. Evaluations through simulation demonstrate that priority probing effectively increases estimation accuracy with a fixed number of probes.
Yu Gu 0004, Guofei Jiang, Vishal K. Singh, Yueping Zhang
INFOCOM2
2010 Evaluating the impact of data center network architectures on application performance in virtualized environments
abstract
In recent years, data center network (DCN) architectures (e.g., DCell, FiConn, BCube, FatTree, and VL2) received a surge of interest from both the industry and academia. However, none of existing efforts provide an in-depth understanding of the impact of these architectures on application performance in practical multi-tier systems under realistic workload. Moreover, it is also unclear how these architectures are affected in virtualized environments. In this paper, we fill this void by conducting an experimental evaluation of FiConn and FatTree, each respectively as a representative of hierarchical and flat architectures, in a three-tier transaction system using virtual machine (VM) based implementation. We observe several fundamental characteristics that are embedded in both classes of network topologies and cast a new light on the implication of virtualization in DCN architectures. Issues observed in this paper are generic and should be properly addressed by any DCN architectures before being considered for actual deployment, especially in mission-critical real-time transaction systems.
Yueping Zhang, Ao-Jan Su, Guofei Jiang
IWQoS3
2010 Supporting System-wide Similarity Queries for networked system management
abstract
Today's networked systems are extensively instrumented for collecting a wealth of monitoring data. In this paper, we propose a framework called System-wide Similarity Query (S2Q) to support a new type of similarity queries on monitoring data for managing complex networked systems. The similarity queries are defined on a novel data model that captures system states, and the implementation includes a streaming algorithm for online state-modeling computation and a companion graph-based indexing technique for fast retrieval of historical system states. S2Q simplifies many systems management tasks through a simple and intuitive query interface available to operators, and two applications are evaluated in the paper: (i) fast diagnosis of repeated failures in enterprise IT systems, and (ii) automated application traffic profiling on computer networks. For the first application, the diagnosis accuracy can reach 95% on a multi-tier web service testbed. For the second application, major network applications were automatically identified in the traffic logs from a large campus wireless network.
Songyun Duan, Hui Zhang 0002, Guofei Jiang, Xiaoqiao Meng
NOMS3
2010 Invariants Based Failure Diagnosis in Distributed Computing Systems
abstract
This paper presents an instance based approach to diagnosing failures in computing systems. Owing to the fact that a large portion of occurred failures are repeated ones, our method takes advantage of past experiences by storing historical failures in a database and retrieving similar instances in the occurrence of failure. We extract the system `invariants' by modeling consistent dependencies between system attributes during the operation, and construct a network graph based on the learned invariants. When a failure happens, the status of invariants network, i.e., whether each invariant link is broken or not, provides a view of failure characteristics. We use a high dimensional binary vector to store those failure evidences, and develop a novel algorithm to efficiently retrieve failure signatures from the database. Experimental results in a web based system have demonstrated the effectiveness of our method in diagnosing the injected failures.
Guofei Jiang, Kenji Yoshihira, Akhilesh Saxena
SRDS2
2010 A Cooperative Sampling Approach to Discovering Optimal Configurations in Large Scale Computing Systems
abstract
With the growing scale of current computing systems, traditional configuration tuning methods become less effective because they usually assume a small number of parameters in the system. In order to handle the scalability issue of configuration tuning, this paper proposes a cooperative optimization framework, which mimics the behavior of team playing to discover the optimal configuration setting in computing systems. We follow a `best of the best' rule to decompose the tuning task into a number of small subtasks with manageable size and complexity. While each decomposed module is responsible for the optimization of its own configuration parameters, all the modules share the performance evaluations of new samples as common feedbacks to enhance their optimization objectives. As a result, the qualities of generated samples become improved during the search, and the cooperative sampling will eventually discover the optimal configurations in the system. Experimental results demonstrate that our proposed cooperative optimization can identify better solutions within limited time periods compared with other state of the art configuration search methods. Such advantage becomes more significant when the number of configuration parameters increases.
Guofei Jiang, Hui Zhang 0002, Kenji Yoshihira
SRDS2
2010 Understanding Internet Video sharing site workload: A view from data center design
Xiaozhu Kang, Hui Zhang 0002, Guofei Jiang, Xiaoqiao Meng, Kenji Yoshihira
J. Vis. Commun. Image Represent.3
2009 Inputs of Coma: Static Detection of Denial-of-Service Vulnerabilities
abstract
As networked systems grow in complexity, they are increasingly vulnerable to denial-of-service (DoS) attacks involving resource exhaustion. A single malicious "input of coma" can trigger high-complexity behavior such as deep recursion in a carelessly implemented server, exhausting CPU time or stack space and making the server unavailable to legitimate clients. These DoS attacks exploit the semantics of the target application, are rarely associated with network traffic anomalies, and are thus extremely difficult to detect using conventional methods.We present SAFER, a static analysis tool for identifying potential DoS vulnerabilities and the root causes of resource-exhaustion attacks before the software is deployed. Our tool combines taint analysis with control dependency analysis to detect high-complexity control structures whose execution can be triggered by untrusted network inputs.When evaluated on real-world networked applications, SAFER discovered previously unknown DoS vulnerabilities in the Expat XML parser and the SQLite library, as well as a new attack on a previously patched version of the wu-ftpd server. This demonstrates the importance of understanding and repairing the root causes of DoS vulnerabilities rather than simply blocking known malicious inputs.
Richard M. Chang, Guofei Jiang, Franjo Ivancic, Sriram Sankaranarayanan 0001, Vitaly Shmatikov
CSF2
2009 Modeling Probabilistic Measurement Correlations for Problem Determination in Large-Scale Distributed Systems
abstract
With the growing complexity in computer systems, it has been a real challenge to detect and diagnose problems in today's large-scale distributed systems. Usually, the correlations between measurements collected across the distributed system contain rich information about the system behaviors, and thus a reasonable model to describe such correlations is crucially important in detecting and locating system problems. In this paper, we propose a transition probability model based on markov properties to characterize pair-wise measurement correlations. The proposed method can discover both the spatial (across system measurements) and temporal (across observation time) correlations, and thus such a model can successfully represent the system normal profiles. Problem determination and localization under this framework is fast and convenient. The framework is general enough to discover any types of correlations (e.g. linear or non-linear). Also, model updating, system problem detection and diagnosis can be conducted effectively and efficiently. Experimental results show that, the proposed method can detect the anomalous events and locate the problematic sources by analyzing the real monitoring data collected from three companies' infrastructures.
Jing Gao 0004, Guofei Jiang, Jiawei Han 0001
ICDCS2
2009 QoEScope: Adaptive IP service management for heterogeneous enterprise networks
abstract
In the recent years, a progressively growing number of computing and communication services have undertaken the migration from their conventional media to the new unified platform, IP networks. As a consequence, business success of service providers becomes largely determined by the effectiveness of their service management schemes, which require rapid identification of problems and resolution of network-related anomalies. However, this is a non-trivial task in heterogeneous enterprise networks due to service providers' invisibility of the health and performance of the underlying carrier network. In addition, the gap between quality of service (QoS) measurements reflecting network performance and quality of experience (QoE) metrics indicating user-perceived service quality further makes effective service management more challenging. In this paper, we present a unified service management system called QoEScope, which combines scalable end-to-end probing, accurate topology inference in the presence of implicit routers, adaptive bridging between QoS measurement and QoE metrics, and intelligent root cause analysis. Extensive testbed emulations and Internet experiments demonstrate that QoEScope is a highly practical and effective IP service management solution for heterogeneous enterprise networks.
Yueping Zhang, Vishal K. Singh, Yu Gu 0004, Guofei Jiang, Yu Ru
IWQoS4
2008 Enabling Information Confidentiality in Publish/Subscribe Overlay Services
abstract
"Alice has a piece of valuable information which she is willing to sell to anyone who is interested in; she is too busy and wants to ask Bob, a professional broker, to sell that information for her; but Alice is in a dilemma where she cannot trust Bob with that information but Bob cannot help her find her customers without knowing that information." In this paper, we propose a security mechanism called information foiling to address new confidentiality problems arising in pub/sub overlay services [1]. Information foiling extends Rivest's "Chaffing and Winnowing" [2], and its basic idea is to carefully generate a set of fake messages to hide an authentic message. Information foiling requires no modification inside the broker network so that the routing/filtering capabilities of broker nodes remains intact. We formally present the information foiling mechanism in the context of publish/subscribe overlay services, and discuss its applicability in other Internet applications. For publish/subscribe applications, we propose a suite of optimal schemes for fake message generation in different scenarios. Real-world data are used in our evaluation to demonstrate the effectiveness of the proposed schemes.
Hui Zhang 0002, Abhishek B. Sharma, Guofei Jiang, Xiaoqiao Meng, Kenji Yoshihira
ICC4
2008 Autotuning Configurations in Distributed Systems for Performance Improvements Using Evolutionary Strategies
abstract
Distributed systems usually have many configurable parameters such as those included in common configuration files. Performance of distributed systems is partially dependent on these system configurations. While operators may choose default settings or manually tune parameters based on their experience and intuition, the resulted settings may not be the optimal one for specific services running on the distributed system. In this paper, we formulate the problem of autotuning configurations as a black-box optimization problem. This problem becomes quite challenging since the joint parameter search space is huge and also no explicit relationship between performance and configurations exists. We propose to use a well known evolutionary algorithm called covariance matrix adaptation (CMA) to automatically tune system parameters. We compare CMA algorithm to another existing techniques called smart hill climbing (SHC) and demonstrate that CMA algorithm outperforms SHC algorithm both on synthetic data and in a real system.
Anooshiravan Saboori, Guofei Jiang
ICDCS2
2008 Exploiting Local and Global Invariants for the Management of Large Scale Information Systems
abstract
This paper presents a data oriented approach to modeling the complex computing systems, in which an ensemble of correlation models are discovered to represent the system status. If the discovered correlations can continually hold under different user scenarios and workloads, they are regarded as invariants of the information system. In our previous work, we have developed an algorithm to automatically search the invariants between any pair of system attributes, which we call local invariants. However that method is unable to deal with the high order dependency models due to the combinatorial explosion of search space. In this paper we use Bayesian regression technique to discover those high order correlation models, called global invariants. We treat each attribute as a response variable in turn and express its dependency with the other attributes in a regression model. By adding the prior constraint of Laplacian distribution to the regression coefficients, we can find the solution in which only the correlated attributes with respect to the response have nonzero regression coefficients. After that we further consider the temporal dependencies of those extracted attributes by incorporating their past observations. We also provide a confidence metric and a validation procedure to measure the reliability of learned models. If the model does not break down in the validation, it is regarded as a true invariant of the system. Experimental results on a real wireless networking system show that the discovered invariants can be used to effectively detect system failures as well as provide valuable information about the failure source.
Haibin Cheng, Guofei Jiang, Kenji Yoshihira
ICDM3
2008 Measurement, Modeling, and Analysis of Internet Video Sharing Site Workload: A Case Study
abstract
In this paper we measured and analyzed the workload on Yahoo! Video, the 2nd largest U.S. video sharing site, to understand its nature and the impact on online video data center design. We discovered interesting statistical properties on both static and temporal dimensions of the workload; they include file duration and popularity distributions, arrival rate dynamics and predictability, and workload stationarity and burstiness. Complemented with queueing-theoretic techniques, we extended our understanding on the measurement data with a virtual data center design assuming the same workload as measured, which reveals results regarding the impact of workload arrival distribution, service level agreements (SLAs) and workload scheduling schemes on the design and operations of such large-scale video distribution systems.
Xiaozhu Kang, Hui Zhang 0002, Guofei Jiang, Xiaoqiao Meng, Kenji Yoshihira
ICWS3
2008 Automatic Profiling of Network Event Sequences: Algorithm and Applications
abstract
The behavior of network entities, such as flows, sessions, hosts, and users, can often be described by communication event sequences in the time domain. For the purpose of many network measurement and monitoring tasks, it is desirable to have an accurate yet information-compact profiling of the behavior of massive event sequences. This paper proposes a new method to achieve this goal. On a given set of event sequences, the proposed method automatically learns a mixture model which fully captures the sequence behavior including both event pattern and duration between events. The learned mixture model is information-compact as it classifies sequences into a set of behavior templates, each of which is described by a Markov Chain. The model parameters are estimated in an iterative procedure which is developed from the Expectation Maximization algorithm. Two network management applications are proposed based on the method: a visualization tool for network administrators to conduct exploratory traffic analysis, and an efficient anomaly detection mechanism. In the evaluation, we validate the method accuracy as well as the usefulness of the two applications by using three networking datasets with different types: TCP packet traces, VoIP calls, and syslog traces in wireless networks.
Xiaoqiao Meng, Guofei Jiang, Hui Zhang 0002, Kenji Yoshihira
INFOCOM2
2008 Correlating real-time monitoring data for mobile network management
abstract
With a proliferation of new mobile data services, the complexity of wireless mobile networks is rapidly growing. While large amount of operational monitoring data such as performance measurement statistics is available, it is a great challenge to correlate such data effectively for real time performance analysis. Meantime, the dynamics of mobile applications and environments introduce another dimension of complexity for us to track the evolving system status. In this paper, we analyze the spatial and temporal correlations of Key Performance Indicators (KPIs) to track and interpret the operational status of wide-area cellular systems. We first correlate large number of raw measurements into limited number of KPIs. Further we exploit spatial and temporal correlations of these KPIs for cellular network management. We use large volume of field data collected from real cellular systems in our analysis. Experimental results demonstrate that it is promising to build a real-time data management and support system by effectively correlating KPIs.
Nanyan Jiang, Guofei Jiang, Kenji Yoshihira
WOWMOM2
2008 Understanding internet video sharing site workload: a view from data center design
abstract
In this paper we measured and analyzed the workload on Yahoo! Video, the 2nd largest U.S. video sharing site, to understand its nature and the impact on online video data center design. We discovered interesting statistical properties on both static and temporal dimensions of the workload including file duration and popularity distributions, arrival rate dynamics and predictability, and workload stationarity and burstiness. Complemented with queueing-theoretic techniques, we further extended our understanding on the measurement data with a virtual design on the workload and capacity management components of a data center assuming the same workload as measured, which reveals key results regarding the impact of Service Level Agreements (SLAs) and workload scheduling schemes on the design and operations of such large-scale video distribution systems.
Xiaozhu Kang, Hui Zhang 0002, Guofei Jiang, Xiaoqiao Meng, Kenji Yoshihira
WWW3
2008 Monitoring High-Dimensional Data for Failure Detection and Localization in Large-Scale Computing Systems
abstract
It is a major challenge to process high-dimensional measurements for failure detection and localization in large-scale computing systems. However, it is observed that in information systems, those measurements are usually located in a low-dimensional structure that is embedded in the high-dimensional space. From this perspective, a novel approach is proposed to model the geometry of underlying data generation and detect anomalies based on that model. We consider both linear and nonlinear data generation models. Two statistics, that is, the Hotelling T2and the squared prediction error (SPE), are used to reflect data variations within and outside the model. We track the probabilistic density of extracted statistics to monitor the system's health. After a failure has been detected, a localization process is also proposed to find the most suspicious attributes related to the failure. Experimental results on both synthetic data and a real e-commerce application demonstrate the effectiveness of our approach in detecting and localizing failures in computing systems.
Guofei Jiang, Kenji Yoshihira
IEEE Trans. Knowl. Data Eng.2
2008 The theory of trackability with applications to sensor networks
abstract
In this article, we formalize the concept of tracking in a sensor network and develop a quantitative theory of trackability of weak models that investigates the rate of growth of the number of consistent tracks given a temporal sequence of observations made by the sensor network. The phenomenon being tracked is modelled by a nondeterministic finite automaton (a weak model) and the sensor network is modelled by an observer capable of detecting events related, typically ambiguously, to the states of the underlying automaton. Formally, an input string of symbols (the sensor network observations) that is presented to a nondeterministic finite automaton, M , (the weak model) determines a set of state sequences (the tracks or hypotheses) that are capable of generating the input string. We study the growth of the size of this candidate set of tracks as a function of the length of the input string. One key result is that for a given automaton and sensor coverage, the worst-case rate of growth is either polynomial or exponential in the number of observations, indicating a kind of phase transition in tracking accuracy. These results have applications to various tracking problems of recent interest involving tracking phenomena using noisy observations of hidden states such as: sensor networks, computer network security, autonomic computing and dynamic social network analysis.
Valentino Crespi, George Cybenko, Guofei Jiang
ACM Trans. Sens. Networks3
2008 Adaptive Sensor Placement and Boundary Estimation for Monitoring Mass Objects
abstract
Sensor networks are widely used in monitoring and tracking a large number of objects. Without prior knowledge on the dynamics of object distribution, their density estimation could be learned in an adaptive manner to support effective sensor placement. After sensors observe the "current" locations of objects, the estimates of object distribution are updated with these new observations through a recursive distributed expectation-maximization algorithm. Based on the real-time estimates of object distribution, an adaptive sensor placement algorithm could be designed to achieve stable and high accuracy in tracking mass objects. This paper constructs a Gaussian mixture model to characterize the mixture distribution of object locations and proposes a novel methodology to adaptively update sensor placement. Our simulation results demonstrate the effectiveness of the proposed algorithm for adaptive sensor placement and boundary estimation of mass objects.
MengChu Zhou, Guofei Jiang
IEEE Trans. Syst. Man Cybern. Part B3
2008 Approximation Modeling for the Online Performance Management of Distributed Computing Systems
abstract
A promising method of automating management tasks in computing systems is to formulate them as control or optimization problems in terms of performance metrics. For an online optimization scheme to be of practical value in a distributed setting, however, it must successfully tackle the curses of dimensionality and modeling. This paper develops a hierarchical control framework to solve performance management problems in distributed computing systems operating in a data center. Concepts from approximation theory are used to reduce the computational burden of controlling such large-scale systems. The relevant approximations are made in the construction of the dynamical models to predict system behavior and in the solution of the associated control equations. Using a dynamic resource-provisioning problem as a case study, we show that a computing system managed by the proposed control framework with approximation models realizes profit gains that are, in the best case, within 1% of a controller using an explicit model of the system.
Dara Kusic, Nagarajan Kandasamy, Guofei Jiang
IEEE Trans. Syst. Man Cybern. Part B3
2007 Real-time Application Monitoring and Diagnosis for Service Hosting Platforms of Black Boxes
abstract
Service hosting platforms typically run a large number of third-party applications that are composed of multiple communicating components distributed on a dynamic set of hosting servers. Understanding the real-time behaviors of these applications and the intricate interactions/dependency relationships among these application components is very important to service management tasks such as load balancing, capacity planning, performance debugging and fault diagnosis. In this paper, we present the scalable real-time application monitoring and diagnosis (SRAMD) tool, for applications consisting of "black box" components: software without source code available, and usually without desired logging instrumentation. SRAMD runs at application layer and requires no modification to existing applications, middleware, or messages. For each application component collocated at its hosting server, a SRAMD monitor traces the component's packet-level traffic unobtrusively, summarizes its local resource utilization and performance (e.g. response time) online, performs interactive queries (e.g. per- request resource utilization) to locate possible bottlenecks on- demand, and discovers inter-component dependence relationships statistically. The SRAMD controller simply aggregates reports from distributed monitors to construct real-time application topologies with rich runtime information. We have developed mechanisms to decentralize the computation overhead and minimize the communication cost in the monitoring and diagnosis process, and two schemes to discover application component dependency relationships in different scenarios. The SRAMD tool offers an alternative to server logs and message-level traces for service monitoring and performance diagnosis.
Huadong Liu, Hui Zhang 0002, Rauf Izmailov, Guofei Jiang, Xiaoqiao Meng
Integrated Network Management4
2007 State space exploration using feedback constraint generation and Monte-Carlo sampling
abstract
The systematic exploration of the space of all the behaviours of a software system forms the basis of numerous approaches to verification. However, existing approaches face many challenges with scalability and precision. We propose a framework for validating programs based on statistical sampling of inputs guided by statically generated constraints, that steer the simulations towards more "desirable" traces.
Sriram Sankaranarayanan 0001, Richard M. Chang, Guofei Jiang, Franjo Ivancic
ESEC/SIGSOFT FSE3
2007 Failure Detection in Large-Scale Internet Services by Principal Subspace Mapping
abstract
Fast and accurate failure detection is becoming essential in managing large scale Internet services. This paper proposes a novel detection approach based on the subspace mapping between system inputs and internal measurements. By exploring these contextual dependencies, our detector can initiate repair actions accurately, increasing the availability of system. While a classical statistical method, the canonical correlation analysis (CCA), is presented in the paper to achieve subspace mapping, we also propose a more advanced technique, the principal canonical correlation analysis (PCCA), to improve the performance of CCA based detector. PCCA extracts a principal subspace from internal measurements that is not only highly correlated with the inputs, but also a significant representative of original measurements. Experimental results on a J2EE based web application demonstrate that such property of PCCA is especially beneficial to failure detection tasks.
Guofei Jiang, Kenji Yoshihira
IEEE Trans. Knowl. Data Eng.2
2007 Efficient and Scalable Algorithms for Inferring Likely Invariants in Distributed Systems
abstract
Distributed systems generate a large amount of monitoring data such as log files to track their operational status. However, it is hard to correlate such monitoring data effectively across distributed systems and along observation time for system management. In previous work, we proposed a concept named flow intensity to measure the intensity with which internal monitoring data reacts to the volume of user requests. We calculated flow intensity measurements from monitoring data and proposed an algorithm to automatically search constant relationships between flow intensities measured at various points across distributed systems. If such relationships hold all the time, we regard them as invariants of the underlying systems. Invariants can be used to characterize complex systems and support various system management tasks. However, the computational complexity of the previous invariant search algorithm is high so that it may not scale well in large systems with thousands of measurements. In this paper, we propose two efficient but approximate algorithms for inferring invariants in large-scale systems. The computational complexity of new randomized algorithms is significantly reduced, and experimental results from a real system are also included to demonstrate the accuracy and efficiency of our new algorithms.
Guofei Jiang, Kenji Yoshihira
IEEE Trans. Knowl. Data Eng.1
2007 Online Tracking of Component Interactions for Failure Detection and Localization in Distributed Systems
abstract
This paper proposes a novel failure-detection approach that can handle high-dimensional observation and frequent system changes. We extract two statistics from the subspace decomposition of observations, and use the mixture of Gaussians to model their probability density. Instead of monitoring the original data, the density model of extracted statistics is adaptively updated and examined regularly to detect failures. We also present a localization method to identify the faulty components once the failure happens. Applying our technique to monitor the component interactions in an e-commerce application shows satisfactory results in detecting a variety of injected failures.
Guofei Jiang, Cristian Ungureanu, Kenji Yoshihira
IEEE Trans. Syst. Man Cybern. Part C2
2007 Multiresolution Abnormal Trace Detection Using Varied-Length n-Grams and Automata
abstract
Detection and diagnosis of faults in a large-scale distributed system is a formidable task. Interest in monitoring and using traces of user requests for fault detection has been on the rise recently. In this paper we propose novel fault detection methods based on abnormal trace detection. One essential problem is how to represent the large amount of training trace data compactly as an oracle. Our key contribution is the novel use of varied-length n-grams and automata to characterize normal traces. A new trace is compared against the learned automata to determine whether it is abnormal. We develop algorithms to automatically extract n-grams and construct multiresolution automata from training data. Further, both deterministic and multihypothesis algorithms are proposed for detection. We inspect the trace constraints of real application software and verify the existence of long n-grams. Our approach is tested in a real system with injected faults and achieves good results in experiments
Guofei Jiang, Cristian Ungureanu, Kenji Yoshihira
IEEE Trans. Syst. Man Cybern. Part C1
2006 Tracking Probabilistic Correlation of Monitoring Data for Fault Detection in Complex Systems
abstract
Due to their growing complexity, it becomes extremely difficult to detect and isolate faults in complex systems. While large amount of monitoring data can be collected from such systems for fault analysis, one challenge is how to correlate the data effectively across distributed systems and observation time. Much of the internal monitoring data reacts to the volume of user requests accordingly when user requests flow through distributed systems. In this paper, we use Gaussian mixture models to characterize probabilistic correlation between flow-intensities measured at multiple points. A novel algorithm derived from Expectation-Maximization (EM) algorithm is proposed to learn the "likely" boundary of normal data relationship, which is further used as an oracle in anomaly detection. Our recursive algorithm can adaptively estimate the boundary of dynamic data relationship and detect faults in real time. Our approach is tested in a real system with injected faults and the results demonstrate its feasibility.
Guofei Jiang, Kenji Yoshihira
DSN2
2006 Semantic Interoperability and Information Fluidity
abstract
Ontologies are developed to describe data semantics on the Semantic Web. Given the distributed nature and scale of the Semantic Web, a large number of ontologies with different terminologies and structures will be created to describe the same concepts and domains. Without semantic mapping, information fluidity within the Web could be blocked at the boundaries of these ontologies. Therefore, ontology mapping is needed to translate datasets represented by disparate ontologies. We believe that over time communities will incrementally build an ontology mapping between select ontologies based on their own communication interests. How will these interest-driven mapping activities eventually change semantic interoperability and information fluidity across the Web? This paper proposes metrics to quantify information fluidity and builds an analytical model with "small-world" graph theory to analyze the growth of the Semantic Web. Further with this model, we analyze how information fluidity can evolve by "market-driven" semantic mapping activities occurring across the Web. Our results can be useful in evaluating mapping efforts needed for large-scale heterogeneous information systems. One conclusion, based on this model, is that the development of decentralized ontology mappings can lead to significant information fluidity within the Semantic Web.
Guofei Jiang, George Cybenko, James A. Hendler
Int. J. Cooperative Inf. Syst.1
2006 Modeling and Tracking of Transaction Flow Dynamics for Fault Detection in Complex Systems
abstract
With the prevalence of Internet services and the increase of their complexity, there is a growing need to improve their operational reliability and availability. While a large amount of monitoring data can be collected from systems for fault analysis, it is hard to correlate this data effectively across distributed systems and observation time. In this paper, we analyze the mass characteristics of user requests and propose a novel approach to model and track transaction flow dynamics for fault detection in complex information systems. We measure the flow intensity at multiple checkpoints inside the system and apply system identification methods to model transaction flow dynamics between these measurements. With the learned analytical models, a model-based fault detection and isolation method is applied to track the flow dynamics in real time for fault detection. We also propose an algorithm to automatically search and validate the dynamic relationship between randomly selected monitoring points. Our algorithm enables systems to have self-cognition capability for system management. Our approach is tested in a real system with a list of injected faults. Experimental results demonstrate the effectiveness of our approach and algorithms
Guofei Jiang, Kenji Yoshihira
IEEE Trans. Dependable Secur. Comput.1
2005 Failure detection and localization in component based systems by online tracking
abstract
The increasing complexity of today's systems makes fast and accurate failure detection essential for their use in mission-critical applications. Various monitoring methods provide a large amount of data about system's behavior. Analyzing this data with advanced statistical methods holds the promise of not only detecting the errors faster, but also detecting errors which are difficult to catch with current monitoring tools. Two challenges to building such detection tools are: the high dimensionality of observation data, which makes the models expensive to apply, and frequent system changes, which make the models expensive to update. In this paper, we present algorithms to reduce the dimensionality of data in a way that makes it easy to adapt to system changes. We decompose the observation data into signal and noise subspaces. Two statistics, the Hotelling T2 score and squared prediction error (SPE) are calculated to represent the data characteristics in signal and noise subspaces respectively. Instead of tracking the original data, we use a sequentially discounting expectation maximization (SDEM) algorithm to learn the distribution of the two extracted statistics. A failure event can then be detected based on the abnormal change of the distribution. Applying our technique to component interaction data in a simple e-commerce application shows better accuracy than building independent profiles for each component. Additionally, experiments on synthetic data show that the detection accuracy is high even for changing systems.
Guofei Jiang, Cristian Ungureanu, Kenji Yoshihira
KDD2
2004 Functional Validation in Grid Computing
Guofei Jiang, George Cybenko
Auton. Agents Multi Agent Syst.1
2003 Semantic depth and markup complexity
abstract
In order to achieve interoperability among heterogeneous systems, markup languages such as XML and DAML are being used to describe distributed systems and data. The ability to successfully interoperate based on semantic markup depends on the ability to create, use and manage shared ontologies of concepts and their interrelationships. Specifically, communicating systems in a networked environment have to achieve a certain level of semantic agreement for them to understand and process exchanged data. A challenging question is how deep the semantic agreement has to be in order to satisfy the communication needs in an environment. Additionally, what is the markup complexity resulting from pursuing that depth of semantic agreement? This paper introduces the concept of semantic depth and markup complexity and proposes models to measure the markup complexity. Furthermore, it is shown that markup complexity can be reduced by employing hierarchical ontologies after partitioning the domain into smaller sub-domains.
Guofei Jiang, George Cybenko, James A. Hendler
SMC1
2002 Performance Analysis of Mobile Agents for Filtering Data Streams on Wireless Networks
David Kotz, George Cybenko, Robert S. Gray, Guofei Jiang, Ronald A. Peterson, Martin O. Hofmann, Daria A. Chacón, Kenneth R. Whitebread, James A. Hendler
Mob. Networks Appl.4
2000 Performance analysis of mobile agents for filtering data streams on wireless networks
abstract
Wireless networks are an ideal environment for mobile agents, because their mobility allows them to move across an unreliable link to reside on a wired host, next to or closer to the resources they need to use. Furthermore, client-specific data transformations can be moved across the wireless link, and run on a wired gateway server, with the goal of reducing bandwidth demands. In this paper we examine the tradeoffs faced when deciding whether to use mobile agents to support a data-filtering application, in which numerous wireless clients filter information from a large data stream arriving across the wired network. We develop an analytical model and use parameters from our own experiments to explore the model's implications.
David Kotz, Guofei Jiang, Robert S. Gray, George Cybenko, Ronald A. Peterson
MSWiM2