Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Wei Xu 0012

dblp:32/1213-12 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
0since 2021 · last 2010
0000-0003-3708-6816ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
3 papers
Software maintenance and evolution · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 100%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software maintenance and evolution
log analysis
0.332010
Detecting Large-Scale System Problems by Mining Console Logs · ICML 2010
Detecting large-scale system problems by mining console logs · SOSP 2009
Online System Problem Detection by Mining Patterns of Console Logs · ICDM 2009
Software maintenance and evolution › log analysis
log-based anomaly detection
0.112009
Online System Problem Detection by Mining Patterns of Console Logs · ICDM 2009
Distributed systems
fault tolerance
0.112009
Detecting large-scale system problems by mining console logs · SOSP 2009
Data mining › pattern mining
log mining
0.012009
Detecting large-scale system problems by mining console logs · SOSP 2009
Data mining
pattern mining
0.012009
Detecting large-scale system problems by mining console logs · SOSP 2009

Methods — techniques the papers use, named apart from their topics

source code analysis · 0.3machine learning · 0.3information retrieval · 0.3decision tree · 0.3log mining · 0.2principal component analysis · 0.1frequent pattern mining · 0.1distribution estimation · 0.1
YearPublicationVenuePosition
2010 Detecting Large-Scale System Problems by Mining Console Logs
Wei Xu 0012, Ling Huang 0001, Armando Fox, David A. Patterson 0001, Michael I. Jordan
ICML1
2009 Online System Problem Detection by Mining Patterns of Console Logs
abstract
We describe a novel application of using data mining and statistical learning methods to automatically monitor and detect abnormal execution traces from console logs in an online setting. Different from existing solutions, we use a two stage detection system. The first stage uses frequent pattern mining and distribution estimation techniques to capture the dominant patterns (both frequent sequences and time duration). The second stage use principal component analysis based anomaly detection technique to identify actual problems. Using real system data from a 203-node Hadoop cluster, we show that we can not only achieve highly accurate and fast problem detection, but also help operators better understand execution patterns in their system.
Wei Xu 0012, Ling Huang 0001, Armando Fox, David A. Patterson 0001, Michael I. Jordan
ICDM1
2009 Detecting large-scale system problems by mining console logs
abstract
Surprisingly, console logs rarely help operators detect problems in large-scale datacenter services, for they often consist of the voluminous intermixing of messages from many software components written by independent developers. We propose a general methodology to mine this rich source of information to automatically detect system runtime problems. We first parse console logs by combining source code analysis with information retrieval to create composite features. We then analyze these features using machine learning to detect operational problems. We show that our method enables analyses that are impossible with previous methods because of its superior ability to create sophisticated features. We also show how to distill the results of our analysis to an operator-friendly one-page decision tree showing the critical messages associated with the detected problems. We validate our approach using the Darkstar online game server and the Hadoop File System, where we detect numerous real problems with high accuracy and few false positives. In the Hadoop case, we are able to analyze 24 million lines of console logs in 3 minutes. Our methodology works on textual console logs of any size and requires no changes to the service software, no human input, and no knowledge of the software's internals.
Wei Xu 0012, Ling Huang 0001, Armando Fox, David A. Patterson 0001, Michael I. Jordan
SOSP1