Li Wei 0001

dblp:w/LiWei · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
0since 2021 · last 2009
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 13 · 6 first-authorArtificial intelligence and machine learning · 6 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
7 papers
Indexing and storage engines · 36% Data mining · 32% Information retrieval · 19%
Computer graphics and multimedia
2 papers
Multimedia analysis and retrieval · 57% Visualization and visual analytics · 43%

Topics — the 13 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Indexing and storage engines › feature-based indexing
shape indexing
0.222009
Supporting exact indexing of arbitrarily rotated shapes and periodic time series under Euclidean and warping distance measures · VLDB J. 2009
LB_Keogh Supports Exact Indexing of Shapes under Rotation Invariance with Arbitrary Representations and Distance Measures · VLDB 2006
Indexing and storage engines › temporal indexing
time series indexing
0.222009
Supporting exact indexing of arbitrarily rotated shapes and periodic time series under Euclidean and warping distance measures · VLDB J. 2009
LB_Keogh Supports Exact Indexing of Shapes under Rotation Invariance with Arbitrary Representations and Distance Measures · VLDB 2006
Data mining › time series analysis
time series classification
0.122006
Semi-supervised time series classification · KDD 2006
Fast time series classification using numerosity reduction · ICML 2006
Data mining › time series analysis
time warping
0.112009
Supporting exact indexing of arbitrarily rotated shapes and periodic time series under Euclidean and warping distance measures · VLDB J. 2009
Multimedia analysis and retrieval › multimedia retrieval › content-based retrieval
shape retrieval
0.112008
Fast Best-Match Shape Searching in Rotation-Invariant Metric Spaces · IEEE Trans. Multim. 2008
Data mining
anomaly detection
0.112006
SAXually Explicit Images: Finding Unusual Shapes · ICDM 2006
Information retrieval › hashing › hashing for nearest neighbor search
locality-sensitive hashing
0.112006
SAXually Explicit Images: Finding Unusual Shapes · ICDM 2006
Information retrieval
similarity search
0.112006
SAXually Explicit Images: Finding Unusual Shapes · ICDM 2006
Spatial and temporal data management
time series data management
0.112006
LB_Keogh Supports Exact Indexing of Shapes under Rotation Invariance with Arbitrary Representations and Distance Measures · VLDB 2006
Visualization and visual analytics › visual encoding
icon-based visualization
0.112006
Intelligent Icons: Integrating Lite-Weight Data Mining and Visualization into GUI Operating Systems · ICDM 2006
Information retrieval
pattern matching
0.112005
Atomic Wedgie: Efficient Query Filtering for Streaming Times Series · ICDM 2005
Data stream processing
streaming time series
0.112005
Atomic Wedgie: Efficient Query Filtering for Streaming Times Series · ICDM 2005
Data mining
visualization
0.012006
Intelligent Icons: Integrating Lite-Weight Data Mining and Visualization into GUI Operating Systems · ICDM 2006

Methods — techniques the papers use, named apart from their topics

similarity arrangement · 0.2lightweight data mining · 0.2metric indexing · 0.1semi-supervised learning · 0.1self-training · 0.1locality-sensitive hashing · 0.1dynamic time warping · 0.1LB_Keogh · 0.1query filtering · 0.1lower bounding · 0.1
YearPublicationVenuePosition
2009 Supporting exact indexing of arbitrarily rotated shapes and periodic time series under Euclidean and warping distance measures
Eamonn J. Keogh, Li Wei 0001, Xiaopeng Xi, Michail Vlachos, Sang-Hee Lee 0003, Pavlos Protopapas
VLDB J.2
2008 Efficiently finding unusual shapes in large image databases
Li Wei 0001, Eamonn J. Keogh, Xiaopeng Xi, Melissa Yoder
Data Min. Knowl. Discov.1
2008 Fast Best-Match Shape Searching in Rotation-Invariant Metric Spaces
abstract
Object recognition and content-based image retrieval systems rely heavily on the accurate and efficient identification of 2-D shapes. Features such as color, texture, positioning etc., are insufficient to convey the information that could be obtained through shape analysis. A fundamental requirement in this analysis is that shape similarities are computed invariantly to basic geometric transformations, e.g., scaling, shifting, and most importantly, rotations. And while scale and shift invariance are easily achievable through a suitable shape representation, rotation invariance is much harder to deal with. In this work, we explore the metric properties of the rotation-invariant distance measures and propose an algorithm for fast similarity search in the shape space. The algorithm can be utilized in a number of important data mining tasks such as shape clustering and classification, or for discovering of motifs and discords in large image collections. The technique is demonstrated to introduce a dramatic speed-up over the current approaches, and is guaranteed to introduce no false dismissals.
Dragomir Yankov, Eamonn J. Keogh, Li Wei 0001, Xiaopeng Xi, Wendy L. Hodges
IEEE Trans. Multim.3
2007 Finding Motifs in a Database of Shapes
abstract
The problem of efficiently finding images that are similar to a target image has attracted much attention in the image processing community and is rightly considered an information retrieval task. However, the problem of finding structure and regularities in large image datasets is an area in which data mining is beginning to make fundamental contributions. In this work, we consider the new problem of discovering shape motifs, which are approximately repeated shapes within (or between) image collections. As we shall show, shape motifs can have applications in tasks as diverse as anthropology, law enforcement, and historical manuscript mining. Brute force discovery of shape motifs could be untenably slow, especially as many domains may require an expensive rotation invariant distance measure. We introduce an algorithm that is two to three orders of magnitude faster than brute force search, and demonstrate the utility of our approach with several real world datasets from diverse domains.
Xiaopeng Xi, Eamonn J. Keogh, Li Wei 0001, Agenor Mafra-Neto
SDM3
2007 Fast Best-Match Shape Searching in Rotation Invariant Metric Spaces
abstract
Object recognition and content-based image retrieval systems rely heavily on the accurate and efficient identification of shapes. A fundamental requirement in the shape analysis process is that shape similarities should be computed invariantly to basic geometric transformations, e.g. scaling, shifting, and most importantly, rotations. And while scale and shift invariance are easily achievable through a suitable shape representation, rotation invariance is much harder to deal with. In this work we explore the metric properties of the rotation invariant distance measures and propose an algorithm for fast similarity search in the shape space. The algorithm can be utilized in a number of important data mining tasks such as shape clustering and classification, or for discovering of motifs and discords in image collections. The technique is demonstrated to introduce a dramatic speed-up over the current approaches, and is guaranteed to introduce no false dismissals.
Dragomir Yankov, Eamonn J. Keogh, Li Wei 0001, Xiaopeng Xi, Wendy L. Hodges
SDM3
2007 Compression-based data mining of sequential data
Eamonn J. Keogh, Stefano Lonardi, Chotirat (Ann) Ratanamahatana, Li Wei 0001, Sang-Hee Lee 0003, John C. Handley
Data Min. Knowl. Discov.4
2007 Experiencing SAX: a novel symbolic representation of time series
Jessica Lin 0001, Eamonn J. Keogh, Li Wei 0001, Stefano Lonardi
Data Min. Knowl. Discov.3
2007 Efficient query filtering for streaming time series with applications to semisupervised learning of time series classifiers
Li Wei 0001, Eamonn J. Keogh, Helga Van Herle, Agenor Mafra-Neto, Russ Abbott
Knowl. Inf. Syst.1
2006 Intelligent Icons: Integrating Lite-Weight Data Mining and Visualization into GUI Operating Systems
abstract
The vast majority of visualization tools introduced so far are specialized pieces of software that run explicitly on a particular dataset at a particular time for a particular purpose. In this work we introduce a novel framework for allowing visualization to take place in the background of normal day-to-day operation of any GUI based operation system. Our system works by replacing the standard file icons with automatically created icons that reflect the contents of the files in a principled way. We call such icons Intelligent Icons. The utility of Intelligent Icons is further enhanced by arranging them in a way that reflects their similarity/differences. We demonstrate the utility of our approach on diverse applications.
Eamonn J. Keogh, Li Wei 0001, Xiaopeng Xi, Stefano Lonardi, Jin Shieh, Scott Sirowy
ICDM2
2006 SAXually Explicit Images: Finding Unusual Shapes
abstract
Over the past three decades, there has been a great deal of research on shape analysis, focusing mostly on shape indexing, clustering, and classification. In this work, we introduce the new problem of finding shape discords, the most unusual shapes in a collection. We motivate the problem by considering the utility of shape discords in diverse domains including zoology, anthropology, and medicine. While the brute force search algorithm has quadratic time complexity, we avoid this by using locality-sensitive hashing to estimate similarity between shapes which enables us to reorder the search more efficiently. An extensive experimental evaluation demonstrates that our approach can speed up computation by three to four orders of magnitude.
Li Wei 0001, Eamonn J. Keogh, Xiaopeng Xi
ICDM1
2006 Fast time series classification using numerosity reduction
abstract
Many algorithms have been proposed for the problem of time series classification. However, it is clear that one-nearest-neighbor with Dynamic Time Warping (DTW) distance is exceptionally difficult to beat. This approach has one weakness, however; it is computationally too demanding for many realtime applications. One way to mitigate this problem is to speed up the DTW calculations. Nonetheless, there is a limit to how much this can help. In this work, we propose an additional technique, numerosity reduction, to speed up one-nearest-neighbor DTW. While the idea of numerosity reduction for nearest-neighbor classifiers has a long history, we show here that we can leverage off an original observation about the relationship between dataset size and DTW constraints to produce an extremely compact dataset with little or no loss in accuracy. We test our ideas with a comprehensive set of experiments, and show that it can efficiently produce extremely fast accurate classifiers.
Xiaopeng Xi, Eamonn J. Keogh, Christian R. Shelton, Li Wei 0001, Chotirat (Ann) Ratanamahatana
ICML4
2006 Semi-supervised time series classification
abstract
The problem of time series classification has attracted great interest in the last decade. However current research assumes the existence of large amounts of labeled training data. In reality, such data may be very difficult or expensive to obtain. For example, it may require the time and expertise of cardiologists, space launch technicians, or other domain specialists. As in many other domains, there are often copious amounts of unlabeled data available. For example, the PhysioBank archive contains gigabytes of ECG data. In this work we propose a semi-supervised technique for building time series classifiers. While such algorithms are well known in text domains, we will show that special considerations must be made to make them both efficient and effective for the time series domain. We evaluate our work with a comprehensive set of experiments on diverse data sources including electrocardiograms, handwritten documents, and video datasets. The experimental results demonstrate that our approach requires only a handful of labeled examples to construct accurate classifiers.
Li Wei 0001, Eamonn J. Keogh
KDD1
2006 LB_Keogh Supports Exact Indexing of Shapes under Rotation Invariance with Arbitrary Representations and Distance Measures
Eamonn J. Keogh, Li Wei 0001, Xiaopeng Xi, Sang-Hee Lee 0003, Michail Vlachos
VLDB2
2005 A Practical Tool for Visualizing and Data Mining Medical Time Series
abstract
The increasing interest in time series data mining has had surprisingly little impact on real world medical applications. Practitioners who work with time series on a daily basis rarely take advantage of the wealth of tools that the data mining community has made available. In this work, we attempt to address this problem by introducing a parameter-light tool that allows users to efficiently navigate through large collections of time series. Our approach extracts features from a time series of arbitrary length and uses information about the relative frequency of these features to color a bitmap in a principled way. By visualizing the similarities and differences within a collection of bitmaps, a user can quickly discover clusters, anomalies, and other regularities within the data collection. We demonstrate the utility of our approach with a set of comprehensive experiments on real datasets from a variety of medical domains.
Li Wei 0001, Nitin Kumar 0002, Venkata Nishanth Lolla, Eamonn J. Keogh, Stefano Lonardi, Chotirat (Ann) Ratanamahatana, Helga Van Herle
CBMS1
2005 Atomic Wedgie: Efficient Query Filtering for Streaming Times Series
abstract
In many applications, it is desirable to monitor a streaming time series for predefined patterns. In domains as diverse as the monitoring of space telemetry, patient intensive care data, and insect populations, where data streams at a high rate and the number of predefined patterns is large, it may be impossible for the comparison algorithm to keep up. We propose a novel technique that exploits the commonality among the predefined patterns to allow monitoring at higher bandwidths, while maintaining a guarantee of no false dismissals. Our approach is based on the widely used envelope-based lower bounding technique. Extensive experiments demonstrate that our approach achieves tremendous improvements in performance in the offline case, and significant improvements in the fastest possible arrival rate of the data stream that can be processed with guaranteed no false dismissal.
Li Wei 0001, Eamonn J. Keogh, Helga Van Herle, Agenor Mafra-Neto
ICDM1
2005 Assumption-Free Anomaly Detection in Time Series
Li Wei 0001, Nitin Kumar 0002, Venkata Nishanth Lolla, Eamonn J. Keogh, Stefano Lonardi, Chotirat (Ann) Ratanamahatana
SSDBM1