Renjie Wu 0001

dblp:191/2467-1 · DBLP profile ↗
← Back
10ranked-venue papers in the field
6as first author
10since 2021 · last 2024
0000-0001-9326-7772ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 6 (6 first)Data Mining & Knowledge Discovery · 4
YearPublicationVenuePosition
2024 C22MP: the marriage of catch22 and the matrix profile creates a fast, efficient and interpretable anomaly detector
Sadaf Tafazoli, Yue Lu 0003, Renjie Wu 0001, Thirumalai Vinjamoor Akhil Srinivas, Hannah Dela Cruz, Ryan Mercer, Eamonn J. Keogh
Knowl. Inf. Syst.3
2023 Matrix Profile XXIX: C22MP, Fusing catch 22 and the Matrix Profile to Produce an Efficient and Interpretable Anomaly Detector
abstract
The Matrix Profile is a data structure that annotates a time series by recording each subsequence’s Euclidean distance to its nearest neighbor. In recent years the community has shown that using the Matrix Profile it is possible to discover many useful properties of a time series, including repeated behaviors, anomalies, evolving patterns, regimes, etc. However, the Matrix Profile is limited to representing the relationship between the subsequence’s shapes. It is known that, for some domains, useful information is conserved not in the subsequence’s shapes, but in the subsequence’s features. In recent years a new set of features for time series called catch22 has revolutionized feature-based mining of time series. Combining these two ideas seems to offer many possibilities for novel data mining applications, however, there are two difficulties in attempting this. A direct application of the Matrix Profile with the catch22 features would be prohibitively slow. Less obviously, as we will demonstrate, in almost all domains, using all twenty-two of the catch22 features produces poor results, and we must somehow select the subset appropriate for the domain. In this work we introduce novel algorithms to solve both problems and demonstrate that for most domains, the proposed $\mathrm{C}^{22}$MP is a state-of-the-art anomaly detector.
Sadaf Tafazoli, Yue Lu 0003, Renjie Wu 0001, Thirumalai Vinjamoor Akhil Srinivas, Hannah Dela Cruz, Ryan Mercer, Eamonn J. Keogh
ICDM3
2023 DAMP: accurate time series anomaly detection on trillions of datapoints and ultra-fast arriving data streams
Yue Lu 0003, Renjie Wu 0001, Abdullah Mueen, Maria A. Zuluaga, Eamonn J. Keogh
Data Min. Knowl. Discov.2
2023 When is Early Classification of Time Series Meaningful?
abstract
Since its introduction two decades ago, there has been increasing interest in the problem of early classification of time series. This problem generalizes classic time series classification to ask if we can classify a time series subsequence with sufficient accuracy and confidence after seeing only some prefix of a target pattern. The idea is that the earlier classification would allow us to take immediate action, in a domain in which some practical interventions are possible. For example, that intervention might be sounding an alarm or applying the brakes in an automobile. In this work, we make a surprising claim. In spite of the fact that there are dozens of papers on early classification of time series, it is not clear that any of them could ever work in a real-world setting. The problem is not with the algorithms per se but with the vague and underspecified problem description. Essentially all algorithms make implicit and unwarranted assumptions about the problem that will ensure that they will be plagued by false positives and false negatives even if their results suggested that they could obtain near-perfect results. We will explain our findings with novel insights and experiments and offer recommendations to the community.
Renjie Wu 0001, Audrey Der, Eamonn J. Keogh
IEEE Trans. Knowl. Data Eng.1
2023 Current Time Series Anomaly Detection Benchmarks are Flawed and are Creating the Illusion of Progress
abstract
Time series anomaly detection has been a perennially important topic in data science, with papers dating back to the 1950s. However, in recent years there has been an explosion of interest in this topic, much of it driven by the success of deep learning in other domains and for other time series tasks. Most of these papers test on one or more of a handful of popular benchmark datasets, created by Yahoo, Numenta, NASA, etc. In this work we make a surprising claim. The majority of the individual exemplars in these datasets suffer from one or more of four flaws. Because of these four flaws, we believe that many published comparisons of anomaly detection algorithms may be unreliable, and more importantly, much of the apparent progress in recent years may be illusionary. In addition to demonstrating these claims, with this paper we introduce the UCR Time Series Anomaly Archive. We believe that this resource will perform a similar role as the UCR Time Series Classification Archive, by providing the community with a benchmark that allows meaningful comparisons between approaches and a meaningful gauge of overall progress.
Renjie Wu 0001, Eamonn J. Keogh
IEEE Trans. Knowl. Data Eng.1
2022 When is Early Classification of Time Series Meaningful? (Extended Abstract)
abstract
The problem of early classification of time series (ETSC) generalizes classic time series classification to ask if we can classify a time series subsequence with sufficient accuracy and confidence after seeing only some prefix of a target pattern. The idea is that the earlier classification would allow us to take immediate actions, such as sounding an alarm or applying the brakes in an automobile. In this work, we make a surprising claim. In spite of the fact that there are dozens of papers on ETSC, it is not clear that any of them could ever work in a real-world setting. The issue is not with the algorithms per se, but with the vague and underspecified problem definition.
Renjie Wu 0001, Audrey Der, Eamonn J. Keogh
ICDE1
2022 Current Time Series Anomaly Detection Benchmarks are Flawed and are Creating the Illusion of Progress (Extended Abstract)
abstract
Most of the time series anomaly detection papers tested on a handful of popular benchmark datasets, created by Yahoo [1], Numenta [2], NASA [3] or Pei's Lab (OMNI) [4], etc. There is a strong implicit assumption that doing well on these public datasets is a sufficient condition to declare an anomaly detection algorithm is useful. In this work, we make a surprising claim. The majority of the individual exemplars in these dataset suffers from one or more of four flaws: triviality, unrealistic anomaly density, mislabeled ground truth and run-to-failure bias. Because of these four flaws, we believe that most published comparisons of anomaly detection algorithms may be unreliable, and more importantly, much of the apparent progress in recent years may be illusionary.
Renjie Wu 0001, Eamonn J. Keogh
ICDE1
2022 Matrix Profile XXIV: Scaling Time Series Anomaly Detection to Trillions of Datapoints and Ultra-fast Arriving Data Streams
abstract
Time series anomaly detection remains one of the most active areas of research in data mining. In spite of the dozens of creative solutions proposed for this problem, recent empirical evidence suggests that time series discords, a relatively simple twenty-year old distance-based technique, remains among the state-of-art techniques. While there are many algorithms for computing the time series discords, they all have limitations. First, they are limited to the batch case, whereas the online case is more actionable. Second, these algorithms exhibit poor scalability beyond tens of thousands of datapoints. In this work we introduce DAMP, a novel algorithm that addresses both these issues. DAMP computes exact left-discords on fast arriving streams, at up to 300,000 Hz using a commodity desktop. This allows us to find time series discords in datasets with trillions of datapoints for the first time. We will demonstrate the utility of our algorithm with the most ambitious set of time series anomaly detection experiments ever conducted.
Yue Lu 0003, Renjie Wu 0001, Abdullah Mueen, Maria A. Zuluaga, Eamonn J. Keogh
KDD2
2022 FastDTW is Approximate and Generally Slower Than the Algorithm it Approximates
abstract
Many time series data mining problems can be solved with repeated use of distance measure. Examples of such tasks include similarity search, clustering, classification, anomaly detection and segmentation. For over two decades it has been known that the Dynamic Time Warping (DTW) distance measure is the best measure to use for most tasks, in most domains. Because the classic DTW algorithm has quadratic time complexity, many ideas have been introduced to reduce its amortized time, or to quickly approximate it. One of the most cited approximate approaches is FastDTW. The FastDTW algorithm has well over a thousand citations and has been explicitly used in several hundred research efforts. In this work, we make a surprising claim. In any realistic data mining application, theapproximateFastDTW is much slower than theexactDTW. This fact clearly has implications for the community that uses this algorithm: allowing it to address much larger datasets, get exact results, and do so in less time.
Renjie Wu 0001, Eamonn J. Keogh
IEEE Trans. Knowl. Data Eng.1
2021 FastDTW is approximate and Generally Slower than the Algorithm it Approximates (Extended Abstract)
abstract
Many time series data mining problems can be solved with repeated use of distance measure. Examples of such tasks include similarity search, clustering, classification, anomaly detection and segmentation. For over two decades it has been known that the Dynamic Time Warping (DTW) distance measure is the best measure to use for most tasks, in most domains. Because the classic DTW algorithm has quadratic time complexity, many ideas have been introduced to reduce its amortized time, or to quickly approximate it. One of the most cited approximate approaches is FastDTW. The FastDTW algorithm has well over a thousand citations and has been explicitly used in several hundred research efforts. In this work, we make a surprising claim. In any realistic data mining application, the approximate FastDTW is much slower than the exact DTW. This fact clearly has implications for the community that uses this algorithm: allowing it to address much larger datasets, get exact results, and do so in less time.
Renjie Wu 0001, Eamonn J. Keogh
ICDE1