EDBT 2026 Demo / reviewers in the wild / expert
Yue Lu 0003
dblp:74/6493-3
· DBLP profile ↗
5ranked-venue papers in the field
3as first author
5since 2021 · last 2024
0000-0003-4812-9658ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 5 (3 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | C22MP: the marriage of catch22 and the matrix profile creates a fast, efficient and interpretable anomaly detector
Sadaf Tafazoli, Yue Lu 0003, Renjie Wu 0001, Thirumalai Vinjamoor Akhil Srinivas, Hannah Dela Cruz, Ryan Mercer, Eamonn J. Keogh |
Knowl. Inf. Syst. | 2 |
| 2023 | Matrix Profile XXX: MADRID: A Hyper-Anytime and Parameter-Free Algorithm to Find Time Series Anomalies of all LengthsabstractIn recent years there has been increasing evidence that one of the simplest time series anomaly detection methods, time series discords, remains one of the most effective methods. However, time series discords have one notable issue; the anomalies discovered depend on the algorithm’s only input parameter, the subsequence length. The obvious way to bypass this issue is to find anomalies at every possible length, however this seems to be untenably slow. In this work we introduce MADRID, an algorithm to efficiently solve the all-discords problem. We show that we can reduce the absolute time to compute all-discords, and that by using a novel computation ordering strategy, MADRID is a Hyper-Anytime Algorithm. We will formally define this term later, but this refers to an anytime algorithm that converges exceptionally fast. In practice this means that for most real-world analytical tasks, the user can interact with their data in real-time. The ability to compute anomalies of all lengths produces the issue of ranking anomalies of different lengths. We further introduce novel algorithms for this task. We demonstrate the utility of MADRID in various domains and show that it allows us to Find anomalies that would otherwise escape our attention. Yue Lu 0003, Thirumalai Vinjamoor Akhil Srinivas, Takaaki Nakamura, Makoto Imamura, Eamonn J. Keogh |
ICDM | 1 |
| 2023 | Matrix Profile XXIX: C22MP, Fusing catch 22 and the Matrix Profile to Produce an Efficient and Interpretable Anomaly DetectorabstractThe Matrix Profile is a data structure that annotates a time series by recording each subsequence’s Euclidean distance to its nearest neighbor. In recent years the community has shown that using the Matrix Profile it is possible to discover many useful properties of a time series, including repeated behaviors, anomalies, evolving patterns, regimes, etc. However, the Matrix Profile is limited to representing the relationship between the subsequence’s shapes. It is known that, for some domains, useful information is conserved not in the subsequence’s shapes, but in the subsequence’s features. In recent years a new set of features for time series called catch22 has revolutionized feature-based mining of time series. Combining these two ideas seems to offer many possibilities for novel data mining applications, however, there are two difficulties in attempting this. A direct application of the Matrix Profile with the catch22 features would be prohibitively slow. Less obviously, as we will demonstrate, in almost all domains, using all twenty-two of the catch22 features produces poor results, and we must somehow select the subset appropriate for the domain. In this work we introduce novel algorithms to solve both problems and demonstrate that for most domains, the proposed $\mathrm{C}^{22}$MP is a state-of-the-art anomaly detector. Sadaf Tafazoli, Yue Lu 0003, Renjie Wu 0001, Thirumalai Vinjamoor Akhil Srinivas, Hannah Dela Cruz, Ryan Mercer, Eamonn J. Keogh |
ICDM | 2 |
| 2023 | DAMP: accurate time series anomaly detection on trillions of datapoints and ultra-fast arriving data streams
Yue Lu 0003, Renjie Wu 0001, Abdullah Mueen, Maria A. Zuluaga, Eamonn J. Keogh |
Data Min. Knowl. Discov. | 1 |
| 2022 | Matrix Profile XXIV: Scaling Time Series Anomaly Detection to Trillions of Datapoints and Ultra-fast Arriving Data StreamsabstractTime series anomaly detection remains one of the most active areas of research in data mining. In spite of the dozens of creative solutions proposed for this problem, recent empirical evidence suggests that time series discords, a relatively simple twenty-year old distance-based technique, remains among the state-of-art techniques. While there are many algorithms for computing the time series discords, they all have limitations. First, they are limited to the batch case, whereas the online case is more actionable. Second, these algorithms exhibit poor scalability beyond tens of thousands of datapoints. In this work we introduce DAMP, a novel algorithm that addresses both these issues. DAMP computes exact left-discords on fast arriving streams, at up to 300,000 Hz using a commodity desktop. This allows us to find time series discords in datasets with trillions of datapoints for the first time. We will demonstrate the utility of our algorithm with the most ambitious set of time series anomaly detection experiments ever conducted. Yue Lu 0003, Renjie Wu 0001, Abdullah Mueen, Maria A. Zuluaga, Eamonn J. Keogh |
KDD | 1 |