EDBT 2026 Demo / reviewers in the wild / expert
Yan Zhu 0014
dblp:82/3167-14
· DBLP profile ↗
17ranked-venue papers
10as first author
1since 2021 · last 2021
0000-0002-5952-2108ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 15 · 9 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
11 papers |
Data mining · 76% Indexing and storage engines · 9% Spatial and temporal data management · 6% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% | |
| Theoretical computer science
2 papers |
Algorithms and data structures · 100% |
Topics — the 21 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › temporal data mining
time series mining |
2.0 | 7 | 2021 | Matrix Profile IX: Admissible Time Series Motif Discovery With Missing Data · IEEE Trans. Knowl. Data Eng. 2021 Matrix Profile XVII: Indexing the Matrix Profile to Allow Arbitrary Range Queries · ICDE 2020 Time Series Chains: A Novel Tool for Time Series Data Mining · IJCAI 2018 |
Data mining › structured data mining › graph mining
motif discovery |
1.5 | 4 | 2021 | Matrix Profile IX: Admissible Time Series Motif Discovery With Missing Data · IEEE Trans. Knowl. Data Eng. 2021 VALMOD: A Suite for Easy and Exact Detection of Variable Length Motifs in Data Series · SIGMOD Conference 2018 Matrix Profile X: VALMOD - Scalable Discovery of Variable-Length Motifs in Data Series · SIGMOD Conference 2018 |
Data mining
pattern mining |
1.0 | 4 | 2018 | VALMOD: A Suite for Easy and Exact Detection of Variable Length Motifs in Data Series · SIGMOD Conference 2018 Matrix Profile X: VALMOD - Scalable Discovery of Variable-Length Motifs in Data Series · SIGMOD Conference 2018 Matrix Profile I: All Pairs Similarity Joins for Time Series: A Unifying View That Includes Motifs, Discords and Shapelets · ICDM 2016 |
Data mining › pattern mining › time series motif discovery
matrix profile |
0.7 | 2 | 2020 | Matrix Profile XVII: Indexing the Matrix Profile to Allow Arbitrary Range Queries · ICDE 2020 Matrix Profile I: All Pairs Similarity Joins for Time Series: A Unifying View That Includes Motifs, Discords and Shapelets · ICDM 2016 |
Data mining › structured data mining › graph mining › motif discovery
variable-length motif discovery |
0.7 | 2 | 2018 | VALMOD: A Suite for Easy and Exact Detection of Variable Length Motifs in Data Series · SIGMOD Conference 2018 Matrix Profile X: VALMOD - Scalable Discovery of Variable-Length Motifs in Data Series · SIGMOD Conference 2018 |
Data mining
anomaly detection |
0.6 | 2 | 2019 | Online Amnestic DTW to allow Real-Time Golden Batch Monitoring · KDD 2019 Matrix Profile I: All Pairs Similarity Joins for Time Series: A Unifying View That Includes Motifs, Discords and Shapelets · ICDM 2016 |
Data integration and cleaning
missing data |
0.5 | 1 | 2021 | Matrix Profile IX: Admissible Time Series Motif Discovery With Missing Data · IEEE Trans. Knowl. Data Eng. 2021 |
Data mining › pattern mining
time series motif discovery |
0.5 | 2 | 2016 | Matrix Profile II: Exploiting a Novel Algorithm and GPUs to Break the One Hundred Million Barrier for Time Series Motifs and Joins · ICDM 2016 Matrix Profile I: All Pairs Similarity Joins for Time Series: A Unifying View That Includes Motifs, Discords and Shapelets · ICDM 2016 |
Indexing and storage engines
range index |
0.4 | 1 | 2020 | Matrix Profile XVII: Indexing the Matrix Profile to Allow Arbitrary Range Queries · ICDE 2020 |
Indexing and storage engines › temporal indexing
time series indexing |
0.4 | 1 | 2020 | Matrix Profile XVII: Indexing the Matrix Profile to Allow Arbitrary Range Queries · ICDE 2020 |
Data mining › time series analysis › time series similarity
dynamic time warping |
0.4 | 1 | 2019 | Online Amnestic DTW to allow Real-Time Golden Batch Monitoring · KDD 2019 |
Spatial and temporal data management
time series data |
0.4 | 1 | 2019 | Online Amnestic DTW to allow Real-Time Golden Batch Monitoring · KDD 2019 |
Audio and music processing
music information retrieval |
0.4 | 1 | 2019 | Fast Similarity Matrix Profile for Music Analysis and Exploration · IEEE Trans. Multim. 2019 |
Audio and music processing › music information retrieval
music similarity |
0.4 | 1 | 2019 | Fast Similarity Matrix Profile for Music Analysis and Exploration · IEEE Trans. Multim. 2019 |
Data mining
time series analysis |
0.3 | 1 | 2018 | Matrix Profile XI: SCRIMP++: Time Series Motif Discovery at Interactive Speeds · ICDM 2018 |
Algorithms and data structures
anytime algorithms |
0.3 | 1 | 2018 | Matrix Profile XI: SCRIMP++: Time Series Motif Discovery at Interactive Speeds · ICDM 2018 |
Query processing and optimization
similarity join |
0.2 | 1 | 2016 | Matrix Profile I: All Pairs Similarity Joins for Time Series: A Unifying View That Includes Motifs, Discords and Shapelets · ICDM 2016 |
Spatial and temporal data management › time series data management
time series join |
0.2 | 1 | 2016 | Matrix Profile II: Exploiting a Novel Algorithm and GPUs to Break the One Hundred Million Barrier for Time Series Motifs and Joins · ICDM 2016 |
Information retrieval
similarity search |
0.1 | 1 | 2019 | Fast Similarity Matrix Profile for Music Analysis and Exploration · IEEE Trans. Multim. 2019 |
GPUs and heterogeneous computing
GPU computing |
0.1 | 1 | 2016 | Matrix Profile II: Exploiting a Novel Algorithm and GPUs to Break the One Hundred Million Barrier for Time Series Motifs and Joins · ICDM 2016 |
Algorithms and data structures › data structure design › search structures
indexing |
0.1 | 1 | 2016 | Matrix Profile I: All Pairs Similarity Joins for Time Series: A Unifying View That Includes Motifs, Discords and Shapelets · ICDM 2016 |
Methods — techniques the papers use, named apart from their topics
matrix profile · 1.7subsequence similarity join · 0.8similarity matrix profile · 0.8dynamic time warping · 0.8amnestic DTW · 0.8STOMP · 0.7STAMP · 0.7subsequence matching · 0.6admissible algorithm · 0.5scalable algorithm design · 0.3lower bounding · 0.2early abandoning · 0.2anytime algorithm · 0.2GPU computing · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Matrix Profile IX: Admissible Time Series Motif Discovery With Missing DataabstractThe discovery of time series motifs has emerged as one of the most useful primitives in time series data mining. Researchers have shown its utility for exploratory data mining, summarization, visualization, segmentation, classification, clustering, and rule discovery. Although there has been more than a decade of extensive research, there is still no technique to allow the discovery of time series motifs in the presence of missing data, despite the well-documented ubiquity of missing data in scientific, industrial, and medical datasets. In this work, we introduce a technique for motif discovery in the presence of missing data. We formally prove that our method is admissible, producing no false negatives. We also show that our method can “piggy-back” off the fastest known motif discovery method with a small constant factor time/space overhead. We will demonstrate our approach on diverse datasets with varying amounts of missing data. Yan Zhu 0014, Abdullah Mueen, Eamonn J. Keogh |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | Matrix Profile XVII: Indexing the Matrix Profile to Allow Arbitrary Range QueriesabstractSince its introduction several years ago, the Matrix Profile has received significant attention for two reasons. First, it is a very general representation, allowing for the discovery of time series motifs, discords, chains, joins, shapelets, segmentations etc. Secondly, it can be computed very efficiently, allowing for fast exact computation and ultra-fast approximate computation. For analysts that use the Matrix Profile frequently, its incremental computability means that they can perform ad-hoc analytics at any time, with almost no delay time. However, they can only issue global queries. That is, queries that consider all the data from time zero to the current time. This is a significant limitation, as they may be interested in localized questions about a contiguous subset of the data. For example, "do we have any unusual motifs that correspond with that unusually cool summer two years ago". Such ad-hoc queries would require recomputing the Matrix Profile for the time period in question. This is not an untenable computation, but it could not be done in interactive time. In this work we introduce a novel indexing framework that allows queries about arbitrary ranges to be answered in quasilinear time, allowing such queries to be interactive for the first time. Yan Zhu 0014, Chin-Chia Michael Yeh, Zachary Schall-Zimmerman, Eamonn J. Keogh |
ICDE | 1 |
| 2020 | Matrix profile goes MAD: variable-length motif and discord discovery in data series
Michele Linardi, Yan Zhu 0014, Themis Palpanas, Eamonn J. Keogh |
Data Min. Knowl. Discov. | 2 |
| 2020 | The Swiss army knife of time series data mining: ten useful things you can do with the matrix profile and ten lines of code
Yan Zhu 0014, Shaghayegh Gharghabi, Diego Furtado Silva, Hoang Anh Dau, Chin-Chia Michael Yeh, Nader Shakibay Senobari, Abdulaziz Almaslukh, Kaveh Kamgar, Zachary Schall-Zimmerman, Gareth J. Funning, Abdullah Mueen, Eamonn J. Keogh |
Data Min. Knowl. Discov. | 1 |
| 2019 | Online Amnestic DTW to allow Real-Time Golden Batch MonitoringabstractIn manufacturing, a golden batch is an idealized realization of the perfect process to produce the desired item, typically represented as a multidimensional time series of pressures, temperatures, flow-rates and so forth. The golden batch is sometimes produced from first-principle models, but it is typically created by recording a batch produced by the most experienced engineers on carefully cleaned and calibrated machines. In most cases, the golden batch is only used in post-mortem analysis of a product with an unexpectedly inferior quality, as plant managers attempt to understand where and when the last production attempt went wrong. In this work, we make two contributions to golden batch processing. We introduce an online algorithm that allows practitioners to understand if the process is currently deviating from the golden batch in real-time, allowing engineers to intervene and potentially save the batch. This may be done, for example, by cooling a boiler that is running unexpectedly hot. In addition, we show that our ideas can greatly expand the purview of golden batch monitoring beyond industrial manufacturing. In particular, we show that golden batch monitoring can be used for anomaly detection, attention focusing, and personalized training/skill assessment in a host of novel domains. Chin-Chia Michael Yeh, Yan Zhu 0014, Hoang Anh Dau, Amirali Darvishzadeh, Mikhail Noskov, Eamonn J. Keogh |
KDD | 2 |
| 2019 | Introducing time series chains: a new primitive for time series data mining
Yan Zhu 0014, Makoto Imamura, Daniel Nikovski, Eamonn J. Keogh |
Knowl. Inf. Syst. | 1 |
| 2019 | Fast Similarity Matrix Profile for Music Analysis and ExplorationabstractMost algorithms for music data mining and retrieval analyze the similarity between feature sets extracted from the raw audio. A conventional approach to assess similarities within or between recordings is to create similarity matrices. However, this method requires quadratic space for each comparison and typically requires costly post-processing of the matrix. We have recently proposed SiMPle, a powerful representation based on subsequence similarity join, which is applicable in several music analysis tasks. In this paper, we propose SiMPle-Fast a highly efficient method for exact computation of SiMPle that is up to one order of magnitude faster than SiMPle. Furthermore, we demonstrate the utility of SiMPle-Fast in cover music recognition and thumbnailing tasks and show that our method is significantly faster and more accurate than the state-of-the-art. Diego Furtado Silva, Chin-Chia Michael Yeh, Yan Zhu 0014, Gustavo Batista, Eamonn J. Keogh |
IEEE Trans. Multim. | 3 |
| 2018 | Matrix Profile XI: SCRIMP++: Time Series Motif Discovery at Interactive SpeedsabstractTime series motif discovery is an important primitive for time series analytics, and is used in domains as diverse as neuroscience, music and sports analytics. In recent years, algorithmic advances (coupled with hardware improvements) have greatly expanded the purview of motif discovery. Nevertheless, we argue that there is an insatiable need for further scalability. This is because more than most types of analytics, motif discovery benefits from interactivity. The two state-of-the-art algorithms to find motifs are STOMP, which requires O(n2) time, and STAMP, which, despite being an O(logn) factor slower, is the preferred solution for most applications, as it is a fast converging anytime algorithm. In favorable scenarios STAMP needs only to be run to a small fraction of completion to provide a very accurate approximation of the top-k motifs. In this work we introduce SCRIMP++, an O(n2) time algorithm that is also an anytime algorithm, combining the best features of STOMP and STAMP. As we shall show, SCRIMP++ maintains all the desirable properties of the original algorithms, but converges much faster, in almost all scenarios producing the correct output after spending a tiny fraction of the full computation time. We argue that for many end-users, this allows motif discovery to be performed in interactive sessions. Moreover, this interactivity can be game changing in terms of the analytics that can be performed. Yan Zhu 0014, Chin-Chia Michael Yeh, Zachary Schall-Zimmerman, Kaveh Kamgar, Eamonn J. Keogh |
ICDM | 1 |
| 2018 | Time Series Chains: A Novel Tool for Time Series Data MiningabstractSince their introduction over a decade ago, time se-ries motifs have become a fundamental tool for time series analytics, finding diverse uses in dozens of domains. In this work we introduce Time Series Chains, which are related to, but distinct from, time series motifs. Informally, time series chains are a temporally ordered set of subsequence patterns, such that each pattern is similar to the pattern that preceded it, but the first and last patterns are arbi-trarily dissimilar. In the discrete space, this is simi-lar to extracting the text chain “hit, hot, dot, dog” from a paragraph. The first and last words have nothing in common, yet they are connected by a chain of words with a small mutual difference. Time Series Chains can capture the evolution of systems, and help predict the future. As such, they potentially have implications for prognostics. In this work, we introduce a robust definition of time series chains, and a scalable algorithm that allows us to discover them in massive datasets. Yan Zhu 0014, Makoto Imamura, Daniel Nikovski, Eamonn J. Keogh |
IJCAI | 1 |
| 2018 | Matrix Profile X: VALMOD - Scalable Discovery of Variable-Length Motifs in Data SeriesabstractIn the last fifteen years, data series motif discovery has emerged as one of the most useful primitives for data series mining, with applications to many domains, including robotics, entomology, seismology, medicine, and climatology. Nevertheless, the state-of-the-art motif discovery tools still require the user to provide the motif length. Yet, in at least some cases, the choice of motif length is critical and unforgiving. Unfortunately, the obvious brute-force solution, which tests all lengths within a given range, is computationally untenable. In this work, we introduce VALMOD, an exact and scalable motif discovery algorithm that efficiently finds all motifs in a given range of lengths. We evaluate our approach with five diverse real datasets, and demonstrate that it is up to 20 times faster than the state-of-the-art. Our results also show that removing the unrealistic assumption that the user knows the correct length, can often produce more intuitive and actionable results, which could have been missed otherwise. Michele Linardi, Yan Zhu 0014, Themis Palpanas, Eamonn J. Keogh |
SIGMOD Conference | 2 |
| 2018 | VALMOD: A Suite for Easy and Exact Detection of Variable Length Motifs in Data SeriesabstractData series motif discovery represents one of the most useful primitives for data series mining, with applications to many domains, such as robotics, entomology, seismology, medicine, and climatology, and others. The state-of-the-art motif discovery tools still require the user to provide the motif length. Yet, in several cases, the choice of motif length is critical for their detection. Unfortunately, the obvious brute-force solution, which tests all lengths within a given range, is computationally untenable, and does not provide any support for ranking motifs at different resolutions (i.e., lengths). We demonstrate VALMOD, our scalable motif discovery algorithm that efficiently finds all motifs in a given range of lengths, and outputs a length-invariant ranking of motifs. Furthermore, we support the analysis process by means of a newly proposed meta-data structure that helps the user to select the most promising pattern length. This demo aims at illustrating in detail the steps of the proposed approach, showcasing how our algorithm and corresponding graphical insights enable users to efficiently identify the correct motifs. Michele Linardi, Yan Zhu 0014, Themis Palpanas, Eamonn J. Keogh |
SIGMOD Conference | 2 |
| 2018 | Time series joins, motifs, discords and shapelets: a unifying view that exploits the matrix profile
Chin-Chia Michael Yeh, Yan Zhu 0014, Liudmila Ulanova, Nurjahan Begum, Yifei Ding, Hoang Anh Dau, Zachary Schall-Zimmerman, Diego Furtado Silva, Abdullah Mueen, Eamonn J. Keogh |
Data Min. Knowl. Discov. | 2 |
| 2018 | Exploiting a novel algorithm and GPUs to break the ten quadrillion pairwise comparisons barrier for time series motifs and joins
Yan Zhu 0014, Zachary Schall-Zimmerman, Nader Shakibay Senobari, Chin-Chia Michael Yeh, Gareth J. Funning, Abdullah Mueen, Philip Brisk, Eamonn J. Keogh |
Knowl. Inf. Syst. | 1 |
| 2017 | Matrix Profile VII: Time Series Chains: A New Primitive for Time Series Data Mining (Best Student Paper Award)abstractSince their introduction over a decade ago, time series motifs have become a fundamental tool for time series analytics, finding diverse uses in dozens of domains. In this work we introduce Time Series Chains, which are related to, but distinct from, time series motifs. Informally, time series chains are a temporally ordered set of subsequence patterns, such that each pattern is similar to the pattern that preceded it, but the first and last patterns are arbitrarily dissimilar. In the discrete space, this is similar to extracting the text chain "hit, hot, dot, dog" from a paragraph. The first and last words have nothing in common, yet they are connected by a chain of words with a small mutual difference. Time series chains can capture the evolution of systems, and help predict the future. As such, they potentially have implications for prognostics. In this work, we introduce a robust definition of time series chains, and a scalable algorithm that allows us to discover them in massive datasets. Yan Zhu 0014, Makoto Imamura, Daniel Nikovski, Eamonn J. Keogh |
ICDM | 1 |
| 2016 | Matrix Profile I: All Pairs Similarity Joins for Time Series: A Unifying View That Includes Motifs, Discords and ShapeletsabstractThe all-pairs-similarity-search (or similarity join) problem has been extensively studied for text and a handful of other datatypes. However, surprisingly little progress has been made on similarity joins for time series subsequences. The lack of progress probably stems from the daunting nature of the problem. For even modest sized datasets the obvious nested-loop algorithm can take months, and the typical speed-up techniques in this domain (i.e., indexing, lower-bounding, triangular-inequality pruning and early abandoning) at best produce one or two orders of magnitude speedup. In this work we introduce a novel scalable algorithm for time series subsequence all-pairs-similarity-search. For exceptionally large datasets, the algorithm can be trivially cast as an anytime algorithm and produce high-quality approximate solutions in reasonable time. The exact similarity join algorithm computes the answer to the time series motif and time series discord problem as a side-effect, and our algorithm incidentally provides the fastest known algorithm for both these extensively-studied problems. We demonstrate the utility of our ideas for two time series data mining problems, including motif discovery and novelty discovery. Chin-Chia Michael Yeh, Yan Zhu 0014, Liudmila Ulanova, Nurjahan Begum, Yifei Ding, Hoang Anh Dau, Diego Furtado Silva, Abdullah Mueen, Eamonn J. Keogh |
ICDM | 2 |
| 2016 | Matrix Profile II: Exploiting a Novel Algorithm and GPUs to Break the One Hundred Million Barrier for Time Series Motifs and JoinsabstractTime series motifs have been in the literature for about fifteen years, but have only recently begun to receive significant attention in the research community. This is perhaps due to the growing realization that they implicitly offer solutions to a host of time series problems, including rule discovery, anomaly detection, density estimation, semantic segmentation, etc. Recent work has improved the scalability to the point where exact motifs can be computed on datasets with up to a million data points in tenable time. However, in some domains, for example seismology, there is an insatiable need to address even larger datasets. In this work we show that a combination of a novel algorithm and a high-performance GPU allows us to significantly improve the scalability of motif discovery. We demonstrate the scalability of our ideas by finding the full set of exact motifs on a dataset with one hundred million subsequences, by far the largest dataset ever mined for time series motifs. Furthermore, we demonstrate that our algorithm can produce actionable insights in seismology and other domains. Yan Zhu 0014, Zachary Schall-Zimmerman, Nader Shakibay Senobari, Chin-Chia Michael Yeh, Gareth J. Funning, Abdullah Mueen, Philip Brisk, Eamonn J. Keogh |
ICDM | 1 |
| 2016 | Irrevocable-choice algorithms for sampling from a stream
Yan Zhu 0014, Eamonn J. Keogh |
Data Min. Knowl. Discov. | 1 |