João Manuel De Almeida Rodrigues

dblp:339/0350 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › temporal data mining
time series mining
0.612022
Matrix Profile XXVI: Mplots: Scaling Time Series Similarity Matrices to Massive Data · ICDM 2022
Data mining › time series analysis
time series segmentation
0.612022
Matrix Profile XXVI: Mplots: Scaling Time Series Similarity Matrices to Massive Data · ICDM 2022
Data mining › time series analysis
time series classification
0.212022
Matrix Profile XXVI: Mplots: Scaling Time Series Similarity Matrices to Massive Data · ICDM 2022

Methods — techniques the papers use, named apart from their topics

multiscale approximation · 0.6matrix profile · 0.6just-in-time recomputation · 0.6
YearPublicationVenuePosition
2022 Matrix Profile XXVI: Mplots: Scaling Time Series Similarity Matrices to Massive Data
abstract
Time series similarity matrices (informally, recurrence plots), are useful tools for time series data mining. They can be used to guide data exploration, and various useful features can be derived from them and then fed into downstream analytics. However, time series similarity matrices suffer from very poor scalability, taxing both time and memory requirements. In this work, we introduce novel ideas that allow us to scale the largest time series similarity matrices that can be examined by several orders of magnitude. The first idea is a novel algorithm to compute the matrices in a way that removes dependency on the subsequence length. This algorithm is so fast that it allows us to now address datasets where the memory limitations begin to dominate. Our second novel contribution is a multiscale algorithm that computes an approximation of the matrix appropriate for the limitations of the user’s memory/screen-resolution, then performs a local, just-in-time recomputation of any region that the user wishes to zoom-in on. Given that we can largely remove time and space barriers, human visual attention then becomes the bottleneck. We further introduce algorithms that search massive matrices with quadrillions of cells and then prioritize regions for later attention by either humans or algorithms. We will demonstrate the utility of our ideas for data exploration, segmentation, and classification in diverse domains.
Maryam Shahcheraghi, Ryan Mercer, João Manuel De Almeida Rodrigues, Audrey Der, Hugo Gamboa, Zachary Schall-Zimmerman, Eamonn J. Keogh
ICDM3