Djamel Edine Yagoubi

dblp:211/2848 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 2 first-authorTheory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Indexing and storage engines · 44% Query processing and optimization · 32% Data mining · 13%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
similarity query processing
0.722020
Massively Distributed Time Series Indexing and Querying · IEEE Trans. Knowl. Data Eng. 2020
DPiSAX: Massively Distributed Partitioned iSAX · ICDM 2017
Indexing and storage engines › temporal indexing
time series indexing
0.722020
Massively Distributed Time Series Indexing and Querying · IEEE Trans. Knowl. Data Eng. 2020
DPiSAX: Massively Distributed Partitioned iSAX · ICDM 2017
Indexing and storage engines
distributed indexing
0.312017
DPiSAX: Massively Distributed Partitioned iSAX · ICDM 2017
Data mining › temporal data mining
time series mining
0.312017
DPiSAX: Massively Distributed Partitioned iSAX · ICDM 2017
Distributed and cloud data management
distributed query processing
0.112020
Massively Distributed Time Series Indexing and Querying · IEEE Trans. Knowl. Data Eng. 2020
Data stream processing
load balancing
0.112020
Massively Distributed Time Series Indexing and Querying · IEEE Trans. Knowl. Data Eng. 2020

Methods — techniques the papers use, named apart from their topics

load balancing · 0.7parallel indexing · 0.4iSAX · 0.3
YearPublicationVenuePosition
2021 BestNeighbor: efficient evaluation of kNN queries on large time series databases
Oleksandra Levchenko, Boyan Kolev, Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Themis Palpanas, Dennis E. Shasha, Patrick Valduriez
Knowl. Inf. Syst.3
2020 Massively Distributed Time Series Indexing and Querying
abstract
Indexing is crucial for many data mining tasks that rely on efficient and effective similarity query processing. Consequently, indexing large volumes of time series, along with high performance similarity query processing, have became topics of high interest. For many applications across diverse domains though, the amount of data to be processed might be intractable for a single machine, making existing centralized indexing solutions inefficient. We propose a parallel indexing solution that gracefully scales to billions of time series, and a parallel query processing strategy that, given a batch of queries, efficiently exploits the index. Our experiments, on both synthetic and real world data, illustrate that our index creation algorithm works on four billion time series in less than five hours, while the state of the art centralized algorithms do not scale and have their limit on 1 billion time series, where they need more than five days. Also, our distributed querying algorithm is able to efficiently process millions of queries over collections of billions of time series, thanks to an effective load balancing mechanism.
Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Themis Palpanas
IEEE Trans. Knowl. Data Eng.1
2019 Distributed Algorithms to Find Similar Time Series
abstract
International audience
Oleksandra Levchenko, Boyan Kolev, Djamel Edine Yagoubi, Dennis E. Shasha, Themis Palpanas, Patrick Valduriez, Reza Akbarinia, Florent Masseglia
ECML/PKDD (3)3
2018 Spark-parSketch: A Massively Distributed Indexing of Time Series Datasets
abstract
A growing number of domains (finance, seismology, internet-of-things, etc.) collect massive time series. When the number of series grow to the hundreds of millions or even billions, similarity queries become intractable on a single machine. Further, naive (quadratic) parallelization won't work well. So, we need both efficient indexing and parallelization. We propose a demonstration of Spark-parSketch, a complete solution based on sketches / random projections to efficiently perform both the parallel indexing of large sets of time series and a similarity search on them. Because our method is approximate, we explore the tradeoff between time and precision. A video showing the dynamics of the demonstration can be found by the link http://parsketch.gforge.inria.fr/video/parSketchdemo_720p.mov.
Oleksandra Levchenko, Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Boyan Kolev, Dennis E. Shasha
CIKM2
2018 ParCorr: efficient parallel methods to identify similar time series pairs across sliding windows
Djamel Edine Yagoubi, Reza Akbarinia, Boyan Kolev, Oleksandra Levchenko, Florent Masseglia, Patrick Valduriez, Dennis E. Shasha
Data Min. Knowl. Discov.1
2017 RadiusSketch: Massively Distributed Indexing of Time Series
abstract
Performing similarity queries on hundreds of millions of time series is a challenge requiring both efficient indexing techniques and parallelization. We propose a sketch/random projection-based approach that scales nearly linearly in parallel environments, and provides high quality answers. We illustrate the performance of our approach, called RadiusSketch, on real and synthetic datasets of up to 1 Terabytes and 500 million time series. The sketch method, as we have implemented, is superior in both quality and response time compared with the state of the art approach, iSAX2+. Already, in the sequential case it improves recall and precision by a factor of two, while giving shorter response times. In a parallel environment with 32 processors, on both real and synthetic data, our parallel approach improves by a factor of up to 100 in index time construction and up to 15 in query answering time. Finally, our data structure makes use of idle computing time to improve the recall and precision yet further.
Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Dennis E. Shasha
DSAA1
2017 DPiSAX: Massively Distributed Partitioned iSAX
abstract
Indexing is crucial for many data mining tasks that rely on efficient and effective similarity query processing. Consequently, indexing large volumes of time series, along with high performance similarity query processing, have became topics of high interest. For many applications across diverse domains though, the amount of data to be processed might be intractable for a single machine, making existing centralized indexing solutions inefficient. We propose a parallel indexing solution that gracefully scales to billions of time series, and a parallel query processing strategy that, given a batch of queries, efficiently exploits the index. Our experiments, on both synthetic and real world data, illustrate that our index creation algorithm works on 1 billion time series in less than 2 hours, while the state of the art centralized algorithms need more than 5 days. Also, our distributed querying algorithm is able to efficiently process millions of queries over collections of billions of time series, thanks to an effective load balancing mechanism.
Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Themis Palpanas
ICDM1