EDBT 2026 Demo / reviewers in the wild / expert
Djamel Edine Yagoubi
dblp:211/2848
· DBLP profile ↗
7ranked-venue papers
4as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 2 first-authorTheory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Indexing and storage engines · 44% Query processing and optimization · 32% Data mining · 13% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
similarity query processing |
0.7 | 2 | 2020 | Massively Distributed Time Series Indexing and Querying · IEEE Trans. Knowl. Data Eng. 2020 DPiSAX: Massively Distributed Partitioned iSAX · ICDM 2017 |
Indexing and storage engines › temporal indexing
time series indexing |
0.7 | 2 | 2020 | Massively Distributed Time Series Indexing and Querying · IEEE Trans. Knowl. Data Eng. 2020 DPiSAX: Massively Distributed Partitioned iSAX · ICDM 2017 |
Indexing and storage engines
distributed indexing |
0.3 | 1 | 2017 | DPiSAX: Massively Distributed Partitioned iSAX · ICDM 2017 |
Data mining › temporal data mining
time series mining |
0.3 | 1 | 2017 | DPiSAX: Massively Distributed Partitioned iSAX · ICDM 2017 |
Distributed and cloud data management
distributed query processing |
0.1 | 1 | 2020 | Massively Distributed Time Series Indexing and Querying · IEEE Trans. Knowl. Data Eng. 2020 |
Data stream processing
load balancing |
0.1 | 1 | 2020 | Massively Distributed Time Series Indexing and Querying · IEEE Trans. Knowl. Data Eng. 2020 |
Methods — techniques the papers use, named apart from their topics
load balancing · 0.7parallel indexing · 0.4iSAX · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | BestNeighbor: efficient evaluation of kNN queries on large time series databases
Oleksandra Levchenko, Boyan Kolev, Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Themis Palpanas, Dennis E. Shasha, Patrick Valduriez |
Knowl. Inf. Syst. | 3 |
| 2020 | Massively Distributed Time Series Indexing and QueryingabstractIndexing is crucial for many data mining tasks that rely on efficient and effective similarity query processing. Consequently, indexing large volumes of time series, along with high performance similarity query processing, have became topics of high interest. For many applications across diverse domains though, the amount of data to be processed might be intractable for a single machine, making existing centralized indexing solutions inefficient. We propose a parallel indexing solution that gracefully scales to billions of time series, and a parallel query processing strategy that, given a batch of queries, efficiently exploits the index. Our experiments, on both synthetic and real world data, illustrate that our index creation algorithm works on four billion time series in less than five hours, while the state of the art centralized algorithms do not scale and have their limit on 1 billion time series, where they need more than five days. Also, our distributed querying algorithm is able to efficiently process millions of queries over collections of billions of time series, thanks to an effective load balancing mechanism. Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Themis Palpanas |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2019 | Distributed Algorithms to Find Similar Time SeriesabstractInternational audience Oleksandra Levchenko, Boyan Kolev, Djamel Edine Yagoubi, Dennis E. Shasha, Themis Palpanas, Patrick Valduriez, Reza Akbarinia, Florent Masseglia |
ECML/PKDD (3) | 3 |
| 2018 | Spark-parSketch: A Massively Distributed Indexing of Time Series DatasetsabstractA growing number of domains (finance, seismology, internet-of-things, etc.) collect massive time series. When the number of series grow to the hundreds of millions or even billions, similarity queries become intractable on a single machine. Further, naive (quadratic) parallelization won't work well. So, we need both efficient indexing and parallelization. We propose a demonstration of Spark-parSketch, a complete solution based on sketches / random projections to efficiently perform both the parallel indexing of large sets of time series and a similarity search on them. Because our method is approximate, we explore the tradeoff between time and precision. A video showing the dynamics of the demonstration can be found by the link http://parsketch.gforge.inria.fr/video/parSketchdemo_720p.mov. Oleksandra Levchenko, Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Boyan Kolev, Dennis E. Shasha |
CIKM | 2 |
| 2018 | ParCorr: efficient parallel methods to identify similar time series pairs across sliding windows
Djamel Edine Yagoubi, Reza Akbarinia, Boyan Kolev, Oleksandra Levchenko, Florent Masseglia, Patrick Valduriez, Dennis E. Shasha |
Data Min. Knowl. Discov. | 1 |
| 2017 | RadiusSketch: Massively Distributed Indexing of Time SeriesabstractPerforming similarity queries on hundreds of millions of time series is a challenge requiring both efficient indexing techniques and parallelization. We propose a sketch/random projection-based approach that scales nearly linearly in parallel environments, and provides high quality answers. We illustrate the performance of our approach, called RadiusSketch, on real and synthetic datasets of up to 1 Terabytes and 500 million time series. The sketch method, as we have implemented, is superior in both quality and response time compared with the state of the art approach, iSAX2+. Already, in the sequential case it improves recall and precision by a factor of two, while giving shorter response times. In a parallel environment with 32 processors, on both real and synthetic data, our parallel approach improves by a factor of up to 100 in index time construction and up to 15 in query answering time. Finally, our data structure makes use of idle computing time to improve the recall and precision yet further. Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Dennis E. Shasha |
DSAA | 1 |
| 2017 | DPiSAX: Massively Distributed Partitioned iSAXabstractIndexing is crucial for many data mining tasks that rely on efficient and effective similarity query processing. Consequently, indexing large volumes of time series, along with high performance similarity query processing, have became topics of high interest. For many applications across diverse domains though, the amount of data to be processed might be intractable for a single machine, making existing centralized indexing solutions inefficient. We propose a parallel indexing solution that gracefully scales to billions of time series, and a parallel query processing strategy that, given a batch of queries, efficiently exploits the index. Our experiments, on both synthetic and real world data, illustrate that our index creation algorithm works on 1 billion time series in less than 2 hours, while the state of the art centralized algorithms need more than 5 days. Also, our distributed querying algorithm is able to efficiently process millions of queries over collections of billions of time series, thanks to an effective load balancing mechanism. Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Themis Palpanas |
ICDM | 1 |