Reza Akbarinia

dblp:25/3618 · DBLP profile ↗
← Back
43ranked-venue papers in the field
8as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 26 (7 first)Data Mining & Knowledge Discovery · 12 (1 first)Information Retrieval & Web Search · 3Knowledge Engineering, Semantic Web & Information Systems · 1Business Process & Enterprise Data · 1
YearPublicationVenuePosition
2026 Efficient Detection of Seasonal Bursts: Applications to Climate Time Series
Guillaume Coulaud, Audrey Brouillet, Reza Akbarinia, Florent Masseglia, Dennis E. Shasha
DEXA (2)3
2025 ClimBurst: A Dynamic Visualization Tool to Display Climatological Anomalies over Time and Space
abstract
Detecting abnormal climate events across temporal and spatial scales is crucial to the understanding of local and regional climate trends. This demonstration introduces ClimBurst, a dynamic tool to detect climate bursts, which are unusually high or low values of one or more climate variables over some time interval. ClimBurst detects bursts without prior assumptions about their temporal duration. The demonstration will allow users to interact directly with our system to see both a summary showing the presence/absence of bursts over a user-specified year and spatial range. The demonstration will also allow users to perform time-travel queries to see how bursts propagate over space and time.
Guillaume Coulaud, Benoit Lange, Dennis E. Shasha, Audrey Brouillet, Reza Akbarinia, Florent Masseglia
CIKM5
2025 Scalable and accurate online multivariate anomaly detection
Rebecca Salles, Benoit Lange, Reza Akbarinia, Florent Masseglia, Eduardo S. Ogasawara, Esther Pacitti
Inf. Syst.3
2024 A One-Health Platform for Antimicrobial Resistance Data Analytics
abstract
Antimicrobial resistance (AMR) poses potentially critical health issues for human and animal populations in the near future. To meet this challenge, we need to adopt a "One Health" strategy, which involves studying and linking information from human and animal populations, as well as from the environment.
Benoit Lange, Reza Akbarinia, Florent Masseglia
CIKM2
2023 kNN matrix profile for knowledge discovery from time series
Tanmoy Mondal, Reza Akbarinia, Florent Masseglia
Data Min. Knowl. Discov.2
2022 Parallel Techniques for Variable Size Segmentation of Time Series Datasets
Lamia Djebour, Reza Akbarinia, Florent Masseglia
ADBIS2
2021 Efficient Incremental Computation of Aggregations over Sliding Windows
abstract
Computing aggregation over sliding windows, i.e., finite subsets of an unbounded stream, is a core operation in streaming analytics. We propose PBA (Parallel Boundary Aggregator), a novel parallel algorithm that groups continuous slices of streaming values into chunks and exploits two buffers, cumulative slice aggregations and left cumulative slice aggregations, to compute sliding window aggregations efficiently. PBA runs in O(1) time, performing at most 3 merging operations per slide while consuming O(n) space for windows with n partial aggregations. Our empirical experiments demonstrate that PBA can improve throughput up to 4X while reducing latency, compared to state-of-the-art algorithms.
Chao Zhang 0045, Reza Akbarinia, Farouk Toumani
KDD2
2021 BestNeighbor: efficient evaluation of kNN queries on large time series databases
Oleksandra Levchenko, Boyan Kolev, Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Themis Palpanas, Dennis E. Shasha, Patrick Valduriez
Knowl. Inf. Syst.4
2020 Massively Distributed Time Series Indexing and Querying
abstract
Indexing is crucial for many data mining tasks that rely on efficient and effective similarity query processing. Consequently, indexing large volumes of time series, along with high performance similarity query processing, have became topics of high interest. For many applications across diverse domains though, the amount of data to be processed might be intractable for a single machine, making existing centralized indexing solutions inefficient. We propose a parallel indexing solution that gracefully scales to billions of time series, and a parallel query processing strategy that, given a batch of queries, efficiently exploits the index. Our experiments, on both synthetic and real world data, illustrate that our index creation algorithm works on four billion time series in less than five hours, while the state of the art centralized algorithms do not scale and have their limit on 1 billion time series, where they need more than five days. Also, our distributed querying algorithm is able to efficiently process millions of queries over collections of billions of time series, thanks to an effective load balancing mechanism.
Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Themis Palpanas
IEEE Trans. Knowl. Data Eng.2
2019 Pipelined Implementation of a Parallel Streaming Method for Time Series Correlation Discovery on Sliding Windows
Boyan Kolev, Reza Akbarinia, Ricardo Jiménez-Peris, Oleksandra Levchenko, Florent Masseglia, Marta Patiño-Martínez, Patrick Valduriez
DATA2
2019 Distributed Algorithms to Find Similar Time Series
abstract
International audience
Oleksandra Levchenko, Boyan Kolev, Djamel Edine Yagoubi, Dennis E. Shasha, Themis Palpanas, Patrick Valduriez, Reza Akbarinia, Florent Masseglia
ECML/PKDD (3)7
2018 Spark-parSketch: A Massively Distributed Indexing of Time Series Datasets
abstract
A growing number of domains (finance, seismology, internet-of-things, etc.) collect massive time series. When the number of series grow to the hundreds of millions or even billions, similarity queries become intractable on a single machine. Further, naive (quadratic) parallelization won't work well. So, we need both efficient indexing and parallelization. We propose a demonstration of Spark-parSketch, a complete solution based on sketches / random projections to efficiently perform both the parallel indexing of large sets of time series and a similarity search on them. Because our method is approximate, we explore the tradeoff between time and precision. A video showing the dynamics of the demonstration can be found by the link http://parsketch.gforge.inria.fr/video/parSketchdemo_720p.mov.
Oleksandra Levchenko, Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Boyan Kolev, Dennis E. Shasha
CIKM3
2018 Answering Top-k Queries over Outsourced Sensitive Data in the Cloud
Sakina Mahboubi, Reza Akbarinia, Patrick Valduriez
DEXA (1)2
2018 A Differentially Private Index for Range Query Processing in Clouds
abstract
Performing non-aggregate range queries on cloud stored data, while achieving both privacy and efficiency is a challenging problem. This paper proposes constructing a differentially private index to an outsourced encrypted dataset. Efficiency is enabled by using a cleartext index structure to perform range queries. Security relies on both differential privacy (of the index) and semantic security (of the encrypted dataset). Our solution, PINED-RQ develops algorithms for building and updating the differentially private index. Compared to state-of-the-art secure index based range query processing approaches, PINED-RQ executes queries in the order of at least one magnitude faster. The security of PINED-RQ is proved and its efficiency is assessed by an extensive experimental validation.
Cetin Sahin, Tristan Allard, Reza Akbarinia, Amr El Abbadi, Esther Pacitti
ICDE3
2018 Top-k Query Processing over Distributed Sensitive Data
abstract
Distributed systems provide users with powerful capabilities to store and process their data in third-party machines. However, the privacy of the outsourced data is not guaranteed. One solution for protecting the user data against privacy attacks is to encrypt the sensitive data before sending to the nodes of the distributed system. Then, the main problem is to evaluate user queries over the encrypted data.
Sakina Mahboubi, Reza Akbarinia, Patrick Valduriez
IDEAS2
2018 ParCorr: efficient parallel methods to identify similar time series pairs across sliding windows
Djamel Edine Yagoubi, Reza Akbarinia, Boyan Kolev, Oleksandra Levchenko, Florent Masseglia, Patrick Valduriez, Dennis E. Shasha
Data Min. Knowl. Discov.2
2017 Massively Distributed Environments and Closed Itemset Mining: The DCIM Approach
Mehdi Zitouni, Reza Akbarinia, Sadok Ben Yahia, Florent Masseglia
CAiSE2
2017 TARDIS: Optimal Execution of Scientific Workflows in Apache Spark
Daniel Gaspar, Fábio Porto 0001, Reza Akbarinia, Esther Pacitti
DaWaK3
2017 RadiusSketch: Massively Distributed Indexing of Time Series
abstract
Performing similarity queries on hundreds of millions of time series is a challenge requiring both efficient indexing techniques and parallelization. We propose a sketch/random projection-based approach that scales nearly linearly in parallel environments, and provides high quality answers. We illustrate the performance of our approach, called RadiusSketch, on real and synthetic datasets of up to 1 Terabytes and 500 million time series. The sketch method, as we have implemented, is superior in both quality and response time compared with the state of the art approach, iSAX2+. Already, in the sequential case it improves recall and precision by a factor of two, while giving shorter response times. In a parallel environment with 32 processors, on both real and synthetic data, our parallel approach improves by a factor of up to 100 in index time construction and up to 15 in query answering time. Finally, our data structure makes use of idle computing time to improve the recall and precision yet further.
Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Dennis E. Shasha
DSAA2
2017 DPiSAX: Massively Distributed Partitioned iSAX
abstract
Indexing is crucial for many data mining tasks that rely on efficient and effective similarity query processing. Consequently, indexing large volumes of time series, along with high performance similarity query processing, have became topics of high interest. For many applications across diverse domains though, the amount of data to be processed might be intractable for a single machine, making existing centralized indexing solutions inefficient. We propose a parallel indexing solution that gracefully scales to billions of time series, and a parallel query processing strategy that, given a batch of queries, efficiently exploits the index. Our experiments, on both synthetic and real world data, illustrate that our index creation algorithm works on 1 billion time series in less than 2 hours, while the state of the art centralized algorithms need more than 5 days. Also, our distributed querying algorithm is able to efficiently process millions of queries over collections of billions of time series, thanks to an effective load balancing mechanism.
Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Themis Palpanas
ICDM2
2017 A highly scalable parallel algorithm for maximally informative k-itemset mining
Saber Salah, Reza Akbarinia, Florent Masseglia
Knowl. Inf. Syst.2
2017 Data placement in massively distributed environments for fast parallel mining of frequent itemsets
Saber Salah, Reza Akbarinia, Florent Masseglia
Knowl. Inf. Syst.2
2016 FP-Hadoop: Efficient processing of skewed MapReduce jobs
Miguel Liroz-Gistau, Reza Akbarinia, Divyakant Agrawal, Patrick Valduriez
Inf. Syst.2
2015 An Efficient Solution for Processing Skewed MapReduce Jobs
Reza Akbarinia, Miguel Liroz-Gistau, Divyakant Agrawal, Patrick Valduriez
DEXA (2)1
2015 Data Partitioning for Fast Mining of Frequent Itemsets in Massively Distributed Environments
Saber Salah, Reza Akbarinia, Florent Masseglia
DEXA (1)2
2015 A Prime Number Based Approach for Closed Frequent Itemset Mining in Big Data
Mehdi Zitouni, Reza Akbarinia, Sadok Ben Yahia, Florent Masseglia
DEXA (1)2
2015 Fast Parallel Mining of Maximally Informative k-Itemsets in Big Data
abstract
The discovery of informative itemsets is a fundamental building block in data analytics and information retrieval. While the problem has been widely studied, only few solutions scale. This is particularly the case when i) the data set is massive, calling for large-scale distribution, and/or ii) the length k of the informative itemset to be discovered is high. In this paper, we address the problem of parallel mining of maximally informative k-itemsets (miki) based on joint entropy. We propose PHIKS (Parallel Highly Informative K-ItemSet) a highly scalable, parallel miki mining algorithm. PHIKS renders the mining process of large scale databases (up to terabytes of data) succinct and effective. Its mining process is made up of only two efficient parallel jobs. With PHIKS, we provide a set of significant optimizations for calculating the joint entropies of miki having different sizes, which drastically reduces the execution time of the mining process. PHIKS has been extensively evaluated using massive real-world data sets. Our experimental results confirm the effectiveness of our proposal by the significant scale-up obtained with high itemsets length and over very large databases.
Saber Salah, Reza Akbarinia, Florent Masseglia
ICDM2
2015 Corrigendum to "Best position algorithms for efficient top-k query processing" [Inf. Syst 36(6) (2011) 973-989]
Reza Akbarinia, Esther Pacitti, Patrick Valduriez
Inf. Syst.1
2015 Profile Diversity for Query Processing using User Recommendations
Maximilien Servajean, Reza Akbarinia, Esther Pacitti, Sihem Amer-Yahia
Inf. Syst.2
2015 FP-Hadoop: Efficient Execution of Parallel Jobs Over Skewed Data
abstract
Big data parallel frameworks, such as MapReduce or Spark have been praised for their high scalability and performance, but show poor performance in the case of data skew. There are important cases where a high percentage of processing in the reduce side ends up being done by only one node. In this demonstration, we illustrate the use of FP-Hadoop , a system that efficiently deals with data skew in MapReduce jobs. In FP-Hadoop, there is a new phase, called intermediate reduce (IR) , in which blocks of intermediate values, constructed dynamically, are processed by intermediate reduce workers in parallel, by using a scheduling strategy. Within the IR phase, even if all intermediate values belong to only one key, the main part of the reducing work can be done in parallel using the computing resources of all available workers. We implemented a prototype of FP-Hadoop, and conducted extensive experiments over synthetic and real datasets. We achieve excellent performance gains compared to native Hadoop, e.g. more than 10 times in reduce time and 5 times in total execution time. During our demonstration, we give the users the possibility to execute and compare job executions in FP-Hadoop and Hadoop. They can retrieve general information about the job and the tasks and a summary of the phases. They can also visually compare different configurations to explore the difference between the approaches.
Miguel Liroz-Gistau, Reza Akbarinia, Patrick Valduriez
Proc. VLDB Endow.2
2014 Entity resolution for probabilistic data
Naser Ayat, Reza Akbarinia, Hamideh Afsarmanesh, Patrick Valduriez
Inf. Sci.2
2013 Fast and Exact Mining of Probabilistic Data Streams
Reza Akbarinia, Florent Masseglia
ECML/PKDD (1)1
2013 Entity resolution for distributed probabilistic data
Naser Ayat, Reza Akbarinia, Hamideh Afsarmanesh, Patrick Valduriez
Distributed Parallel Databases2
2013 Efficient Evaluation of SUM Queries over Probabilistic Data
abstract
SUM queries are crucial for many applications that need to deal with uncertain data. In this paper, we are interested in the queries, called ALL_SUM, that return all possible sum values and their probabilities. In general, there is no efficient solution for the problem of evaluating ALL_SUM queries. But, for many practical applications, where aggregate values are small integers or real numbers with small precision, it is possible to develop efficient solutions. In this paper, based on a recursive approach, we propose a new solution for those applications. We implemented our solution and conducted an extensive experimental evaluation over synthetic and real-world data sets; the results show its effectiveness.
Reza Akbarinia, Patrick Valduriez, Guillaume Verger
IEEE Trans. Knowl. Data Eng.1
2012 Dynamic Workload-Based Partitioning for Large-Scale Databases
Miguel Liroz-Gistau, Reza Akbarinia, Esther Pacitti, Fábio Porto 0001, Patrick Valduriez
DEXA (2)2
2011 Efficient Early Top-k Query Processing in Overloaded P2P Systems
William Kokou Dedzoe, Philippe Lamarre, Reza Akbarinia, Patrick Valduriez
DEXA (1)3
2011 Best position algorithms for efficient top-k query processing
Reza Akbarinia, Esther Pacitti, Patrick Valduriez
Inf. Syst.1
2011 Building a peer-to-peer content distribution network with high performance, scalability and robustness
Manal El Dick, Esther Pacitti, Reza Akbarinia, Bettina Kemme
Inf. Syst.3
2009 DHTJoin: processing continuous join queries using DHT networks
Wenceslao Palma, Reza Akbarinia, Esther Pacitti, Patrick Valduriez
Distributed Parallel Databases2
2008 P2P logging and timestamping for reconciliation
abstract
In this paper, we address data reconciliation in peer-to-peer (P2P) collaborative applications. We propose P2P-LTR (Logging and Timestamping for Reconciliation) which provides P2P logging and timestamping services for P2P reconciliation over a distributed hash table (DHT). While updating at collaborating peers, updates are timestamped and stored in a highly available P2P log. During reconciliation, these updates are retrieved in total order to enforce eventual consistency. In this paper, we first give an overview of P2P-LTR with its model and its main procedures. We then present our prototype used to validate P2P-LTR. To demonstrate P2P-LTR, we propose several scenarios that test our solutions and measure performance. In particular, we demonstrate how P2P-LTR handles the dynamic behavior of peers with respect to the DHT.
Mounir Tlili, William Kokou Dedzoe, Esther Pacitti, Patrick Valduriez, Reza Akbarinia, Pascal Molli, Gérôme Canals, Stéphane Laurière
Proc. VLDB Endow.5
2007 Data currency in replicated DHTs
abstract
Distributed Hash Tables (DHTs) provide a scalable solution for data sharing in P2P systems. To ensure high data availability, DHTs typically rely on data replication, yet without data currency guarantees. Supporting data currency in replicated DHTs is difficult as it requires the ability to return a current replica despite peers leaving the network or concurrent updates. In this paper, we give a complete solution to this problem. We propose an Update Management Service (UMS) to deal with data availability and efficient retrieval of current replicas based on timestamping. For generating timestamps, we propose a Key-based Timestamping Service (KTS) which performs distributed timestamp generation using local counters. Through probabilistic analysis, we compute the expected number of replicas which UMS must retrieve for finding a current replica. Except for the cases where the availability of current replicas is very low, the expected number of retrieved replicas is typically small, e.g. if at least 35% of available replicas are current then the expected number of retrieved replicas is less than 3. We validated our solution through implementation and experimentation over a 64-node cluster and evaluated its scalability through simulation up to 10,000 peers using SimJava. The results show the effectiveness of our solution. They also show that our algorithm used in UMS achieves major performance gains, in terms of response time and communication cost, compared with a baseline algorithm.
Reza Akbarinia, Esther Pacitti, Patrick Valduriez
SIGMOD Conference1
2007 Best Position Algorithms for Top-k Queries
Reza Akbarinia, Esther Pacitti, Patrick Valduriez
VLDB1
2006 Reducing network traffic in unstructured P2P systems using Top-k queries
Reza Akbarinia, Esther Pacitti, Patrick Valduriez
Distributed Parallel Databases1