Hannaneh Najdataei

dblp:209/9626 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data stream processing · 46% Data mining · 30% Distributed and cloud data management · 23%
Computer networks
1 paper
Vehicular, aerial and satellite networks · 44% Internet of things and sensor networks · 44% Edge and fog computing · 13%
Software engineering, system software, and programming languages
1 paper
Concurrent programming · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data stream processing › stream processing systems
elastic stream processing
0.612022
STRETCH: Virtual Shared-Nothing Parallelism for Scalable and Elastic Stream Processing · IEEE Trans. Parallel Distributed Syst. 2022
Data stream processing
parallel stream processing
0.612022
STRETCH: Virtual Shared-Nothing Parallelism for Scalable and Elastic Stream Processing · IEEE Trans. Parallel Distributed Syst. 2022
Distributed and cloud data management › parallel data processing
shared-nothing parallelism
0.612022
STRETCH: Virtual Shared-Nothing Parallelism for Scalable and Elastic Stream Processing · IEEE Trans. Parallel Distributed Syst. 2022
Data mining › clustering › online clustering
data stream clustering
0.412019
DRIVEN: a Framework for Efficient Data Retrieval and Clustering in Vehicular Networks · ICDE 2019
Data mining › clustering
distance-based clustering
0.412019
DRIVEN: a Framework for Efficient Data Retrieval and Clustering in Vehicular Networks · ICDE 2019
Internet of things and sensor networks › wireless sensor network
data collection
0.412019
DRIVEN: a Framework for Efficient Data Retrieval and Clustering in Vehicular Networks · ICDE 2019
Vehicular, aerial and satellite networks
vehicular networks
0.412019
DRIVEN: a Framework for Efficient Data Retrieval and Clustering in Vehicular Networks · ICDE 2019
Concurrent programming
shared memory
0.212022
STRETCH: Virtual Shared-Nothing Parallelism for Scalable and Elastic Stream Processing · IEEE Trans. Parallel Distributed Syst. 2022
Edge and fog computing › mobile edge computing
vehicular edge computing
0.112019
DRIVEN: a Framework for Efficient Data Retrieval and Clustering in Vehicular Networks · ICDE 2019

Methods — techniques the papers use, named apart from their topics

virtual shared-nothing parallelism · 1.1elasticity · 1.1streaming clustering · 0.8piecewise linear approximation · 0.8
YearPublicationVenuePosition
2022 STRETCH: Virtual Shared-Nothing Parallelism for Scalable and Elastic Stream Processing
abstract
Stream processing applications extract value from raw data through Directed Acyclic Graphs of data analysis tasks. Shared-nothing (SN) parallelism is the de-facto standard to scale stream processing applications. Given an application, SN parallelism ins9tantiates several copies of each analysis task, making each instance responsible for a dedicated portion of the overall analysis, and relies on dedicated queues to exchange data among connected instances. On the one hand, SN parallelism can scale the execution of applications both up and out since threads can run task instances within and across processes/nodes. On the other hand, its lack of sharing can cause unnecessary overheads and hinder the scaling up when threads operate on data that could be jointly accessed in shared memory. This trade-off motivated us in studying a way for stream processing applications to leverage shared memory and boost the scale up (before the scale out) while adhering to the widely-adopted and SN-based APIs for stream processing applications. We introduceSTRETCH, a framework that maximizes the scale up and offers instantaneous elastic reconfigurations (without state transfer) for stream processing applications. We propose the concept of Virtual Shared-Nothing (VSN) parallelism and elasticity and provide formal definitions and correctness proofs for the semantics of the analysis tasks supported bySTRETCH, showing they extend the ones found in common Stream Processing Engines. We also provide a fully implemented prototype and show thatSTRETCH's performance exceeds that of state-of-the-art frameworks such as Apache Flink and offers, to the best of our knowledge, unprecedented ultra-fast reconfigurations, taking less than 40 ms even when provisioning tens of new task instances.
Vincenzo Gulisano, Hannaneh Najdataei, Yiannis Nikolakopoulos, Alessandro Vittorio Papadopoulos, Marina Papatriantafilou, Philippas Tsigas
IEEE Trans. Parallel Distributed Syst.2
2020 DRIVEN: A framework for efficient Data Retrieval and clustering in Vehicular Networks
Bastian Havers, Romaric Duvignau, Hannaneh Najdataei, Vincenzo Gulisano, Marina Papatriantafilou, Ashok Chaitanya Koppisetty
Future Gener. Comput. Syst.3
2019 Adaptive Stream-based Shifting Bottleneck Detection in IoT-based Computing Architectures
abstract
Cloud computing is revolutionizing the backbone of data analysis applications, including industrial ones. One of its main pillars is the separation of the logic with which data is accessed (e.g., to study the efficiency of a manufacturing system) from the actual hardware (e.g., server) that maintains and analyses the data. Large distributed cyber-physical systems enabled by, among other technologies, the Internet of Things (IoT), made nonetheless clear that “what to do” with the data and “where to do it” are not disjoint problems; i.e., cloud computing on its own is not enough. Fog and edge computing have emerged as complementary options, to distribute the analysis, helping with challenges by means of close-to-the-source data analysis.We show for a key problem for industrial processes, that of shifting bottleneck detection, how to take advantage of such multi-tier computing architectures, to perform continuous and configurable analysis of data from Manufacturing Execution Systems. We propose a processing framework, STRATUM, and an algorithm, AMBLE, for continuous, data stream processing. STRATUM seamlessly distributes and parallelizes the processing across the tiers and AMBLE guarantees consistent analysis in spite of timing fluctuations, which are commonly introduced due to e.g. the communication system; it also achieves efficiency through appropriate data structures for in-memory processing. The experimental study on a real-world dataset, taken from a production line over two years and including 8.5 million entries, shows the benefits of the proposed solution in enabling configurable and efficient analysis.
Hannaneh Najdataei, Mukund Subramaniyan, Vincenzo Gulisano, Anders Skoogh, Marina Papatriantafilou
ETFA1
2019 DRIVEN: a Framework for Efficient Data Retrieval and Clustering in Vehicular Networks
abstract
Applications for adaptive (sometimes also called smart) Cyber-Physical Systems are blossoming thanks to the large volumes of data, sensed in a continuous fashion, in large distributed systems. The benefits of these applications come nonetheless with a price: the need for jointly addressing challenges in efficient data communication and analysis (among others). The goal of the DRIVEN framework, presented here, is to address these challenges for a data gathering and distance-based clustering tool in the context of vehicular networks. Because of the limited communication bandwidth (compared to the volume of sensed data) of vehicular networks and the monetary costs of data transmission, the intuition behind DRIVEN is to avoid gathering the data to be clustered in a raw format from each vehicle, but rather to allow for a streaming-based error-bounded approximation, through Piecewise Linear Approximation, to compress the volumes of data to be gathered. At the same time, rather than relying on a batch-based clustering algorithm that requires all the data to be first gathered (and then clustered), DRIVEN relies on and extends a streaming-based clustering algorithm that leverages the inherent ordering of the spatial and temporal data being collected, to perform the clustering in an online fashion, while data is being retrieved. As we show, based on our prototype implementation using Apache Flink and our evaluation with real-world data such as GPS and LiDAR, the accuracy loss for the clustering performed on the reconstructed data can be small, even when the raw data is compressed to 10-35% of its original size, and the transferring of data itself can be completed in up to one-tenth of the duration observed when gathering raw data.
Bastian Havers, Romaric Duvignau, Hannaneh Najdataei, Vincenzo Gulisano, Ashok Chaitanya Koppisetty, Marina Papatriantafilou
ICDE3
2018 Continuous and Parallel LiDAR Point-Cloud Clustering
abstract
In distributed digitalized environments in the context of the Internet of Things, we often need to do an analysis of big data originating at high rate-sensors at the edge of the infrastructure. A characteristic example is the light detection and ranging (LiDAR) technology, that allows sensing surrounding objects with fine-grained resolution in large areas. Their data (known as point clouds), generated continuously at very high rates, through appropriate analysis can provide information to support automated functionality in distributed cyber-physical? systems; clustering of point clouds is a key problem to extract this type of information. Methods for solving the problem in a continuous fashion can facilitate improved processing in fog architectures, through enabling low-latency, efficient continuous and streaming processing of data close to the sources; moreover, parallelism is a key requirement to exploit a variety of computing architectures in this context. We proposeLisco, a single-pass continuous Euclidean-distance-based clustering of LiDAR point clouds, that maximizes the granularity of the data processing pipeline and thus shows the potential for data-and pipeline-parallelism. We further present its parallel version, P-Lisco, that is architecture-independent and exploits the parallelism revealed byLisco'salgorithmic approach. Besides their algorithmic analysis, we provide a thorough experimental evaluation on architectures representative of high-end servers and of resource-constrained embedded devices and highlight the multiplicative improvements and scalability benefits of the proposed algorithms compared to the baseline, using both real-world datasets as well as synthetic ones to fully explore a wide spectrum of stress-levels for the algorithms.
Hannaneh Najdataei, Yiannis Nikolakopoulos, Vincenzo Gulisano, Marina Papatriantafilou
ICDCS1