VLDB 2026 Research / reviewers in the wild / expert
Isam Mashhour Aljawarneh
dblp:205/5499 · also Isam Mashhour Al Jawarneh
· DBLP profile ↗
11ranked-venue papers
9as first author
6since 2021 · last 2025
0000-0002-4796-2181ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 8 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Air Quality Forecasting using Time-Series Interpolation with Simple Moving Average and Deep Learning-Based ModelsabstractAir pollution, particularly fine particulate matter (PM2.5), poses significant environmental and public health risks in urban areas. In New York City (NYC), the dynamic nature of air pollution introduces both temporal and spatial complexities, making accurate forecasting a challenging task especially since the data is hyperlocal. Traditional machine learning and statistical models often fail to capture these complexities effectively, leading to poor predictive performance. Many existing deep learning models struggle with spatial variability and temporal inconsistencies, limiting their ability to provide robust PM2.5 predictions tailored to NYC’s unique urban landscape. To address these challenges, we propose an enhanced PM2.5 forecasting framework that integrates geohash encoding and geospatial joins to augment the dataset with additional geospatial features to make the models spatially aware and leverage linear interpolation with Simple Moving Average (SMA) to ensure temporal consistency. We then train 3 CNN-LSTM hybrid models inspired by ResNet, InceptionNet and EfficientNet and compare their performance to a Temporal Convolutional Network (TCN). CNNs excel at feature extraction, while LSTMs specialize in modeling temporal dependencies, making them well-suited for time-series forecasting. Additionally, TCNs provide an alternative temporal modeling approach with parallel processing advantages. Our approach enhances the ability to model both spatial heterogeneity and temporal dependencies, leading to more accurate and dependable PM2.5 forecasts for hyperlocal NYC. Experimental results demonstrate that the TCN model outperforms the other models, offering improved prediction accuracy and robustness against spatial and temporal variations. This study highlights the importance of integrating geospatial encoding and advanced deep learning architectures for air quality forecasting, paving the way for more effective urban pollution management strategies. Madyan Bagosher, Domenico Scotece, Isam Mashhour Aljawarneh |
ISCC | 3 |
| 2025 | LSTM based Method for Forecasting Hyperlocal Air Quality in Metropolitan CitiesabstractWith the advancement of industrialization, air pollution has become a significant concern. This study presents a predictive analysis of PM2.5 in New York City using Long ShortTerm Memory (LSTM) neural networks to forecast pollution concentration levels. The key challenge is that the LSTM model lacks the spatial awareness needed for inference on a hyperlocal street level. To overcome this limitation, we enrich the dataset with additional geospatial features using geohash referencing and a geospatial join operation to capture PM2.5 patterns using features with varying spatial granularity. This enables the model to incorporate spatial context, shifting the problem from purely temporal to spatial temporal. We evaluated the models using several statistical metrics, including Mean Absolute Error (MAE), Mean Squared Error (MSE), and Root Mean Squared Error (RMSE), for the LSTM model, other machine learning (ML) models and networks, such as RNN, MLP, SVR, and LR. A comparison of the results indicates that the LSTM model with a geohash precision of 9 outperformed other models in predicting pollution levels. Madyan Bagosher, Domenico Scotece, Isam Mashhour Aljawarneh |
ISCC | 3 |
| 2024 | SpatialSSJP: QoS-Aware Adaptive Approximate Stream-Static Spatial Join ProcessorabstractThe widespread adoption of Internet of Things (IoT) motivated the emergence of mixed workload scenarios in smart cities, where fast arriving geo-referenced massive amounts data streams need to be joined with archive tables, at scale. This aims at enriching streams with descriptive attributes that enable deeper insightful analytics. More applications are now relying on finding, in real-time, to which geographical region each data streaming spatially-tagged tuple belongs. This problem requires a computationally intensive stream-static join operation, where one side of join is a dynamic stream while the other is a disk-resident static table. Even with emergence of some libraries that solve this problem in static-static fashion, their adoption for live scenarios is challenging because join operations are expensive in real-time. In addition, the time-varying nature of fluctuation and skewness in the geospatial data loads arriving online calls for an approximate solution that can trade-off QoS constraints in a way which ensures that the system survives sudden spikes in data loads. In this paper, we present SpatialSSJP, an adaptive spatial-aware approximate query processing system that specifically focuses on stream-static joins in a way that guarantees achieving an agreed set of Quality-of-Service goals and maintains geo-statistics of stateful online aggregations over stream-static join results. SpatialSSJP employs a state-of-art stratified-like sampling design to select well-balanced representative geospatial data stream samples and serve them to a stream-static geospatial join operator downstream. We implemented a prototype atop Spark Structured Streaming. Our extensive evaluations on big real datasets show that our system can survive and mitigate harsh join workloads and outperform state-of-art baselines by significant magnitudes, without risking rigorous error bounds in terms of the accuracy of the output results. SpatialSSJP achieves a relative accuracy gain against plain Spark joins of approximately 10% in worst cases but reaching up to 50% in best case scenarios. Isam Mashhour Aljawarneh, Paolo Bellavista, Antonio Corradi, Luca Foschini 0001, Rebecca Montanari |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | Efficient Geospatial Analytics on Time Series Big DataabstractIn smart city advanced analytical scenarios, tremendous amounts of georeferenced big time series data arrive continuously to time series databases, requiring the shared analytics on both geospatial and time dimensions. Mostly, the focus has been given to optimizing the storage and processing of each workload alone, either geospatial or time dimensions. To close this gap, in this paper, we have designed a pyramid-like indexing scheme that we term as geoTSI (short for geo time series index) which twists two dimensionality reduction geospatial encoding methods (geohash and S2) sequentially with a time series index to efficiently enable such mixed workload scenarios. This method enables geospatial and time indexes to collaborate synergistically in an aim to reduce the time required for accessing the disk and retrieving the time series data that comprises the answer for the mixed workload query. We show how our indexing scheme can be efficiently exploited to run a hybrid geospatial proximity query on time series data. Also, we evaluate our index on real-world georeferenced time series data, where we obtain, on average, a significant 34 % reduction in the query running time by applying our method against the baseline. Isam Mashhour Aljawarneh, Paolo Bellavista, Antonio Corradi, Luca Foschini 0001, Rebecca Montanari |
ICC | 1 |
| 2021 | Context Incorporation Techniques for Social Recommender SystemsabstractThe problem of information overloading is prevalent in recommendations websites and social networks. Users seek relevant recommendations from like-minded connections. User-item interactions (i.e., ratings) are prevalent in recommendation websites such as Netflix, whereas user-user connections are the interaction sought in social websites such as Twitter. Social recommender systems seek to generate recommendations for users based on similar preferences of their close friends. Because social networks do not normally contain user-item interactions, social recommender systems are typically hybridized with other recommenders (e.g., website recommenders such as Netflix) that provide such interaction. However, current systems are unaware of the user’s additional contextual information when coupled with social counterparts. In this paper, we propose a context-aware deep learning-based recommender system, US-NCF, in support for social recommender systems. Our experiments show US-NCF outperforms state-of-art counterparts. Isam Mashhour Aljawarneh, Paolo Bellavista, Antonio Corradi, Luca Foschini 0001, Rebecca Montanari |
ICC | 1 |
| 2021 | Efficient QoS-Aware Spatial Join Processing for Scalable NoSQL Storage FrameworksabstractCurrent cloud-enabled NoSQL database frameworks support flexible and scalable storage of huge amounts of data arriving through various and often heterogeneous channels. However, they do not natively provide optimised processing of spatial data, thus making it more difficult to perform accurate data analytics needed in many smart city application scenarios. To improve the performance of spatial data computation in the NoSQL MongoDB storage framework, this article proposes a novel data partitioning method based on dimensionality reduction. The underlying key idea is to reduce a spatial data representation from multi to single dimensionality, by still maintaining its geometrical meaning and by employing a specific geo-encoding scheme, i.e., a geohash string. In particular, the geohash string is used as a sharding key in order to store geometrically-nearby objects into the same chunks (and consequently into the same shard). In addition, as a distinctive feature, we have extended the MongoDB framework with a custom spatial QoS-aware optimizer that exploits our novel partitioning scheme to support two, typically expensive, types of spatial queries with QoS guarantees. Those queries are containment (and consequently top-N) and proximity. The paper also contributes to the existing literature with extensive experimental results about the performance of both our partitioning method and query optimizer; the reported results show that our solutions outperform baselines by orders of magnitude. Isam Mashhour Aljawarneh, Paolo Bellavista, Antonio Corradi, Luca Foschini 0001, Rebecca Montanari |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2020 | Locality-Preserving Spatial Partitioning for Geo Big Data Analytics in Main Memory FrameworksabstractThe easily reachable IoT edge devices have caused the accumulation of vast amounts of geo-referenced data traces that can help in performing deep insightful analytics. Geospatial data in real geometries are normally clumped into batches and has strong autocorrelation properties which can be exploited in discovering interesting insights. Current plain Cloud computing frameworks are not attuned to the shape of data. Most importantly, data splitting is an important precursor in data parallelization mechanisms. Current systems mostly focus on general data workloads, thus are giving attention mostly to load balancing while splitting the data to Cloud computing resources. However, many benefits can be reaped by being attuned to the spatial characteristics while distributing the data, thus striking a plausible balance between load balancing and spatial data locality preservation normally leads to achieving better time-based QoS goals, which then leads to an optimized provisioning of Cloud computing resources. In this paper, we have designed a spatial batch processing engine that comprises a custom spatial data locality aware partitioning method for disseminating spatial data loads in Cloud computing clusters. We have also extended a state-of-art benchmark density-based clustering method that is known as DBSCAN-MR and implemented a standard compliant prototype on top of a best-in-breed de facto Cloud-based main memory processing framework, Apache Spark. Our results show that our partitioning method with the associated spatial query optimizers can achieve gains that significantly outperform baselines. Isam Mashhour Aljawarneh, Paolo Bellavista, Antonio Corradi, Luca Foschini 0001, Rebecca Montanari |
GLOBECOM | 1 |
| 2019 | Spatial-Aware Approximate Big Data Stream ProcessingabstractThe widespread adoption of ubiquitous IoT edge devices and modern telemetry spewing out unprecedented avalanches of spatially-tagged datasets that if could interactively be explored would offer deep insights into interesting natural phenomena, which might remain otherwise illusive. Online application of spatial queries is expensive, a problem that is further inflated by the fact that we, more than often, do not have access to a full dataset population in non- stationary settings. As a way of coping up, sampling stands out as a natural solution for approximating estimators such as averages and totals of some interesting correlated parameters. In any sampling design, representativeness remains the main issue upon which a method is regarded good or bad. In a loose way, in a spatial context, this means fairly sampling quantities in a way that preserves spatial characteristics so as to provide more accurate approximates for spatial query responses. Current big data management systems either do not offer over-the-counter spatial-aware online sampling solutions or, at best, rely on randomness, which causes too many imponderables for an overall estimation. We herein have designed a QoS- spatial-aware online sampling method that outperforms vanilla baselines by statically significant magnitudes. Our method sits atop Apache Spark Structured Streaming's codebase and have been tested against a benchmark that is consisting of millions-records of spatially- augmented dataset. Isam Mashhour Aljawarneh, Paolo Bellavista, Luca Foschini 0001, Rebecca Montanari |
GLOBECOM | 1 |
| 2019 | Container Orchestration Engines: A Thorough Functional and Performance ComparisonabstractIn the last decade, novel software architectural patterns, such as microservices, have emerged to improve application modularity and to streamline their development, testing, scaling, and component replacement. To support these new trends, new practices as DevOps methodologies and tools, promoting better cooperation between software development and operations teams, have emerged to support automation and monitoring throughout the whole software construction lifecycle. That affected positively several IT companies, but also helped the transition to the softwarization of complex telco infrastructures in the last years. Container-based technologies played a crucial role by enabling microservice fast deployment and their scalability at low overhead; however, modern container-based applications may easily consist of hundreds of microservices services with complex interdependencies and call for advanced orchestration capabilities. While there are several emerging container orchestration engines, such as Docker Swarm, Kubernetes, Apache Mesos, and Cattle, a thorough functional and performance assessment to help IT managers in the selection of the most appropriate orchestration solution is still missing. This paper aims to fill that gap. Collected experimental results show that Kubernetes outperforms its counterparts for very complex application deployments, while other engines can be a better choice for simpler deployments. Isam Mashhour Aljawarneh, Paolo Bellavista, Filippo Bosi, Luca Foschini 0001, Giuseppe Martuscelli, Rebecca Montanari, Amedeo Palopoli |
ICC | 1 |
| 2018 | Cost-Effective Strategies for Provisioning NoSQL Storage Services in Support for Industry 4.0abstractThe advancement of networking and sensor-enabled devices have motivated the emergence of unprecedented initiatives, including Industry 4.0 and smart cities. Those are entwined in a way that makes their operation duly interconnected. Industry 4.0 will sooner become the biggest consumer of smart city big data. That data is geo-referenced, and its storage and processing need spatial-awareness, which is currently absent within the constellation of biggest big data management players of the market. We aim to fill this gap by providing spatial-aware big data management strategies in support for Industry 4.0 main principles. Our experimental results show that our strategies outperform those of state-of-the-art by orders of magnitude. Isam Mashhour Aljawarneh, Paolo Bellavista, Francesco Casimiro, Antonio Corradi, Luca Foschini 0001 |
ISCC | 1 |
| 2017 | Efficient spark-based framework for big geospatial data query processing and analysisabstractThe exponential amount of geospatial data that has been accumulated in an accelerated pace has inevitably motivated the scientific community to examine novel parallel technologies for tuning the performance of spatial queries. Managing spatial data for an optimized query performance is particularly a challenging task. This is due to the growing complexity of geometric computations involved in querying spatial data, where traditional systems failed to beneficially expand. However, the use of large-scale and parallel-based computing infrastructures based on cost-effective commodity clusters and cloud computing environments introduces new management challenges to avoid bottlenecks such as overloading scarce computing resources, which may be caused by an unbalanced loading of parallel tasks. In this paper, we aim to fill those gaps by introducing a generic framework for optimizing the performance of big spatial data queries on top of Apache Spark. Our framework also supports advanced management functions including a unique self-adaptable load-balancing service to self-tune framework execution. Our experimental evaluation shows that our framework is scalable and efficient for querying massive amounts of real spatial datasets. Isam Mashhour Aljawarneh, Paolo Bellavista, Antonio Corradi, Rebecca Montanari, Luca Foschini 0001, Andrea Zanotti |
ISCC | 1 |