Salman Ahmed Shaikh

dblp:50/11145 · DBLP profile ↗
← Back
21ranked-venue papers
13as first author
8since 2021 · last 2026
0000-0002-2204-8561ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 18 · 11 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Design Space of Iterative Graph Stream Processing on DAG Constrained Systems
Komal Mariam, Salman Ahmed Shaikh, Hiroyuki Kitagawa, Akiyoshi Matono
DEXA (1)2
2025 LPStream: Fine-grained Lazy Provenance for Stream Processing
abstract
Stream processing enables real-time data analysis. Recent stream processing engines (SPEs) execute stream processing in a distributed manner for real-time analysis of massive amounts of data produced by IoT devices and sensors. It has been widely adopted in various applications that support critical decision making. To explain the results of stream processing, ensuring provenance is indispensable. Provenance clarifies the relationship between input data and output data in the processing. With provenance, we can understand what input data contributed to the output. Existing frameworks for providing provenance for stream processing generate provenance or additional information to construct provenance at runtime. However, these approaches impose substantial overhead in ordinary stream processing. In this paper, we propose a new framework, named LPStream, for fine-grained lazy provenance. LPStream is the first framework to support lazy provenance for stream processing. In the ordinary execution mode, LPStream executes stream processing with checkpointing but without provenance generation. If provenance is necessary for some target output tuples, it replays the processing from an appropriate checkpoint and generates the provenance for the target tuple. We explain the design and implementation of LPStream and evaluate its performance by comparing LPStream with stream processing without provenance and with eager provenance. The experimental results demonstrate the effectiveness of our proposal.
Masaya Yamada, Hiroyuki Kitagawa, Salman Ahmed Shaikh, Toshiyuki Amagasa, Akiyoshi Matono
Proc. ACM Manag. Data3
2024 Distributed, Continuous and Real-time Trajectory Similarity Search
abstract
Moving objects’ trajectory data is becoming increasingly available with the omnipresent GPS-equipped devices, e.g., smart phones, smart watches, etc. Identification of similar trajectories is a fundamental requirement of many real-world applications, for instance, car pooling, road planning, epidemic contact tracing. There exist a number of studies focusing on trajectory similarity search, where given a query trajectory, similar trajectories are searched from a trajectory store. GPS devices generate trajectories as a continuous data stream, highlighting the critical need for real-time and continuous trajectory similarity search. Thus, given a trajectories stream ($S_{\Gamma}$) and a trajectory store (R), we propose a distributed, continuous and real-time similarity search between $S_{\Gamma}$ and R. The proposed method employs an index-based strategy to facilitate effective similarity search, with the added capability of execution in distributed environments to ensure scalable processing. A comprehensive empirical study is provided to demonstrate the efficiency and pruning capability of the proposed approach.
Salman Ahmed Shaikh, Hiroyuki Kitagawa, Akiyoshi Matono
ISPDC1
2023 Efficient Missing Value Imputation by Maximum Distance Likelihood
abstract
Predicting missing attribute values in data is extremely important in improving the accuracy in many applications. Existing algorithms ignore the difference between the records used for learning and predicting. The accuracy is not good enough and can be further improved. This paper proposes two solutions: (1) Maximization-based approach (MP) and (2) Distance-ratio-based approach (DP). MP and DP ensure that the incomplete records with the missed values are similar to the records used to learn the parameters as much as possible. MP and DP learn all possible parameters not only from the k nearest neighboring set (k-NN) but from the k-Sets, which are all possible combinations of k complete records. The parameters learnt from the records that are most similar to the repaired candidates of the incomplete records are chosen. Experimentally, MP and DP significantly outperform the existing approaches.
Savong Bou, Toshiyuki Amagasa, Hiroyuki Kitagawa, Salman Ahmed Shaikh, Akiyoshi Matono
IEEE Big Data4
2023 TraPM: A Framework for Online Pattern Matching Over Trajectory Streams
Rina Trisminingsih, Salman Ahmed Shaikh, Toshiyuki Amagasa, Hiroyuki Kitagawa, Akiyoshi Matono
iiWAS2
2022 TStream: a framework for real-time and scalable trajectory stream processing and analysis
abstract
Recent advances in location-aware devices have resulted in an exponential increase in the trajectory data streams. A number of applications require real-time processing and analysis of massive moving objects' trajectories. For instance, route guidance in emergency evacuation, patients tracking, etc. Existing scalable trajectory management systems lack support for real-time processing, while the real-time systems do not natively support spatial trajectory processing. This work presents TStream, a real-time and scalable trajectory stream processing and analysis framework. TStream utilizes grid index to support efficient processing of continuous range, kNN and join queries.
Salman Ahmed Shaikh, Hiroyuki Kitagawa, Akiyoshi Matono, Kyoung-Sook Kim 0001
SIGSPATIAL/GIS1
2022 PR-MVI: Efficient Missing Value Imputation over Data Streams by Distance Likelihood
Savong Bou, Toshiyuki Amagasa, Hiroyuki Kitagawa, Salman Ahmed Shaikh, Akiyoshi Matono
iiWAS4
2022 Streaming Augmented Lineage: Traceability of Complex Stream Data Analysis
Masaya Yamada, Hiroyuki Kitagawa, Salman Ahmed Shaikh, Toshiyuki Amagasa, Akiyoshi Matono
iiWAS3
2020 GeoFlink: A Distributed and Scalable Framework for the Real-time Processing of Spatial Streams
abstract
Apache Flink is an open-source system for scalable processing of batch and streaming data. Flink does not natively support efficient processing of spatial data streams, which is a requirement of many applications dealing with spatial data. Besides Flink, other scalable spatial data processing platforms including GeoSpark, Spatial Hadoop, etc. do not support streaming workloads and can only handle static/batch workloads. To fill this gap, we present GeoFlink, which extends Apache Flink to support spatial data types, indexes and continuous queries over spatial data streams. To enable efficient processing of spatial continuous queries and for the effective data distribution across Flink cluster nodes, a gird-based index is introduced. GeoFlink currently supports spatial range, spatial kNN and spatial join queries on point data type. An experimental study on real spatial data streams shows that GeoFlink achieves significantly higher query throughput than ordinary Flink processing.
Salman Ahmed Shaikh, Komal Mariam, Hiroyuki Kitagawa, Kyoung-Sook Kim 0001
CIKM1
2019 A Robust and Scalable Pipeline for the Real-time Processing and Analysis of Massive 3D Spatial Streams
abstract
With the increase in the use of 3D scanner to sample the earth surface, there is a surge in the availability of 3D spatial data. 3D spatial data contains a wealth of information and can be of potential use if integrated, processed and analyzed in real-time. The 3D spatial data is generated as continuous data stream, however due to its size, velocity and inherent noise, it is processed offline. Many applications require real-time processing and analysis of spatial stream, for-instance, forest fire management, real-time road traffic analysis, disaster engulfed areas monitoring, etc., however they suffer from slow offline processing of traditional systems. This paper presents and demonstrates a robust and scalable pipeline for the real-time processing and analysis of 3D spatial streams. An experimental evaluation is also presented to prove the effectiveness of the proposed framework.
Salman Ahmed Shaikh, Jun Lee 0002, Akiyoshi Matono, Kyoung-Sook Kim 0001
iiWAS1
2019 Smart scheme: an efficient query execution scheme for event-driven stream processing
Salman Ahmed Shaikh, Yousuke Watanabe, Hiroyuki Kitagawa
Knowl. Inf. Syst.1
2017 Smart distributed query execution over data streams
abstract
Current era is witnessing a tremendous growth in the volume of data that is being generated in the form of data streams by the omnipresent sensors, micro-blogs, e-businesses, etc. Many organizations require on-line processing of their data for real time analysis and actionable alerts. It is not possible to process such voluminous and velocious data in real time using the traditional centralized stream processing engines. Hence distributed stream processing has emerged to facilitate such large scale real time processing. In this work we present a smart distributed event-driven stream processing approach. In contrast to the ordinary stream processing, event-driven stream processing generates query results on the occurrence of specified events only. In the basic event-driven stream processing, even when no event is raised input stream tuples are continuously processed by query operators, though they do not generate any query result. This results in increased system load and wastage of system resources. Whereas in the smart event-driven stream processing scheme, incoming tuples are processed in the presence of events only resulting in reduced system load. The proposed smart distributed event-driven stream processing utilizes the concept of smart query execution to distribute the data stream among the distributed worker nodes in the presence of events only; while in the absence of events no data is distributed as it can not generate query output. This smart data distribution can significantly reduce the network traffic in the absence of events and ultimately results in improved overall system throughput. Detailed experiments are performed to prove the effectiveness of the proposed framework.
Salman Ahmed Shaikh, Hiroyuki Kitagawa
IEEE BigData1
2017 StreamingCube: A Unified Framework for Stream Processing and OLAP Analysis
abstract
In most streaming applications, the data streams need to be analyzed continuously to make instant decisions exploiting latest information. Often data streams are multidimensional and are at the low-level of abstraction, whereas analysts are interested in multi-level interactive analysis of data streams across several dimensions. On-line analytical processing (OLAP) is a proven technique for such analysis of static data and has also been studied by some researchers for data streams. Traditionally this is achieved by coupling a stream processing engine with an OLAP engine. We believe that coupling multiple systems is not an efficient solutions as it results in lower performance (due to the transfer of data between multiple systems), resource wastage (due to replication of data for each coupled system) and increased complexity and maintenance cost. To this end, we present StreamingCube, a unified framework for data stream processing and its interactive OLAP analysis. The proposed framework possesses all the essential operators to process data streams and introduces a new operator, cubify, to maintain OLAP lattice nodes (materialized views) incrementally. The novelty of the introduced cubify operator lies in the incremental maintenance of the materialized views. To demonstrate StreamingCube, a web-based GUI has been developed which enables users to register continuous queries (CQs). Once a CQ has been registered, users can perform different OLAP operations through the GUI for the interactive analysis. The results of the OLAP queries/operations are displayed in the form of tables and graphs.
Salman Ahmed Shaikh, Hiroyuki Kitagawa
CIKM1
2017 Approximate OLAP on Sustained Data Streams
Salman Ahmed Shaikh, Hiroyuki Kitagawa
DASFAA (2)1
2017 SOLA: Stream OLAP-based Analytical Framework for Roadway Maintenance
abstract
Maintaining infrastructures (e.g., roadway) is a critical issue for local governments. Data from physical devices and reports from citizens through social networks are helpful to observe conditions of infrastructures. This paper proposes a framework called SOLA for integrating and analysing data from multiple sources including streaming data and static data for roadway management. The framework integrates data from multiple sources in the way of stream OLAP architecture, and analyses the integrated data in terms of OLAP analysis. This paper applies the framework to support roadway managements of local governments, and develops the application called SOLAR. SOLAR aims at providing historical views of roadway patrols as well as roadway statuses for assisting in determining roadway patrolling schedules. The real-world use case on a city exhibits the applicability of SOLAR with positive feedbacks from city officers. SOLA is a promising framework for big data analysis and smart city applications, as the number, amount, and speed of generating data increase in the era of big data and smart city.
Takahiro Komamizu, Toshiyuki Amagasa, Salman Ahmed Shaikh, Hiroaki Shiokawa, Hiroyuki Kitagawa
MEDES3
2016 Incremental Continuous Query Processing over Streams and Relations with Isolation Guarantees
Salman Ahmed Shaikh, Dong Chao, Kazuya Nishimura, Hiroyuki Kitagawa
DEXA (1)1
2015 An architecture for stream OLAP exploiting SPE and OLAP engine
abstract
Explosive increase of real-time data sources, so-called "data streams" (or just "steams") and increasing demands for real-time analysis over streams give rise to realtime analysis over streams. However, developing tailor-made systems for such applications is not always desirable due to high developing costs and long developing periods. To cope with this problem, this paper proposes a novel architecture for online analytical processing (OLAP) over streams exploiting off-the-shelf stream processing engine (SPE) combined with OLAP engine. It allows users to perform OLAP analysis over streams for the latest time period, called Interval of Interest (Iol). The system in the meantime processes multiple continuous query language (CQL) queries corresponding to different aggregation levels in cube lattice. To cover arbitrary aggregation levels using limited system's memory, we propose to partially deploy CQL queries for those with higher reference frequencies, whereas the results are dynamically calculated using existing aggregation results with the help of OLAP engine. For optimal CQL query deployment, we propose a cost-based optimization method that maximizes the performance. The experimental results show that the proposed architecture is feasible enough to realize stream OLAP by combining an SPE and an OLAP engine. Also, the proposed system significantly outperforms other comparative methods by generating optimized query deployment plans.
Kousuke Nakabasami, Toshiyuki Amagasa, Salman Ahmed Shaikh, Franck Gass, Hiroyuki Kitagawa
IEEE BigData3
2014 MOOD: Moving Objects Outlier Detection
Salman Ahmed Shaikh, Hiroyuki Kitagawa
APWeb1
2014 Efficient distance-based outlier detection on uncertain datasets of Gaussian distribution
Salman Ahmed Shaikh, Hiroyuki Kitagawa
World Wide Web1
2013 Fast Top-k Distance-Based Outlier Detection on Uncertain Data
Salman Ahmed Shaikh, Hiroyuki Kitagawa
WAIM1
2012 Distance-Based Outlier Detection on Uncertain Data of Gaussian Distribution
Salman Ahmed Shaikh, Hiroyuki Kitagawa
APWeb1