Md. Saiful Islam 0003

dblp:04/3572-3 · DBLP profile ↗
← Back
25ranked-venue papers in the field
8as first author
7since 2021 · last 2025
0000-0001-7181-5328ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 11 (6 first)Big Data, Cloud & Distributed Data Systems · 6Information Retrieval & Web Search · 4 (1 first)Data Mining & Knowledge Discovery · 2Business Process & Enterprise Data · 1 (1 first)Other / Interdisciplinary · 1
YearPublicationVenuePosition
2025 Evaluation of Large Language Models for Understanding Counterfactual Reasoning in Texts
S. I. M. Adnan, Abrar Hameem, Shikha Anirban, Md. Saiful Islam 0003, Md Musfique Anwar
IEEE Big Data4
2025 Adaptive Semi-Supervised Federated Learning with Selective Knowledge Distillation
Syeda Faiza Ahmed, Kaies Al Mahmud, Md. Saiful Islam 0003, Lafifa Jamal
IEEE Big Data3
2025 An Experimental Evaluation of Mixup Ensemble Meta-Classifiers for Improving GNN-Based Node Classification
Venkata Kadiyala, Md. Saiful Islam 0003, M. A. Hakim Newton
IEEE Big Data2
2025 An Integrated Deep Learning Framework for Trading Card Grading: Design, Deployment, and Evaluation
abstract
This paper presents an integrated deep learning framework that unifies fine-grained corner, edge, and surface evaluation into a coherent grading pipeline, enabling consistent, scalable, and explainable trading card grading for real-world deployment. The proposed system decomposes each card into multiple high-resolution regions—eight corners (front and back), edge patches, and surface segments—each evaluated by specialized deep learning models optimized for distinct defect types. Model outputs are systematically aggregated to produce an overall grade, ensuring both granularity and consistency in evaluation. The modular design supports real-time inference and seamless deployment via FastAPI, enabling scalable industrial adoption. We detail the architecture, deployment pipeline, and performance evaluation, and share key lessons from production integration. Tested on a large industrial dataset, the system achieves high accuracy and exhibits strong agreement with expert human graders. By delivering explainable, consistent, and costefficient grading, this work provides a practical, production-ready alternative to traditional manual grading, addressing a critical bottleneck in the collectible card industry.
Lutfun Nahar, Md. Saiful Islam 0003, Mohammad Awrangjeb, Rob Verhoeve
IEEE Big Data2
2025 Feature Drift-Guided Adaptive ML Retraining: An MLOps Approach for Big Data Analytics
Md Nahid Parves Shakil, Md. Saiful Islam 0003, Nasimul Noman, Marcella Papini
IEEE Big Data2
2024 Edge Grading in Trading Cards Using Transfer Learning: Methods, Experiments, and Evaluation
abstract
The trading card market is a dynamic industry where card values depend on meticulous grading of condition, authenticity, and quality. Traditionally, grading is performed manually by expert assessors, but this process is prone to subjectivity and inconsistency. Automated grading systems could ensure greater objectivity and uniformity in assessments. Key grading factors include centering, edges, corners, and surface quality. While our previous work focused on corner grading, this paper explores edge grading using CNN and transfer learning models such as DenseNet, ResNet, and VGG. By fine-tuning, ResNet50 achieved 93% accuracy on a dataset from our industry partner. To address uncertainty in grading, various calibration methods are employed, and a human-in-the-loop approach enhances robustness. A final scoring method provides an objective edge grading for each card.
Lutfun Nahar, Md. Saiful Islam 0003, Mohammad Awrangjeb, Rob Verhoeve
IEEE Big Data2
2023 Experimental Evaluation of Indexing Techniques for Shortest Distance Queries on Road Networks
abstract
Shortest distance calculation between two locations in road networks is an important problem and has many applications. This problem has been widely researched for over two decades. Several advanced algorithms have been developed since the last formal evaluation. This paper provides a comprehensive experimental evaluation of these state-of-the-art algorithms. Our evaluation provides several important insights on the advantage/disadvantages of these algorithms, and it enables us to recommend the most suitable algorithm for some application scenarios. We are able to confirm some previous experimental results and raise questions on some others. We also evaluate the effect of a simple path compression technique on these algorithms.
Shikha Anirban, Junhu Wang, Md. Saiful Islam 0003
ICDE3
2020 Rice Leaf Diseases Recognition Using Convolutional Neural Networks
Syed Mohammad Minhaz Hossain, Md. Monjur Morhsed Tanjil, Mohammed Abser Bin Ali, Mohammad Zihadul Islam, Md. Saiful Islam 0003, Sabrina Mobassirin, Iqbal H. Sarker, S. M. Riazul Islam
ADMA5
2020 A Hybrid Index for Distance Queries
Junhu Wang, Shikha Anirban, Toshiyuki Amagasa, Hiroaki Shiokawa, Zhiguo Gong, Md. Saiful Islam 0003
WISE (1)6
2020 Efficient processing of reverse nearest neighborhood queries in spatial databases
Md. Saiful Islam 0003, Bojie Shen, Can Wang 0004, David Taniar, Junhu Wang
Inf. Syst.1
2019 A Causality Driven Approach to Adverse Drug Reactions Detection in Tweets
Humayun Kayesh, Md. Saiful Islam 0003, Junhu Wang
ADMA2
2019 Multi-level Graph Compression for Fast Reachability Detection
Shikha Anirban, Junhu Wang, Md. Saiful Islam 0003
DASFAA (2)3
2019 Modular Decomposition-Based Graph Compression for Fast Reachability Detection
abstract
Fast reachability detection is one of the key problems in graph applications. Most of the existing works focus on creating an index and answering reachability based on that index. For these approaches, the index construction time and index size can become a concern for large graphs. More recently query-preserving graph compression has been proposed, and searching reachability over the compressed graph has been shown to be able to significantly improve query performance as well as reducing the index size. In this paper, we introduce a multilevel compression scheme for DAGs, which builds on existing compression schemes, but can further reduce the graph size for many real-world graphs. We propose an algorithm to answer reachability queries using the compressed graph. Extensive experiments with four existing state-of-the-art reachability algorithms and 12 real-world datasets demonstrate that our approach outperforms the existing methods. Experiments with synthetic datasets ensure the scalability of this approach. We also provide a discussion on possible compression for k-reachability.
Shikha Anirban, Junhu Wang, Md. Saiful Islam 0003
Data Sci. Eng.3
2019 XSnippets: Exploring semi-structured data via snippets
Mehdi Naseriparsa, Md. Saiful Islam 0003, Chengfei Liu, Lu Chen 0008
Data Knowl. Eng.2
2018 Capturing the Spatiotemporal Evolution in Road Traffic Networks
abstract
The urban road networks undergo frequent traffic congestions during the peak hours and around the city center. Capturing the spatiotemporal evolution of the congestion scenario in real-time in an urban-scale can aid in developing smart traffic management systems, and guiding commuters in making informed decision about route choice. The congestion scenario is often represented by a set of distinguishable network partitions that have a homogeneous level of congestion inside them but are heterogeneous to others. Due to the dynamic nature of traffic, these partitions evolve with time in terms of their structure and location. In this paper, we propose a comprehensive framework to capture the evolution by incrementally updating the partitions in an efficient manner using a two-layer approach. The physical layer maintains a set of small-sized road network building blocks in a fine granularity, and performs low-level computations to incrementally update them, whereas the logical layer performs high-level computations in order to serve as an interface to query the physical layer about the congested partitions in a coarse granularity. We also propose an in-memory index calledBinthat compactly stores the historical sets of building blocks in the main memory with no information loss, and facilitates their efficient retrieval. Our experimental results show that the proposed method is much efficient than the existing re-partitioning methods without significant sacrifice in accuracy. The proposedBinconsumes a minimum space with least redundancy at different time stamps.
Tarique Anwar, Chengfei Liu, Hai Le Vu 0001, Md. Saiful Islam 0003, Timos K. Sellis
IEEE Trans. Knowl. Data Eng.4
2017 Computing Influence of a Product through Uncertain Reverse Skyline
abstract
Understanding the influence of a product is crucially important for making informed business decisions. This paper introduces a new type of skyline queries, called uncertain reverse skyline, for measuring the influence of a probabilistic product in uncertain data settings. More specifically, given a dataset of probabilistic products P and a set of customers C, an uncertain reverse skyline of a probabilistic product q retrieves all customers c ∈ C which include q as one of their preferred products. We present efficient pruning ideas and techniques for processing the uncertain reverse skyline query of a probabilistic product using R-Tree data index. We also present an efficient parallel approach to compute the uncertain reverse skyline and influence score of a probabilistic product. Our approach significantly outperforms the baseline approach derived from the existing literature. The efficiency of our approach is demonstrated by conducting experiments with both real and synthetic datasets.
Md. Saiful Islam 0003, Wenny Rahayu, Chengfei Liu, Tarique Anwar, Bela Stantic
SSDBM1
2016 Q+Tree: An Efficient Quad Tree based Data Indexing for Parallelizing Dynamic and Reverse Skylines
abstract
Skyline queries play an important role in multi-criteria decision making applications of many areas. Given a dataset of objects, a skyline query retrieves data objects that are not dominated by any other data object in the dataset. Unlike standard skyline queries where the different aspects of data objects are compared directly, dynamic and reverse skyline queries adhere to the around-by semantics, which is realized by comparing the relative distances of the data objects w.r.t. a given query. Though, there are a number of works on parallelizing the standard skyline queries, only a few works are devoted to the parallel computation of dynamic and reverse skyline queries. This paper presents an efficient quad-tree based data indexing scheme, called Q+Tree, for parallelizing the computations of the dynamic and reverse skyline queries. We compare the performance of Q+Tree with an existing quad-tree based indexing scheme. We also present several optimization heuristics to improve the performance of both of the indexing schemes further. Experimentation with both real and synthetic datasets verifies the efficiency of the proposed indexing scheme and optimization heuristics.
Md. Saiful Islam 0003, Chengfei Liu, Wenny Rahayu, Tarique Anwar
CIKM1
2016 Tracking the Evolution of Congestion in Dynamic Urban Road Networks
abstract
The congestion scenario on a road network is often represented by a set of differently congested partitions having homogeneous level of congestion inside. Due to the changing traffic, these partitions evolve with time. In this paper, we propose a two-layer method to incrementally update the differently congested partitions from those at the previous time point in an efficient manner, and thus track their evolution. The physical layer performs low-level computations to incrementally update a set of small-sized road network building blocks, and the logical layer provides an interface to query the physical layer about the congested partitions. At each time point, the unstable road segments are identified and moved to their most suitable building blocks. Our experimental results on different datasets show that the proposed method is much efficient than the existing re-partitioning methods without significant sacrifice in accuracy.
Tarique Anwar, Chengfei Liu, Hai Le Vu 0001, Md. Saiful Islam 0003
CIKM4
2016 Efficient answering of why-not questions in similar graph matching
abstract
Graph data management and matching similar graphs are very important for many applications including bioinformatics, computer vision, VLSI design, bug localization, road networks, social and communication networking. Many graph indexing and similarity matching techniques have already been proposed for managing and querying graph data. In similar graph matching, a user is returned with the database graphs whose distances with the query graph are below a threshold. In such query settings, a user may not receive certain database graphs that are very similar to the query graph if the initial query graph is inappropriate/imperfect for the expected answer set. To exemplify this, consider a drug designer who is looking for chemical compounds that could be the target of her hypothetical drug before realizing it. In response to her query, the traditional search system may return the structures from the database that are most similar to the query graph. However, she may get surprised if some of the expected targets are missing in the answer set. She may then seek assistance from the system by asking “Is there other query graph that can match my expected answer set?”. The system may then modify her initial query graph to include the missing answers in the new answer set. Here, we study this kind of problem of answering why-not questions in similar graph matching for graph databases.
Md. Saiful Islam 0003, Chengfei Liu, Jianxin Li 0001
ICDE1
2016 Know your customer: computing k-most promising products for targeted marketing
Md. Saiful Islam 0003, Chengfei Liu
VLDB J.1
2015 RoadRank: Traffic Diffusion and Influence Estimation in Dynamic Urban Road Networks
abstract
With the rapidly growing population in urban areas, these days the urban road networks are expanding at a faster rate. The frequent movement of people on them leads to traffic congestions. These congestions originate from some crowded road segments, and diffuse towards other parts of the urban road networks creating further congestions. This behavior of road networks motivates the need to understand the influence of individual road segments on others in terms of congestion. In this work, we propose RoadRank, an algorithm to compute the influence scores of each road segment in an urban road network, and rank them based on their overall influence. It is an incremental algorithm that keeps on updating the influence scores with time, by feeding with the latest traffic data at each time point. The method starts with constructing a directed graph called influence graph, which is then used to iteratively compute the influence scores using probabilistic diffusion theory. We show promising preliminary experimental results on real SCATS traffic data of Melbourne.
Tarique Anwar, Chengfei Liu, Hai Le Vu 0001, Md. Saiful Islam 0003
CIKM4
2015 Efficient Answering of Why-Not Questions in Similar Graph Matching
abstract
Answeringwhy-notquestions in databases is promised to have wide application prospect in many areas and thereby, has attracted recent attention in the database research community. This paper addresses the problem of answering these so-calledwhy-notquestions in similar graph matching for graph databases. Given a set of answer graphs of an initial query graph$q$and a set of missing (why-not) graphs, we aim to modify$q$into a new query graph$q^*$such that the missing graphs are included in the new answer set of$q^*$. We present an approximate solution to address the above as the optimal solution is NP-hard to compute. In our approach, we first compute the bounded search space and the distance to be minimized for$q^*$. Then, we present a two-phase algorithm to find the new query$q^*$. In the first phase, we generate a set of candidate edges to be added/deleted into/from the initial query$q$within the bounded search space and in the second phase, we select a subset of candidate edges generated in the first phase to minimize the distance for$q^*$. We also demonstrate the effectiveness and efficiency of our approach by conducting extensive experiments on two real datasets.
Md. Saiful Islam 0003, Chengfei Liu, Jianxin Li 0001
IEEE Trans. Knowl. Data Eng.1
2014 Keyword-based correlated network computation over large social media
abstract
Recent years have witnessed an unprecedented proliferation of social media, e.g., millions of blog posts, micro-blog posts, and social networks on the Internet. This kind of social media data can be modeled in a large graph where nodes represent the entities and edges represent relationships between entities of the social media. Discovering keyword-based correlated networks of these large graphs is an important primitive in data analysis, from which users can pay more attention about their concerned information in the large graph. In this paper, we propose and define the problem of keyword-based correlated network computation over a massive graph. To do this, we first present a novel tree data structure that only maintains the shortest path of any two graph nodes, by which the massive graph can be equivalently transformed into a tree data structure for addressing our proposed problem. After that, we design efficient algorithms to build the transformed tree data structure from a graph offline and compute the γ-bounded keyword matched subgraphs based on the pre-built tree data structure on the fly. To further improve the efficiency, we propose weighted shingle-based approximation approaches to measure the correlation among a large number of γ-bounded keyword matched subgraphs. At last, we develop a merge-sort based approach to efficiently generate the correlated networks. Our extensive experiments demonstrate the efficiency of our algorithms on reducing time and space cost. The experimental results also justify the effectiveness of our method in discovering correlated networks from three real datasets.
Jianxin Li 0001, Chengfei Liu, Md. Saiful Islam 0003
ICDE3
2013 On answering why-not questions in reverse skyline queries
abstract
This paper aims at answering the so called why-not questions in reverse skyline queries. A reverse skyline query retrieves all data points whose dynamic skylines contain the query point. We outline the benefit and the semantics of answering why-not questions in reverse skyline queries. In connection with this, we show how to modify the why-not point and the query point to include the why-not point in the reverse skyline of the query point. We then show, how a query point can be positioned safely anywhere within a region (i.e., called safe region) without losing any of the existing reverse skyline points. We also show how to answer why-not questions considering the safe region of the query point. Our approach efficiently combines both query point and data point modification techniques to produce meaningful answers. Experimental results also demonstrate that our approach can produce high quality explanations for why-not questions in reverse skyline queries.
Md. Saiful Islam 0003, Rui Zhou 0001, Chengfei Liu
ICDE1
2012 User Feedback Based Query Refinement by Exploiting Skyline Operator
Md. Saiful Islam 0003, Chengfei Liu, Rui Zhou 0001
ER1