EDBT 2026 Demo / reviewers in the wild / expert
Shaikh Quader
dblp:224/6596
· DBLP profile ↗
10ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0009-4264-0528ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LearnedWMP: Workload Memory Prediction Using Distribution of Query Templates
Shaikh Quader, Andres Jaramillo, Sumona Mukhopadhyay, Ghadeer AbuOda, Calisto Zuzarte, David Kalmuk, Marin Litoiu, Manos Papagelis |
EDBT | 1 |
| 2022 | Semi-Supervised Ensemble Learning for Dealing with Inaccurate and Incomplete SupervisionabstractIn real-world tasks, obtaining a large set of noise-free data can be prohibitively expensive. Therefore, recent research tries to enable machine learning to work with weakly supervised datasets, such as inaccurate or incomplete data. However, the previous literature treats each type of weak supervision individually, although, in most cases, different types of weak supervision tend to occur simultaneously. Therefore, in this article, we present Smart MEnDR, a Classification Model that applies Ensemble Learning and Data-driven Rectification to deal with inaccurate and incomplete supervised datasets. The model first applies a preliminary phase of ensemble learning in which the noisy data points are detected while exploiting the unlabelled data. The phase employs a semi-supervised technique with maximum likelihood estimation to decide on the disagreement rate. Second, the proposed approach applies an iterative meta-learning step to tackle the problem of knowing which points should be made correct to improve the performance of the final classifier. To evaluate the proposed framework, we report the classification performance, noise detection, and the labelling accuracy of the proposed method against state-of-the-art techniques. The experimental results demonstrate the effectiveness of the proposed framework in detecting noise, providing correct labels, and attaining high classification performance. Mona Nashaat, Aindrila Ghosh, James Miller 0001, Shaikh Quader |
ACM Trans. Knowl. Discov. Data | 4 |
| 2022 | Interpretation of Structural Preservation in Low-Dimensional EmbeddingsabstractDespite being commonly used in big-data analytics; the outcome of dimensionality reduction remains a black-box to most of its users. Understanding the quality of a low-dimensional embedding is important as not only it enables trust in the transformed data, but it can also help to select the most appropriate dimensionality reduction algorithm in a given scenario. As existing research primarily focuses on the visual exploration of embeddings, there is still a need for enhancing interpretability of such algorithms. To bridge this gap, we propose two novel interactive explanation techniques for low-dimensional embeddings obtained fromanydimensionality reduction algorithm. The first technique LAPS produces a local approximation of the neighborhood structure to generate interpretable explanations on the preserved locality for a single instance. The second method GAPS explains the retained global structure of a high-dimensional dataset in its embedding, by combining non-redundant local-approximations from a coarse discretization of the projection space. We demonstrate the applicability of the proposed techniques using 16 real-life tabular, text, image, and audio datasets. Our extensive experimental evaluation shows the utility of the proposed techniques in interpreting the quality of low-dimensional embeddings, as well as with selecting the most suitable dimensionality reduction algorithm for any given dataset. Aindrila Ghosh, Mona Nashaat, James Miller 0001, Shaikh Quader |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | VisExPreS: A Visual Interactive Toolkit for User-Driven Evaluations of EmbeddingsabstractAlthough popularly used in big-data analytics, dimensionality reduction is a complex, black-box technique whose outcome is difficult to interpret and evaluate. In recent years, a number of quantitative and visual methods have been proposed for analyzing low-dimensional embeddings. On the one hand, quantitative methods associate numeric identifiers to qualitative characteristics of these embeddings; and, on the other hand, visual techniques allow users to interactively explore these embeddings and make decisions. However, in the former case, users do not have control over the analysis, while in the latter case. assessment decisions are entirely dependent on the user's perception and expertise. In order to bridge the gap between the two, in this article, we present VisExPreS, a visual interactive toolkit that enables a user-driven assessment of low-dimensional embeddings. VisExPreS is based on three novel techniques namely PG-LAPS, PG-GAPS, and RepSubset, that generate interpretable explanations of the preserved local and global structures in embeddings. In the first two techniques, the VisExPreS system proactively guides users during every step of the analysis. We demonstrate the utility of VisExPreS in interpreting, analyzing, and evaluating embeddings from different dimensionality reduction algorithms using multiple case studies and an extensive user study. Aindrila Ghosh, Mona Nashaat, James Miller 0001, Shaikh Quader |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2021 | Context-Based Evaluation of Dimensionality Reduction Algorithms - Experiments and Statistical Significance AnalysisabstractDimensionality reduction is a commonly used technique in data analytics. Reducing the dimensionality of datasets helps not only with managing their analytical complexity but also with removing redundancy. Over the years, several such algorithms have been proposed with their aims ranging from generating simple linear projections to complex non-linear transformations of the input data. Subsequently, researchers have defined several quality metrics in order to evaluate the performances of different algorithms. Hence, given a plethora of dimensionality reduction algorithms and metrics for their quality analysis, there is a long-existing need for guidelines on how to select the most appropriate algorithm in a given scenario. In order to bridge this gap, in this article, we have compiled 12 state-of-the-art quality metrics and categorized them into 5 identified analytical contexts. Furthermore, we assessed 15 most popular dimensionality reduction algorithms on the chosen quality metrics using a large-scale and systematic experimental study. Later, using a set of robust non-parametric statistical tests, we assessed the generalizability of our evaluation on 40 real-world datasets. Finally, based on our results, we present practitioners’ guidelines for the selection of an appropriate dimensionally reduction algorithm in the present analytical contexts. Aindrila Ghosh, Mona Nashaat, James Miller 0001, Shaikh Quader |
ACM Trans. Knowl. Discov. Data | 4 |
| 2020 | Author Correction: Customer support ticket escalation prediction using feature engineering
Lloyd Montgomery, Daniela E. Damian, Tyson Bulmer, Shaikh Quader |
Requir. Eng. | 4 |
| 2019 | M-Lean: An end-to-end development framework for predictive models in B2B scenarios
Mona Nashaat, Aindrila Ghosh, James Miller 0001, Shaikh Quader, Chad Marston |
Inf. Softw. Technol. | 4 |
| 2018 | Hybridization of Active Learning and Data Programming for Labeling Large Industrial DatasetsabstractModern machine learning (ML) models are being used heavily in business domains to build effective decision support systems. As a primary requirement, supervised ML models need large labeled datasets. However, obtaining a high volume of labeled training data is both expensive and time-consuming. Researchers have proposed several labeling approaches to avoid manual labeling efforts. Active learning (AL) and Data Programming (DP) are two state-of-the-art techniques used to label datasets. Nevertheless, both approaches have their strengths and weaknesses. For example, AL is computationally expensive to apply on large industrial datasets; and labels generated by DP are often inaccurate and difficult to interpret. To address these challenges, in this paper, we propose a novel hybrid method that integrates the scalability of DP with the user engagement and accuracy of AL. The proposed approach aims at optimizing the labeling process by applying DP to generate initial noisy training data and then use AL to query the user to label only those points that maximize the accuracy of the final labels with a minimum annotation cost. To evaluate the proposed approach, we have used five open source datasets and a real-world business dataset of 1.5 million records. We use traditional active learning and data programming techniques as baselines to compare the performance and annotation cost of our proposed approach. The results show that the proposed method can achieve higher labeling accuracy than data programming. It also can minimize the labeling cost in real-world business scenarios, while delivering a comparable level of performance (accuracy) with active learning. Mona Nashaat, Aindrila Ghosh, James Miller 0001, Shaikh Quader, Chad Marston, Jean-François Puget |
IEEE BigData | 4 |
| 2018 | Customer support ticket escalation prediction using feature engineering
Lloyd Montgomery, Daniela E. Damian, Tyson Bulmer, Shaikh Quader |
Requir. Eng. | 4 |
| 2018 | A comprehensive review of tools for exploratory analysis of tabular industrial datasetsabstractExploratory data analysis plays a major role in obtaining insights from data. Over the last two decades, researchers have proposed several visual data exploration tools that can assist with each step of the analysis process. Nevertheless, in recent years, data analysis requirements have changed significantly. With constantly increasing size and types of data to be analyzed, scalability and analysis duration are now among the primary concerns of researchers. Moreover, in order to minimize the analysis cost, businesses are in need of data analysis tools that can be used with limited analytical knowledge. To address these challenges, traditional data exploration tools have evolved within the last few years. In this paper, with an in-depth analysis of an industrial tabular dataset, we identify a set of additional exploratory requirements for large datasets. Later, we present a comprehensive survey of the recent advancements in the emerging field of exploratory data analysis. We investigate 50 academic and non-academic visual data exploration tools with respect to their utility in the six fundamental steps of the exploratory data analysis process. We also examine the extent to which these modern data exploration tools fulfill the additional requirements for analyzing large datasets. Finally, we identify and present a set of research opportunities in the field of visual exploratory data analysis. Aindrila Ghosh, Mona Nashaat, James Miller 0001, Shaikh Quader, Chad Marston |
Vis. Informatics | 4 |