VLDB 2026 Research / reviewers in the wild / expert
Mona Nashaat
dblp:208/0750
· DBLP profile ↗
12ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0002-7580-5757ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 5 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An enhanced transformer-based framework for interpretable code clone detection
Mona Nashaat, Reem Amin, Ahmad Hosny Eid, Rabab F. Abdel-Kader |
J. Syst. Softw. | 1 |
| 2025 | Refining software defect prediction through attentive neural models for code understanding
Mona Nashaat, James Miller 0001 |
J. Syst. Softw. | 1 |
| 2024 | Towards Efficient Fine-Tuning of Language Models With Organizational Data for Automated Software ReviewabstractLarge language models like BERT and GPT possess significant capabilities and potential impacts across various applications. Software engineers often use these models for code-related tasks, including generating, debugging, and summarizing code. Nevertheless, large language models still have several flaws, including model hallucination. (e.g., generating erroneous code and producing outdated and inaccurate programs) and the substantial computational resources and energy required for training and fine-tuning. To tackle these challenges, we propose CodeMentor, a framework for few-shot learning to train large language models with the data available within the organization. We employ the framework to train a language model for code review activities, such as code refinement and review generation. The framework utilizes heuristic rules and weak supervision techniques to leverage available data, such as previous review comments, issue reports, and related code updates. Then, the framework employs the constructed dataset to fine-tune LLMs for code review tasks. Additionally, the framework integrates domain expertise by employing reinforcement learning with human feedback. This allows domain experts to assess the generated code and enhance the model performance. Also, to assess the performance of the proposed model, we evaluate it with four state-of-the-art techniques in various code review tasks. The experimental results attest that CodeMentor enhances the performance in all tasks compared to the state-of-the-art approaches, with an improvement of up to 22.3%, 43.4%, and 24.3% in code quality estimation, review generation, and bug report summarization tasks, respectively. Mona Nashaat, James Miller 0001 |
IEEE Trans. Software Eng. | 1 |
| 2022 | Semi-Supervised Ensemble Learning for Dealing with Inaccurate and Incomplete SupervisionabstractIn real-world tasks, obtaining a large set of noise-free data can be prohibitively expensive. Therefore, recent research tries to enable machine learning to work with weakly supervised datasets, such as inaccurate or incomplete data. However, the previous literature treats each type of weak supervision individually, although, in most cases, different types of weak supervision tend to occur simultaneously. Therefore, in this article, we present Smart MEnDR, a Classification Model that applies Ensemble Learning and Data-driven Rectification to deal with inaccurate and incomplete supervised datasets. The model first applies a preliminary phase of ensemble learning in which the noisy data points are detected while exploiting the unlabelled data. The phase employs a semi-supervised technique with maximum likelihood estimation to decide on the disagreement rate. Second, the proposed approach applies an iterative meta-learning step to tackle the problem of knowing which points should be made correct to improve the performance of the final classifier. To evaluate the proposed framework, we report the classification performance, noise detection, and the labelling accuracy of the proposed method against state-of-the-art techniques. The experimental results demonstrate the effectiveness of the proposed framework in detecting noise, providing correct labels, and attaining high classification performance. Mona Nashaat, Aindrila Ghosh, James Miller 0001, Shaikh Quader |
ACM Trans. Knowl. Discov. Data | 1 |
| 2022 | Interpretation of Structural Preservation in Low-Dimensional EmbeddingsabstractDespite being commonly used in big-data analytics; the outcome of dimensionality reduction remains a black-box to most of its users. Understanding the quality of a low-dimensional embedding is important as not only it enables trust in the transformed data, but it can also help to select the most appropriate dimensionality reduction algorithm in a given scenario. As existing research primarily focuses on the visual exploration of embeddings, there is still a need for enhancing interpretability of such algorithms. To bridge this gap, we propose two novel interactive explanation techniques for low-dimensional embeddings obtained fromanydimensionality reduction algorithm. The first technique LAPS produces a local approximation of the neighborhood structure to generate interpretable explanations on the preserved locality for a single instance. The second method GAPS explains the retained global structure of a high-dimensional dataset in its embedding, by combining non-redundant local-approximations from a coarse discretization of the projection space. We demonstrate the applicability of the proposed techniques using 16 real-life tabular, text, image, and audio datasets. Our extensive experimental evaluation shows the utility of the proposed techniques in interpreting the quality of low-dimensional embeddings, as well as with selecting the most suitable dimensionality reduction algorithm for any given dataset. Aindrila Ghosh, Mona Nashaat, James Miller 0001, Shaikh Quader |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | VisExPreS: A Visual Interactive Toolkit for User-Driven Evaluations of EmbeddingsabstractAlthough popularly used in big-data analytics, dimensionality reduction is a complex, black-box technique whose outcome is difficult to interpret and evaluate. In recent years, a number of quantitative and visual methods have been proposed for analyzing low-dimensional embeddings. On the one hand, quantitative methods associate numeric identifiers to qualitative characteristics of these embeddings; and, on the other hand, visual techniques allow users to interactively explore these embeddings and make decisions. However, in the former case, users do not have control over the analysis, while in the latter case. assessment decisions are entirely dependent on the user's perception and expertise. In order to bridge the gap between the two, in this article, we present VisExPreS, a visual interactive toolkit that enables a user-driven assessment of low-dimensional embeddings. VisExPreS is based on three novel techniques namely PG-LAPS, PG-GAPS, and RepSubset, that generate interpretable explanations of the preserved local and global structures in embeddings. In the first two techniques, the VisExPreS system proactively guides users during every step of the analysis. We demonstrate the utility of VisExPreS in interpreting, analyzing, and evaluating embeddings from different dimensionality reduction algorithms using multiple case studies and an extensive user study. Aindrila Ghosh, Mona Nashaat, James Miller 0001, Shaikh Quader |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | Context-Based Evaluation of Dimensionality Reduction Algorithms - Experiments and Statistical Significance AnalysisabstractDimensionality reduction is a commonly used technique in data analytics. Reducing the dimensionality of datasets helps not only with managing their analytical complexity but also with removing redundancy. Over the years, several such algorithms have been proposed with their aims ranging from generating simple linear projections to complex non-linear transformations of the input data. Subsequently, researchers have defined several quality metrics in order to evaluate the performances of different algorithms. Hence, given a plethora of dimensionality reduction algorithms and metrics for their quality analysis, there is a long-existing need for guidelines on how to select the most appropriate algorithm in a given scenario. In order to bridge this gap, in this article, we have compiled 12 state-of-the-art quality metrics and categorized them into 5 identified analytical contexts. Furthermore, we assessed 15 most popular dimensionality reduction algorithms on the chosen quality metrics using a large-scale and systematic experimental study. Later, using a set of robust non-parametric statistical tests, we assessed the generalizability of our evaluation on 40 real-world datasets. Finally, based on our results, we present practitioners’ guidelines for the selection of an appropriate dimensionally reduction algorithm in the present analytical contexts. Aindrila Ghosh, Mona Nashaat, James Miller 0001, Shaikh Quader |
ACM Trans. Knowl. Discov. Data | 2 |
| 2019 | The current state of software license renewals in the I.T. industry
Aindrila Ghosh, Mona Nashaat, James Miller 0001 |
Inf. Softw. Technol. | 2 |
| 2019 | M-Lean: An end-to-end development framework for predictive models in B2B scenarios
Mona Nashaat, Aindrila Ghosh, James Miller 0001, Shaikh Quader, Chad Marston |
Inf. Softw. Technol. | 1 |
| 2018 | Hybridization of Active Learning and Data Programming for Labeling Large Industrial DatasetsabstractModern machine learning (ML) models are being used heavily in business domains to build effective decision support systems. As a primary requirement, supervised ML models need large labeled datasets. However, obtaining a high volume of labeled training data is both expensive and time-consuming. Researchers have proposed several labeling approaches to avoid manual labeling efforts. Active learning (AL) and Data Programming (DP) are two state-of-the-art techniques used to label datasets. Nevertheless, both approaches have their strengths and weaknesses. For example, AL is computationally expensive to apply on large industrial datasets; and labels generated by DP are often inaccurate and difficult to interpret. To address these challenges, in this paper, we propose a novel hybrid method that integrates the scalability of DP with the user engagement and accuracy of AL. The proposed approach aims at optimizing the labeling process by applying DP to generate initial noisy training data and then use AL to query the user to label only those points that maximize the accuracy of the final labels with a minimum annotation cost. To evaluate the proposed approach, we have used five open source datasets and a real-world business dataset of 1.5 million records. We use traditional active learning and data programming techniques as baselines to compare the performance and annotation cost of our proposed approach. The results show that the proposed method can achieve higher labeling accuracy than data programming. It also can minimize the labeling cost in real-world business scenarios, while delivering a comparable level of performance (accuracy) with active learning. Mona Nashaat, Aindrila Ghosh, James Miller 0001, Shaikh Quader, Chad Marston, Jean-François Puget |
IEEE BigData | 1 |
| 2018 | A comprehensive review of tools for exploratory analysis of tabular industrial datasetsabstractExploratory data analysis plays a major role in obtaining insights from data. Over the last two decades, researchers have proposed several visual data exploration tools that can assist with each step of the analysis process. Nevertheless, in recent years, data analysis requirements have changed significantly. With constantly increasing size and types of data to be analyzed, scalability and analysis duration are now among the primary concerns of researchers. Moreover, in order to minimize the analysis cost, businesses are in need of data analysis tools that can be used with limited analytical knowledge. To address these challenges, traditional data exploration tools have evolved within the last few years. In this paper, with an in-depth analysis of an industrial tabular dataset, we identify a set of additional exploratory requirements for large datasets. Later, we present a comprehensive survey of the recent advancements in the emerging field of exploratory data analysis. We investigate 50 academic and non-academic visual data exploration tools with respect to their utility in the six fundamental steps of the exploratory data analysis process. We also examine the extent to which these modern data exploration tools fulfill the additional requirements for analyzing large datasets. Finally, we identify and present a set of research opportunities in the field of visual exploratory data analysis. Aindrila Ghosh, Mona Nashaat, James Miller 0001, Shaikh Quader, Chad Marston |
Vis. Informatics | 2 |
| 2017 | Detecting Security Vulnerabilities in Object-Oriented PHP ProgramsabstractPHP is one of the most popular web development tools in use today. A major concern though is the improper and insecure uses of the language by application developers, motivating the development of various static analyses that detect security vulnerabilities in PHP programs. However, many of these approaches do not handle recent, important PHP features such as object orientation, which greatly limits the use of such approaches in practice. In this paper, we present OOPIXY, a security analysis tool that extends the PHP security analyzer PIXY to support reasoning about object-oriented features in PHP applications. Our empirical evaluation shows that OOPIXY detects 88% of security vulnerabilities found in micro benchmarks. When used on real-world PHP applications, OOPIXY detects security vulnerabilities that could not be detected using state-of-the-art tools, retaining a high level of precision. We have contacted the maintainers of those applications, and two applications' development teams verified the correctness of our findings. They are currently working on fixing the bugs that lead to those vulnerabilities. Mona Nashaat, Karim Ali 0001, James Miller 0001 |
SCAM | 1 |