EDBT 2026 Demo / reviewers in the wild / expert
Mojtaba Mostafavi Ghahfarokhi
dblp:239/6274
· DBLP profile ↗
6ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0003-0217-5616ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QualCode: A Data-Driven Framework for Predicting Software Maintainability Based on ISO/IEC 25010
Elham Azhir, Morteza Zakeri Nasrabadi, Yasaman Abedini, Mojtaba Mostafavi Ghahfarokhi |
Sci. Comput. Program. | 4 |
| 2025 | Predicting the understandability of computational notebooks through code metrics analysis
Mojtaba Mostafavi Ghahfarokhi, Alireza Asadi, Arash Asgari, Bardia Mohammadi, Abbas Heydarnoori |
Empir. Softw. Eng. | 1 |
| 2024 | A Roadmap for Enriching Jupyter Notebooks Documentation with Kaggle DataabstractRecent advancements in AI and data science have led to the increased use of Jupyter notebooks. As such, various AI-Based automated tools have been also developed to automatically document notebooks. However, a key challenge is the absence of suitable datasets for training AI models. In this paper, we outline a valuable roadmap for developing a dataset of (markdown, code) pairs centered on functions in Jupyter notebooks. The roadmap encompasses four high-level steps: structural filtering, structural processing, conceptual filtering, and conceptual processing. Our proposed roadmap leads to providing a quality dataset for training AI models on Jupyter notebooks. Mojtaba Mostafavi Ghahfarokhi, Hamed Jahantigh, Alireza Asadi, Sepehr Kianiangolafshani, Ashkan Khademian, Abbas Heydarnoori |
CAIN | 1 |
| 2024 | Beyond Syntax: Unleashing the Power of Computational Notebooks Code Metrics in Documentation GenerationabstractComputational notebooks, like Kaggle notebooks, offer an integrated platform for coding and documentation, yet the latter's quality often falls short as scientists may neglect this crucial aspect. This paper addresses the need for improved and efficient code documentation generation in computational notebooks. As recent literature emphasizes integrating code's inherent structure into documentation generation models, our research explores unutilized structural characteristics, incorporating metrics from code sequences to enable better code documentation suggestions. Evidenced by the improved BLEU scores, our proposed method significantly outperforms the conventional model in a preliminary 10-fold cross-validation experiment and further provides a flexible foundation for integrating source code metrics into diverse code documentation generation models. Mojtaba Mostafavi Ghahfarokhi, Ashkan Khademian, Sepehr Kianiangolafshani, Alireza Asadi, Hamed Jahantigh, Abbas Heydarnoori |
CAIN | 1 |
| 2024 | Can Code Metrics Enhance Documentation Generation for Computational Notebooks?abstractIn software development, code documentation is crucial for collaboration and maintenance, especially as projects become more complex. However, it is often neglected due to the tedious effort it requires. This paper explores automating documentation generation for computational notebooks, focusing on the impact of code metrics such as lines of code, API popularity, and complexity on this task. Using a dataset of 22K code-documentation pairs, we compare deep learning models with and without code metric augmentation. The results show that incorporating these metrics significantly improves the accuracy of documentation generation, underscoring the connection between code metrics and quality documentation. Mojtaba Mostafavi Ghahfarokhi, Hamed Jahantigh, Sepehr Kianiangolafshani, Ashkan Khademian, Alireza Asadi, Abbas Heydarnoori |
ASE | 1 |
| 2024 | DistilKaggle: A Distilled Dataset of Kaggle Jupyter NotebooksabstractJupyter notebooks have become indispensable tools for data analysis and processing in various domains. However, despite their widespread use, there is a notable research gap in understanding and analyzing the contents and code metrics of these notebooks. This gap is primarily attributed to the absence of datasets that encompass both Jupyter notebooks and extracted their code metrics. To address this limitation, we introduce DistilKaggle, a unique dataset specifically curated to facilitate research on code metrics in Jupyter notebooks, utilizing the Kaggle repository as a prime source. Through an extensive study, we identify thirty-four code metrics that significantly impact Jupyter notebook code quality. These features such as lines of code cell, mean number of words in markdown cells, performance tier of developer, etc., are crucial for understanding and improving the overall effectiveness of computational notebooks. The DistilKaggle dataset which is derived from a vast collection of notebooks constitutes two distinct datasets: (i) Code Cells and Markdown Cells Dataset which is presented in two CSV files, allowing for easy integration into researchers' workflows as dataframes. It provides a granular view of the content structure within 542,051 Jupyter notebooks, enabling detailed analysis of code and markdown cells; and (ii) The Notebook Code Metrics Dataset focused on the identified code metrics of notebooks. Researchers can leverage this dataset to access Jupyter notebooks with specific code quality characteristics, surpassing the limitations of filters available on the Kaggle website. Furthermore, the reproducibility of the notebooks in our dataset is ensured through the code cells and markdown cells datasets, offering a reliable foundation for researchers to build upon. Given the substantial size of our datasets, it becomes an invaluable resource for the research community, surpassing the capabilities of individual Kaggle users to collect such extensive data. For accessibility and transparency, both the dataset and the code utilized in crafting this dataset are publicly available at https://github.com/ISE-Research/DistilKaggle. Mojtaba Mostafavi Ghahfarokhi, Arash Asgari, Mohammad Abolnejadian, Abbas Heydarnoori |
MSR | 1 |