VLDB 2026 Research / reviewers in the wild / expert
Ahmed Helal
dblp:245/1796
· DBLP profile ↗
4ranked-venue papers
2as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | KGLiDS: A Platform for Semantic Abstraction, Linking, and Automation of Data ScienceabstractIn recent years, we have witnessed the growing interest from academia and industry in applying data science technologies to analyze large amounts of data. In this process, a myriad of artifacts (datasets, pipeline scripts, etc.) are created. However, there has been no systematic attempt to holistically collect and exploit all the knowledge and experiences that are implicitly contained in those artifacts. Instead, data scientists recover information and expertise from colleagues or learn via trial and error. Hence, this paper presents a scalable platform, KGLiDS, that employs machine learning and knowledge graph technologies to abstract and capture the semantics of data science artifacts and their connections. Based on this information, KGLiDS enables various downstream applications, such as data discovery and pipeline automation. Our comprehensive evaluation covers use cases in data discovery, data cleaning, transformation, and AutoML. It shows that KGLiDS is significantly faster with a lower memory footprint than the state-of-the-art systems while achieving comparable or better accuracy. Mossad Helali, Niki Monjazeb, Shubham Vashisth, Philippe Carrier, Ahmed Helal, Antonio Cavalcante, Khaled Ammar, Katja Hose, Essam Mansour 0001 |
ICDE | 5 |
| 2021 | Data Lakes Empowered by Knowledge Graph TechnologiesabstractWith the emergence of open data, governments [19, 25], Kaggle [13], OpenML [26], and organizations [2, 20] have started making data available on the web. This data represents an opportunity for Artificial Intelligence (AI) practitioners to accomplish their tasks including but not limited to improving the performance of models in Machine Learning or predicting insights about the future in Scenario Planning. However with the proliferation of available data, finding the most relevant one is time-consuming and cumbersome. Practitioners tend to spend considerable time looking for data to enrich their existing datasets to accomplish their tasks. For instance, in Deep Learning, engineers need a lot of data to train their models [27]. Moreover, they face several problems such as increasing the size of their datasets with similar data or including more features that can contribute to better results [3, 8, 27]. These problems emerge because data lakes are schema-agnostic repositories [30]. Ahmed Helal |
SIGMOD Conference | 1 |
| 2021 | A Demonstration of KGLac: A Data Discovery and Enrichment Platform for Data ScienceabstractData science growing success relies on knowing where a relevant dataset exists, understanding its impact on a specific task, finding ways to enrich a dataset, and leveraging insights derived from it. With the growth of open data initiatives, data scientists need an extensible set of effective discovery operations to find relevant data from their enterprise datasets accessible via data discovery systems or open datasets accessible via data portals. Existing portals and systems suffer from limited discovery support and do not track the use of a dataset and insights derived from it. We will demonstrate KGLac, a system that captures metadata and semantics of datasets to construct a knowledge graph (GLac) interconnecting data items, e.g., tables and columns. KGLac supports various data discovery operations via SPARQL queries for table discovery, unionable and joinable tables, plus annotation with related derived insights. We harness a broad range of Machine Learning (ML) approaches with GLac to enable automatic graph learning for advanced and semantic data discovery. The demo will showcase how KGLac facilitates data discovery and enrichment while developing an ML pipeline to evaluate potential gender salary bias in IT jobs. Ahmed Helal, Mossad Helali, Khaled Ammar, Essam Mansour 0001 |
Proc. VLDB Endow. | 1 |
| 2019 | Chi squared feature selection over Apache SparkabstractWe live in the age of big data and distributed computing. The current large scale computation frameworks are based on a scaling-out approach for distributing tasks over a cluster of commodity machines. Apache Spark is one of these frameworks that has excelled in many computational tasks. Implementation of statistical learning algorithms over Spark is a challenging task. A bad implementation may lead to a significant decrease in performance and a waste of cluster time and money. Poor performance is mostly due to a lack of understanding of the data in hand and Spark's underlying mechanisms more than it is due to a deficit in the framework itself. In this paper, we consider the use case of X2 feature selection which is very popular in supervised learning pipelines. Our implementation follows the algorithm of the Scikit-learn Python machine learning library which is different than the algorithm used by the Spark machine learning library. The Spark ML library implementation of X2 feature selection accepts only categorical features. Our alternative implementation is more suitable for numerical features. We experiment in particular with features of high sparsity such as n-gram counts. We study the best partitioning scheme of the data and the optimal number of partitions. Our experiments are run over the Databricks platform. Mohamed Nassar 0001, Haïdar Safa, Alaa Al Mutawa, Ahmed Helal, Iskander Gaba |
IDEAS | 4 |