Mohamed Khalefa

dblp:213/1360 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 A JSON document algebra for query optimization
Tomás F. Llano-Ríos, Mohamed Khalefa, Antonio Badia
Inf. Syst.2
2024 Adaptive Benchmarking for Data System using LLMs
abstract
Benchmarking is crucial for understanding the complexity of big data systems. However, designing and running benchmarks is often tedious, time-intensive, and can be costly where the large data volumes required for benchmarks drive up storage needs and costs.Moreover, standard benchmarks may not always reflect user requirements, as users might want to adjust data generation methods or introduce new query types. In this work, we leverage Large Language Models (LLMs) to transform user prompts into detailed benchmarking plans. These plans can: (1) load and configure data systems, (2) generate benchmark data, (3) execute benchmarks while monitoring runtime, and (4) adapt the next steps based on the gathered metrics and user prompt. We use DSPy framework to guarantee the generated plan adheres to the user and system requirements and optimize the LLM-based workflow.
Mohamed Khalefa
IEEE Big Data1
2024 A JSON Document Algebra for Query Optimization
Antonio Badia, Mohamed Khalefa, Tomás F. Llano-Ríos
DOLAP2
2020 Evaluating NoSQL Systems for Decision Support: An Experimental Approach
abstract
We design and implement an experimental analysis comparing two relational systems (PostgreSQL and MariaDB) and two document-based NoSQL systems (MongoDB and CouchBase). We compare their performance on a single server, Decision Support (DSS) scenario. We argue that DSS is becoming an important case study for NoSQL. We experiment with several database designs and several query translations in order to investigate the effect of physical design and query optimization in document-based stores. Our results show that design is very important for MongoDB's performance, and that query optimization over documents is much less sophisticated on document-based stores than in relational data bases and needs to improve. Our results also offer some ideas to guide further development in this area.
Tomás F. Llano-Ríos, Mohamed Khalefa, Antonio Badia
IEEE BigData2
2017 Towards a distributed infrastructure for data-driven discoveries & analysis
abstract
Big data analytics traditionally involves download of massive amounts of datasets to common server/cluster for processing. Analytic process gets slower with increasing size of required data and network conditions. Data scientists also need explicit access to data locations to download required data. Explicit access to required data may not always be granted due to security reasons. To simplify and accelerate the analytics process on distributed big data with security considerations, we proposed the Virtual Information Fabric Infrastructure (VIFI) for data driven discoveries. Instead of moving large amounts of data to a common place of processing, VIFI allows automatic transfer of required analytics programs to the distributed data locations for in-place processing of relevant data. VIFI allows data scientists to conduct and coordinate complex analytics processes on distributed data repositories using containerization technology and open-source workflow design tools. VIFI alleviates users from having detailed knowledge of distributed data locations, as well as required dependencies, installation and configuration of analytical libraries. In this paper, we demonstrate our current and future work to improve the VIFI architecture using previous and additional uses cases, data management layer that simplifies search of relevant data sets through addition of metadata, integration with security policies at different institutions with the proposed VIFI security layer, and the use of a user-friendly web interface to carry different VIFI activities.
Mohammed Elshambakey, Mohamed Khalefa, William J. Tolone, Sreyasee Das Bhattacharjee, Huikyo Lee, Luca Cinquini, Shannon Schlueter, Isaac Cho, Wenwen Dou, Daniel J. Crichton
IEEE BigData2