Abdurrahman Ghanem

dblp:178/3717 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Knowledge graphs · 32% Data mining · 19% Graph data management · 19%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 91% Parallel and multicore computing · 9%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge graphs
knowledge graph querying
1.022022
RDFFrames: knowledge graph access for machine learning tools · VLDB J. 2022
RDFFrames: Knowledge Graph Access for Machine Learning Tools · Proc. VLDB Endow. 2020
Machine learning and data management › data management for machine learning
data preparation for machine learning
0.612022
RDFFrames: knowledge graph access for machine learning tools · VLDB J. 2022
Graph data management
RDF data management
0.612022
RDFFrames: knowledge graph access for machine learning tools · VLDB J. 2022
Query processing and optimization › query optimization › transformation-based optimization
query pushdown
0.412020
PushdownDB: Accelerating a DBMS Using S3 Computation · ICDE 2020
Knowledge graphs › knowledge graph querying
SPARQL query generation
0.412020
RDFFrames: Knowledge Graph Access for Machine Learning Tools · Proc. VLDB Endow. 2020
Cloud and datacenter computing › cloud data management
cloud query execution
0.412020
PushdownDB: Accelerating a DBMS Using S3 Computation · ICDE 2020
Cloud and datacenter computing
cloud storage
0.412020
PushdownDB: Accelerating a DBMS Using S3 Computation · ICDE 2020
Data mining › pattern mining › graph pattern mining
clique enumeration
0.312017
Graph Data Mining with Arabesque · SIGMOD Conference 2017
Data mining › structured data mining
graph mining
0.312017
Graph Data Mining with Arabesque · SIGMOD Conference 2017
Graph data management
motif counting
0.312017
Graph Data Mining with Arabesque · SIGMOD Conference 2017
Data mining › structured data mining › graph mining
subgraph mining
0.312017
Graph Data Mining with Arabesque · SIGMOD Conference 2017
Machine learning and data management
data management for machine learning
0.112020
RDFFrames: Knowledge Graph Access for Machine Learning Tools · Proc. VLDB Endow. 2020
Parallel and multicore computing › graph processing
parallel graph analytics
0.112017
Graph Data Mining with Arabesque · SIGMOD Conference 2017

Methods — techniques the papers use, named apart from their topics

s3 select · 0.9graph traversal · 0.4API translation · 0.4
YearPublicationVenuePosition
2022 RDFFrames: knowledge graph access for machine learning tools
abstract
Abstract Knowledge graphs represented as RDF datasets are integral to many machine learning applications. RDF is supported by a rich ecosystem of data management systems and tools, most notably RDF database systems that provide a SPARQL query interface. Surprisingly, machine learning tools for knowledge graphs do not use SPARQL, despite the obvious advantages of using a database system. This is due to the mismatch between SPARQL and machine learning tools in terms of data model and programming style. Machine learning tools work on data in tabular format and process it using an imperative programming style, while SPARQL is declarative and has as its basic operation matching graph patterns to RDF triples. We posit that a good interface to knowledge graphs from a machine learning software stack should use an imperative, navigational programming paradigm based on graph traversal rather than the SPARQL query paradigm based on graph patterns. In this paper, we present RDFFrames, a framework that provides such an interface. RDFFrames provides an imperative Python API that gets internally translated to SPARQL, and it is integrated with the PyData machine learning software stack. RDFFrames enables the user to make a sequence of Python calls to define the data to be extracted from a knowledge graph stored in an RDF database system, and it translates these calls into a compact SPQARL query, executes it on the database system, and returns the results in a standard tabular format. Thus, RDFFrames is a useful tool for data preparation that combines the usability of PyData with the flexibility and performance of RDF database systems.
Aisha Mohamed, Ghadeer AbuOda, Abdurrahman Ghanem, Zoi Kaoudi, Ashraf Aboulnaga
VLDB J.3
2020 PushdownDB: Accelerating a DBMS Using S3 Computation
abstract
This paper studies the effectiveness of pushing parts of DBMS analytics queries into the Simple Storage Service (S3) of Amazon Web Services (AWS), using a recently released capability called S3 Select. We show that some DBMS primitives (filter, projection, and aggregation) can always be cost-effectively moved into S3. Other more complex operations (join, top-K, and group-by) require reimplementation to take advantage of S3 Select and are often candidates for pushdown. We demonstrate these capabilities through experimentation using a new DBMS that we developed, PushdownDB. Experimentation with a collection of queries including TPC-H queries shows that PushdownDB is on average 30% cheaper and 6.7× faster than a baseline that does not use S3 Select.
Xiangyao Yu, Matt Youill, Matthew E. Woicik, Abdurrahman Ghanem, Marco Serafini, Ashraf Aboulnaga, Michael Stonebraker
ICDE4
2020 RDFFrames: Knowledge Graph Access for Machine Learning Tools
abstract
Knowledge graphs represented in RDF are becoming increasingly popular and are essential to many machine learning applications. A rich ecosystem of RDF data management systems and tools has evolved over the years, most notably RDF database management systems that support the SPARQL query language. Surprisingly, machine learning tools for knowledge graphs typically do not use SPARQL despite the obvious advantages of using a database system. This is due to the mismatch between SPARQL and machine learning tools in terms of expected data model and interface style. Machine learning tools work on data in tabular format and process it using imperative relational API calls, while SPARQL matches graph patterns to RDF triples. To access knowledge graphs for machine learning, we observe that it is more natural to use a navigational paradigm based on graph traversal rather than the SPARQL paradigm based on triple patterns. We demonstrate RDFFrames, a framework that bridges the gap between machine learning tools and RDF database systems by offering the usability and flexibility of machine learning tools together with the performance of a database system. RDFFrames enables the user to make a sequence of Python calls to define the data to be extracted from a knowledge graph stored in an RDF database system, and it translates these calls into a compact SPARQL query, executes it on the database system, and returns the results in a standard tabular format.
Aisha Mohamed, Ghadeer AbuOda, Abdurrahman Ghanem, Zoi Kaoudi, Ashraf Aboulnaga
Proc. VLDB Endow.3
2017 Graph Data Mining with Arabesque
abstract
Graph data mining is defined as searching in an input graph for all subgraphs that satisfy some property that makes them interesting to the user. Examples of graph data mining problems include frequent subgraph mining, counting motifs, and enumerating cliques. These problems differ from other graph processing problems such as PageRank or shortest path in that graph data mining requires searching through an exponential number of subgraphs. Most current parallel graph analytics systems do not provide good support for graph data mining. One notable exception is Arabesque, a system that was built specifically to support graph data mining. Arabesque provides a simple programming model to express graph data mining computations, and a highly scalable and efficient implementation of this model, scaling to billions of subgraphs on hundreds of cores. This demonstration will showcase the Arabesque system, focusing on the end-user experience and showing how Arabesque can be used to simply and efficiently solve practical graph data mining problems that would be difficult with other systems.
Eslam Hussein, Abdurrahman Ghanem, Vinícius Vitor dos Santos Dias, Carlos H. C. Teixeira, Ghadeer AbuOda, Marco Serafini, Georgos Siganos, Gianmarco De Francisci Morales, Ashraf Aboulnaga, Mohammed J. Zaki
SIGMOD Conference2