Kemafor Anyanwu

dblp:51/1028 · DBLP profile ↗
← Back
30ranked-venue papers
4as first author
2since 2021 · last 2025
0000-0002-9528-061XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 25 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 3 first-authorArtificial intelligence and machine learning · 6Systems, architecture and hardware · 1Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
8 papers
Query processing and optimization · 35% Graph data management · 20% Information retrieval · 13%

Topics — the 19 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Database theory
ontology-mediated queries
0.412019
Semantic query transformations for increased parallelization in distributed knowledge graph query processing · SC 2019
Query processing and optimization
parallel query processing
0.412019
Semantic query transformations for increased parallelization in distributed knowledge graph query processing · SC 2019
Graph data management › RDF data management
RDF query processing
0.322013
PrefixSolve: efficiently solving multi-source multi-destination path queries on RDF graphs by sharing suffix computations · WWW 2013
From SPARQL to MapReduce: The Journey Using a Nested TripleGroup Algebra · Proc. VLDB Endow. 2011
Query processing and optimization
query rewriting
0.312017
Type-based Semantic Optimization for Scalable RDF Graph Pattern Matching · WWW 2017
Graph data management › graph data model
RDF data model
0.122017
Type-based Semantic Optimization for Scalable RDF Graph Pattern Matching · WWW 2017
?-Queries: enabling querying for semantic associations on the semantic web · WWW 2003
Distributed and cloud data management
mapreduce
0.112011
From SPARQL to MapReduce: The Journey Using a Nested TripleGroup Algebra · Proc. VLDB Endow. 2011
Data models and query languages
query language
0.112011
From SPARQL to MapReduce: The Journey Using a Nested TripleGroup Algebra · Proc. VLDB Endow. 2011
Information retrieval › cross-language information retrieval
query translation
0.112011
From SPARQL to MapReduce: The Journey Using a Nested TripleGroup Algebra · Proc. VLDB Endow. 2011
Data models and query languages › RDF query language
SPARQL
0.112011
From SPARQL to MapReduce: The Journey Using a Nested TripleGroup Algebra · Proc. VLDB Endow. 2011
Information retrieval › query reformulation
query expansion
0.112019
Semantic query transformations for increased parallelization in distributed knowledge graph query processing · SC 2019
Query processing and optimization
cardinality estimation
0.112007
Estimating the cardinality of RDF graph patterns · WWW 2007
Data models and query languages › query language design
query language extension
0.112007
SPARQ2L: towards support for subgraph extraction queries in rdf databases · WWW 2007
Graph data management
RDF data management
0.112007
SPARQ2L: towards support for subgraph extraction queries in rdf databases · WWW 2007
Information retrieval
ranking
0.112005
SemRank: ranking complex relationship search results on the semantic web · WWW 2005
Information retrieval › ranking › text ranking › semantic ranking
semantic association ranking
0.112005
SemRank: ranking complex relationship search results on the semantic web · WWW 2005
Graph data management
graph query processing
0.012007
SPARQ2L: towards support for subgraph extraction queries in rdf databases · WWW 2007
Information retrieval › retrieval models › language model
relevance model
0.012005
SemRank: ranking complex relationship search results on the semantic web · WWW 2005
Information retrieval
retrieval models
0.012005
SemRank: ranking complex relationship search results on the semantic web · WWW 2005
Graph data management
graph pattern
0.012003
?-Queries: enabling querying for semantic associations on the semantic web · WWW 2003

Methods — techniques the papers use, named apart from their topics

semantic query transformation · 0.4ontology inference · 0.4integrity constraints · 0.3indexing · 0.3suffix computation sharing · 0.2nested triplegroup algebra · 0.1pattern-based summarization · 0.1algebraic path computation · 0.1information-theoretic ranking · 0.1heuristics · 0.1
YearPublicationVenuePosition
2025 Taming the Beast of User-Programmed Transactions on Blockchains: A Declarative Transaction Approach
Nodirbek Korchiev, Akash Pateria, Vodelina Samatova, Sogolsadat Mansouri, Kemafor Anyanwu
EDBT5
2025 Towards Declarative Blockchains: A SHACL-Based Model for Robust and Efficient Transactions
Kemafor Anyanwu, Sogolsadat Mansouri, David Adei
ICBC1
2020 Efficient Constrained Subgraph Extraction for Exploratory Discovery in Large Knowledge Graphs
abstract
Knowledge graphs which often integrate heterogeneous data can be exploited for serendipitous knowledge discovery using appropriate integration paradigms. We posit that a semi-structured querying model which blends the benefits of structured and unstructured querying could offer a sweetspot. However, there is a need for effective algorithmic techniques for such query processing.In this paper, we propose a class of constrained subgraph connection structure discovery queries whose specification is only partially structured. Graph theoretically, these amount subgraph homeomorphism problems that tolerate flexibility in graph structure matching. Central to achieving the goals of performance and scale of query evaluation is the use of a path algebraic framework rather than a graph theoretic framework. The path algebraic framework is coupled with some efficient data encoding, representation and indexing. Together, these allow more effective querying than using the traditional graph traversal style algorithms, demonstrated by a comparative evaluation.
Sidan Gao, Nodirbek Korchiev, Vodelina Samatova, Kemafor Anyanwu
IEEE BigData4
2019 Semantic query transformations for increased parallelization in distributed knowledge graph query processing
abstract
Ontologies have become an increasingly popular semantic layer for integrating multiple heterogeneous datasets. However, significant challenges remain with supporting efficient and scalable processing of queries with data linked with ontologies (ontological queries). Ontological query processing queries requires explicitly defined query patterns be expanded to capture implicit ones, based on available ontology inference axioms. However, in practice such as in the biomedical domain, the complexity of the ontological axioms results in significantly large query expansions which present day query processing infrastructure cannot support. In particular, it remains unclear how to effectively parallelize such queries.
HyeongSik Kim 0001, Abhisha Bhattacharyya, Kemafor Anyanwu
SC3
2018 Scalable Exploratory Search on Knowledge Graphs Using Apache Spark
abstract
Faceted search is a popular exploratory search paradigm on Big Knowledge Graphs. Translating exploration steps into database queries for processing leads to several joins when dealing with knowledge graphs as opposed to filter conditions when dealing with structured data. Further, existing engines handle each exploration step as independent queries in spite of data dependencies that often exist between steps. In this work, we propose an incremental query execution model RAPIDFacet, that exploits the iterative nature of faceted search and reuses intermediate results. The approach is built on top of Apache Spark which naturally supports iterative models and the Nested Triplegroup Data Model and Algebra (NTGA) which uses a coarse grained data model to avoid joins. Evaluations showed up to 150× faster execution than existing approaches.
Avimanyu Mukhopadhyay, HyeongSik Kim 0001, Kemafor Anyanwu
WETICE3
2017 A semantics-aware storage framework for scalable processing of knowledge graphs on Hadoop
abstract
Knowledge graphs are graph-based data models which employ named nodes and edges to capture differentiation among entities and relationships in richly diverse data collections such as in the biomedical domain. The flexibility of knowledge graphs allows for heterogeneous collections to be linked and integrated in precise ways. However, resulting data models often have irregular structures which are not easy to manage using platforms for structured, schema-first data models like the relational model. To facilitate exchange, inter-operability and reuse of data, standards such as Resource Description Framework (RDF) have been increasingly adopted for representation. Domains such as the biomedical now have large collections of publicly available RDF graphs as well as benchmark workloads. To achieve scalability in data processing, some efforts are being made to build on distributed processing platforms such as Hadoop and Spark. However, while some distributed graph platforms have emerged for certain classes of mining workloads for non-semantic graphs (without typed edges and nodes), knowledge graph processing, which often involves ontological inferencing, continues to be plagued by scalability and efficiency challenges. In this paper, we present the design of a Hadoop-based storage architecture for knowledge graphs that overcomes some of the challenges of big RDF data processing. The rationale of the design strategy is to go beyond the traditional approach of exploiting structural properties of graphs while storing to include exploitation of semantic properties of knowledge graphs. Our system SemStorm is a Hadoop-based indexed, polymorphic, signatured file organization that supports efficient storage of data collections with significant data heterogeneity. Naive storage models for such data place more demands for meta-data management than traditional systems can support. The polymorphic file organization is further coupled with a nested, column-oriented file format to enable discriminatory data access based on queries. A major hallmark of SemStorm is the enabling of semantic-awareness in storage framework. The idea is to exploit the knowledge represented in ontologies that accompany data for optimizing data storage models such as identifying and managing data (sometimes implicit) redundancies. Another major advantage of SemStorm is that it derives optimized storage models for data autonomically, i.e., without user input. Extensive experiments conducted on real-world and synthetic benchmark datasets show that SemStorm is up to 10X faster than existing approaches.
HyeongSik Kim 0001, Padmashree Ravindra, Kemafor Anyanwu
IEEE BigData3
2017 Type-based Semantic Optimization for Scalable RDF Graph Pattern Matching
abstract
Scalable query processing relies on early and aggressive determination and pruning of query-irrelevant data. Besides the traditional space-pruning techniques such as indexing, type-based optimizations that exploit integrity constraints defined on the types can be used to rewrite queries into more efficient ones. However, such optimizations are only applicable in strongly-typed data and query models which make it a challenge for semi-structured models such as RDF. Consequently, developing techniques for enabling typebased query optimizations will contribute new insight to improving the scalability of RDF processing systems.
HyeongSik Kim 0001, Padmashree Ravindra, Kemafor Anyanwu
WWW3
2016 Optimization of Complex SPARQL Analytical Queries
Padmashree Ravindra, HyeongSik Kim 0001, Kemafor Anyanwu
EDBT3
2015 Rewriting complex SPARQL analytical queries for efficient cloud-based processing
abstract
Many emerging Semantic Web applications combine and aggregate data across domains for analysis. Such analytical queries compute aggregates over multiple groupings of data, resulting in query plans with complex grouping-aggregation constraints. In the context of an RDF analytical query, each such grouping maps to a graph pattern subquery with multiple join operations, and related groups often result in overlapping graph patterns within the same query. In this paper, we propose a holistic approach to optimize RDF analytical queries by refactoring queries to achieve shared execution of common subexpressions that enables parallel evaluation of groupings as well as aggregations. Such a rewriting enables shorter execution workflows, particularly beneficial for scale-out processing on distributed Cloud systems with multiple I/O phases. Experiments on real-world and synthetic benchmarks confirm that such a rewriting can achieve more efficient execution plans when compared to relational-style SPARQL query plans executed on popular Cloud systems.
Padmashree Ravindra, HyeongSik Kim 0001, Kemafor Anyanwu
IEEE BigData3
2015 Scaling Unbound-Property Queries on Big RDF Data Warehouses using MapReduce
Padmashree Ravindra, Kemafor Anyanwu
EDBT2
2014 SYRql: A Dataflow Language for Large Scale Processing of RDF Data
Fadi Maali, Padmashree Ravindra, Kemafor Anyanwu, Stefan Decker
ISWC (1)3
2014 Nesting Strategies for Enabling Nimble MapReduce Dataflows for Large RDF Data
abstract
Graph and semi-structured data are usually modeled in relational processing frameworks as “thin” relations (node, edge, node) and processing such data involves a lot of join operations. Intermediate results of joins with multi-valued attributes or relationships, contain redundant subtuples due to repetition of single-valued attributes. The amount of redundant content is high for real-world multi-valued relationships in social network (millions of Twitter followers of popular celebrities) or biological (multiple references to related proteins) datasets. In MapReduce-based platforms such as Apache Hive and Pig, redundancy in intermediate results contributes avoidable costs to the overall I/O, sorting, and network transfer overhead of join-intensive workloads due to longer workflows. Consequently, providing techniques for dealing with such redundancy will enable more nimble execution of such workflows. This paper argues for the use of a nested data model for representing intermediate data concisely using nesting-aware dataflow operators that allow for lazy and partial unnesting strategies. This approach reduces the overall I/O and network footprint of a workflow by concisely representing intermediate results during most of a workflow's execution, until complete unnesting is absolutely necessary. The proposed strategies are integrated into Apache Pig and experimental evaluation over real-world and synthetic benchmark datasets confirms their superiority over relational-style MapReduce systems such as Apache Pig and Hive.
Padmashree Ravindra, Kemafor Anyanwu
Int. J. Semantic Web Inf. Syst.2
2013 Scaling concurrency of personalized Semantic search over Large RDF data
abstract
Recent keyword search techniques on Semantic Web are moving away from shallow, information retrieval-style approaches that merely find “keyword matches” towards more interpretive approaches that attempt to induce structure from keyword queries. The process of query interpretation is usually guided by structures in data, and schema and is often supported by a graph exploration procedure. However, graph exploration-based interpretive techniques are impractical for multi-tenant scenarios for large databases because separate expensive graph exploration states need to be maintained for different user queries. This leads to significant memory overhead in situations of large numbers of concurrent requests. This limitation could negatively impact the possibility of achieving the ultimate goal of personalizing search. In this paper, we propose a lightweight interpretation approach that employs indexing to improve throughput and concurrency with much less memory overhead. It is also more amenable to distributed or partitioned execution. The approach is implemented in a system called “SKI” and an experimental evaluation of SKI's performance on the DBPedia and Billion Triple Challenge datasets shows orders-of-magnitude performance improvement over existing techniques.
Haizhou Fu, HyeongSik Kim 0001, Kemafor Anyanwu
IEEE BigData3
2013 Optimizing queries over semantically integrated datasets on MapReduce platforms
abstract
Life science databases generally consist of multiple heterogeneous datasets that have been integrated using complex ontologies. Querying such databases typically involves complex graph patterns, and evaluating such patterns poses challenges when MapReduce-based platforms are used to scale up processing, translating to long execution workflows with large amount of disk and network I/O costs. In this poster, we focus on optimizing UNION queries (e.g., unions of conjunctives for inference) and present an algebraic interpretation of the query rewritings which are more amenable to efficient processing on MapReduce.
HyeongSik Kim 0001, Kemafor Anyanwu
IEEE BigData2
2013 PrefixSolve: efficiently solving multi-source multi-destination path queries on RDF graphs by sharing suffix computations
abstract
Uncovering the "nature" of the connections between a set of entities e.g. passengers on a flight and organizations on a watchlist can be viewed as a Multi-Source Multi-Destination (MSMD) Path Query problem on labeled graph data models such as RDF. Using existing graph-navigational path finding techniques to solve MSMD problems will require queries to be decomposed into multiple single-source or destination path subqueries, each of which is solved independently. Navigational techniques on disk-resident graphs typically generate very poor I/O access patterns for large, disk-resident graphs and for MSMD path queries, such poor access patterns may be repeated if common graph exploration steps exist across subqueries.
Sidan Gao, Kemafor Anyanwu
WWW2
2012 Scan-Sharing for Optimizing RDF Graph Pattern Matching on MapReduce
abstract
Recently, the number and size of RDF data collections has increased rapidly making the issue of scalable processing techniques crucial. The MapReduce model has become a de facto standard for large scale data processing using a cluster of machines in the cloud. Generally, RDF query processing creates join-intensive workloads, resulting in lengthy MapReduce workflows with expensive I/O, data transfer, and sorting costs. However, the MapReduce computation model provides limited static optimization techniques used in relational databases (e.g., indexing and cost-based optimization). Consequently, dynamic optimization techniques for such join-intensive tasks on MapReduce need to be investigated. In some previous efforts, we propose a Nested Triple Group data model and Algebra (NTGA) for efficient graph pattern query processing in the cloud. Here, we extend this work with a scan-sharing technique that is used to optimize the processing of graph patterns with repeated properties. Specifically, our scan-sharing technique eliminates the need for repeated scanning of input relations when properties are used repeatedly in graph patterns. A formal foundation underlying this scan sharing technique is discussed as well as an implementation strategy that has been integrated in the Apache Pig framework is presented. We also present a comprehensive evaluation demonstrating performance benefits of our NTGA plus scan-sharing approach.
HyeongSik Kim 0001, Padmashree Ravindra, Kemafor Anyanwu
IEEE CLOUD3
2012 HIP: Information Passing for Optimizing Join-Intensive Data Processing Workloads on Hadoop
Seokyong Hong, Kemafor Anyanwu
DEXA (2)2
2011 Efficiently Evaluating Skyline Queries on RDF Databases
Sidan Gao, Kemafor Anyanwu
ESWC (2)3
2011 An Intermediate Algebra for Optimizing RDF Graph Pattern Matching on MapReduce
Padmashree Ravindra, HyeongSik Kim 0001, Kemafor Anyanwu
ESWC (2)3
2011 Effectively Interpreting Keyword Queries on RDF Databases with a Rear View
Haizhou Fu, Kemafor Anyanwu
ISWC (1)2
2011 From SPARQL to MapReduce: The Journey Using a Nested TripleGroup Algebra
HyeongSik Kim 0001, Padmashree Ravindra, Kemafor Anyanwu
Proc. VLDB Endow.3
2010 Scheduling Hadoop Jobs to Meet Deadlines
abstract
User constraints such as deadlines are important requirements that are not considered by existing cloud-based data processing environments such as Hadoop. In the current implementation, jobs are scheduled in FIFO order by default with options for other priority based schedulers. In this paper, we extend real time cluster scheduling approach to account for the two-phase computation style of MapReduce. We develop criteria for scheduling jobs based on user specified deadline constraints and discuss our implementation and preliminary evaluation of a Deadline Constraint Scheduler for Hadoop which ensures that only jobs whose deadlines can be met are scheduled for execution.
Kamal Kc, Kemafor Anyanwu
CloudCom2
2010 An Agglomerative Query Model for Discovery in Linked Data: Semantics and Approach
abstract
Data on the Web is increasingly being used for discovery and exploratory tasks. Unlike traditional fact-finding tasks that require only the typical single-query and response paradigm, these tasks involve a multistage search process in which bits of information are accumulated over a series of related queries. The ability and effectiveness of users to connect the dots between these pieces of information are crucial to enable discovery. In this paper, we introduce the notion of agglomerative querying for supporting "search processes" and present its motivation, challenges and formalization. We focus on a specific class of agglomerative querying called association agglomerative querying which is very natural for linked data models such as RDF. We present a preliminary implementation approach for processing such queries and discuss its relationship with SPARQL query processing. Finally, we present empirical results for proving the effectiveness of our approach on the DBLP dataset and future directions.
Sidan Gao, Haizhou Fu, Kemafor Anyanwu
WebDB3
2009 RAPID: Enabling Scalable Ad-Hoc Analytics on the Semantic Web
Radhika Sridhar, Padmashree Ravindra, Kemafor Anyanwu
ISWC3
2008 Graph Summaries for Subgraph Frequency Estimation
Angela Maduko, Kemafor Anyanwu, Amit P. Sheth, Paul Schliekelman
ESWC2
2007 SPARQ2L: towards support for subgraph extraction queries in rdf databases
abstract
Many applications in analytical domains often have the need to "connect the dots" i.e., query about the structure of data. In bioinformatics for example, it is typical to want to query about interactions between proteins. The aim of such queries is to "extract" relationships between entities i.e. paths from a data graph. Often, such queries will specify certain constraints that qualifying results must satisfy e.g. paths involving a set of mandatory nodes. Unfortunately, most present day Semantic Web query languages including the current draft of the anticipated recommendation SPARQL, lack the ability to express queries about arbitrary path structures in data. In addition, many systems that support some limited form of path queries rely on main memory graph algorithms limiting their applicability to very large scale graphs. In this paper, we present an approach for supporting Path Extraction queries. Our proposal comprises (i) a query language SPARQ2L which extends SPARQL with path variables and path variable constraint expressions, and (ii) a novel query evaluation framework based on efficient algebraic techniques for solving path problems which allows for path queries to be efficiently evaluated on disk resident RDF graphs. The effectiveness of our proposal is demonstrated by a performance evaluation of our approach on both real world based and synthetic dataset.
Kemafor Anyanwu, Angela Maduko, Amit P. Sheth
WWW1
2007 Estimating the cardinality of RDF graph patterns
abstract
Most RDF query languages allow for graph structure search through a conjunction of triples which is typically processed using join operations. A key factor in optimizing joins is determining the join order which depends on the expected cardinality of intermediate results. This work proposes a pattern-based summarization framework for estimating the cardinality of RDF graph patterns. We present experiments on real world and synthetic datasets which confirm the feasibility of our approach.
Angela Maduko, Kemafor Anyanwu, Amit P. Sheth, Paul Schliekelman
WWW2
2005 SemRank: ranking complex relationship search results on the semantic web
abstract
While the idea that querying mechanisms for complex relationships (otherwise known as Semantic Associations) should be integral to Semantic Web search technologies has recently gained some ground, the issue of how search results will be ranked remains largely unaddressed. Since it is expected that the number of relationships between entities in a knowledge base will be much larger than the number of entities themselves, the likelihood that Semantic Association searches would result in an overwhelming number of results for users is increased, therefore elevating the need for appropriate ranking schemes. Furthermore, it is unlikely that ranking schemes for ranking entities (documents, resources, etc.) may be applied to complex structures such as Semantic Associations.In this paper, we present an approach that ranks results based on how predictable a result might be for users. It is based on a relevance model SemRank, which is a rich blend of semantic and information-theoretic techniques with heuristics that supports the novel idea of modulative searches, where users may vary their search modes to effect changes in the ordering of results depending on their need. We also present the infrastructure used in the SSARK system to support the computation of SemRank values for resulting Semantic Associations and their ordering.
Kemafor Anyanwu, Angela Maduko, Amit P. Sheth
WWW1
2005 Semantic Association Identification and Knowledge Discovery for National Security Applications
abstract
Public and private organizations have access to a vast amount of internal, deep Web and open Web information. Transforming this heterogeneous and distributed information into actionable and insightful information is the key to the emerging new classes of business intelligence and national security applications. Although the role of semantics in search and integration has been often talked about, in this paper we discuss semantic approaches to support analytics on vast amounts of heterogeneous data. In particular, we bring together novel academic research and commercialized Semantic Web technology. The academic research related to semantic association identification is built upon commercial Semantic Web technology for semantic metadata extraction. A prototypical demonstration of this research and technology is presented in the context of an aviation security application of significance to national security.
Amit P. Sheth, Boanerges Aleman-Meza, Ismailcem Budak Arpinar, Clemens Bertram, Yashodhan S. Warke, Cartic Ramakrishnan, Christian Halaschek-Wiener, Kemafor Anyanwu, David Avant, Fatma Sena Arpinar, Krys J. Kochut
J. Database Manag.8
2003 ?-Queries: enabling querying for semantic associations on the semantic web
abstract
This paper presents the notion of Semantic Associations as complex relationships between resource entities. These relationships capture both a connectivity of entities as well as similarity of entities based on a specific notion of similarity called r-isomorphism. It formalizes these notions for the RDF data model, by introducing a notion of a Property Sequence as a type. In the context of a graph model such as that for RDF, Semantic Associations amount to specific certain graph signatures. Specifically, they refer to sequences (i.e. directed paths) here called Property Sequences, between entities, networks of Property Sequences (i.e. undirected paths), or subgraphs of r-isomorphic Property Sequences.The ability to query about the existence of such relationships is fundamental to tasks in analytical domains such as national security and business intelligence, where tasks often focus on finding complex yet meaningful and obscured relationships between entities. However, support for such queries is lacking in contemporary query systems, including those for RDF.
Kemafor Anyanwu, Amit P. Sheth
WWW1