Jeff Heflin

dblp:94/1154 · DBLP profile ↗
← Back
47ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-7290-1495ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 37 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 23 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Computer networks · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 An Evaluation of Approaches to Train Embeddings for Logical Inference (Student Abstract)
abstract
Knowledge bases traditionally require manual optimization to ensure reasonable performance when answering queries. We build on previous neurosymbolic approaches by improving the training of an embedding model for logical statements that maximizes similarity between unifying atoms and minimizes similarity of non-unifying atoms. In particular, we evaluate different approaches to training this model.
Yasir White, Jevon Lipsey, Jeff Heflin
AAAI3
2025 Efficient Image Similarity Search with Quadtrees (Student Abstract)
abstract
In this paper, we present a new image similarity search algorithm designed to enhance traditional information retrieval(IR) by adding an image search capability. Our approach uses a quadtree data structure to organize image data, significantly reducing search space and improving retrieval efficiency. We describe an indexing strategy and two query algorithms that can be implemented in any IR system. We tested our method on a 70K material microscopy image dataset, achieving a 25 times improvement in retrieval speed with only a 20% reduction in ranking accuracy.
Jeff Heflin
AAAI2
2025 High Quality Embeddings for Horn Logic Reasoning
abstract
Neural networks can be trained to rank the choices made by logical reasoners, resulting in more efficient searches for answers. A key step in this process is creating useful embeddings, i.e., numeric representations of logical statements. This paper introduces and evaluates several approaches to creating embeddings that result in better downstream results. We train embeddings using triplet loss, which requires examples consisting of an anchor, a positive example, and a negative example. We introduce three ideas: generating anchors that are more likely to have repeated terms, generating positive and negative examples in a way that ensures a good balance between easy, medium, and hard examples, and periodically emphasizing the hardest examples during training. We conduct several experiments to evaluate this approach, including a comparison of different embeddings across different knowledge bases, in an attempt to identify what characteristics make an embedding well-suited to a particular reasoning task.
Yasir White, Dean Clark, Joseph Sanchez, Jevon Lipsey, Ashely Hirst, Jeff Heflin
NeSy7
2023 An Evaluation of Strategies to Train More Efficient Backward-Chaining Reasoners
abstract
Knowledge bases traditionally require manual optimization to ensure reasonable performance when answering queries. We build on previous work on training a deep learning model to learn heuristics for answering queries by comparing different representations of the sentences contained in knowledge bases. We decompose the problem into issues of representation, training, and control and propose solutions for each subproblem. We evaluate different configurations on three synthetic knowledge bases. In particular we compare a novel representation approach based on learning to maximize similarity of logical atoms that unify and minimize similarity of atoms that do not unify, to two vectorization strategies taken from the automated theorem proving literature: a chain-based and a 3-term-walk strategy. We also evaluate the efficacy of pruning the search by ignoring rules with scores below a threshold.
Yue-Bo Jia, Gavin Johnson, Alex Arnold, Jeff Heflin
K-CAP4
2022 DAME: Domain Adaptation for Matching Entities
abstract
Entity matching (EM) identifies data records that refer to the same real-world entity. Despite the effort in the past years to improve the performance in EM, the existing methods still require a huge amount of labeled data in each domain during the training phase. These methods treat each domain individually, and capture the specific signals for each dataset in EM, and this leads to overfitting on just one dataset. The knowledge that is learned from one dataset is not utilized to better understand the EM task in order to make predictions on the unseen datasets with fewer labeled samples. In this paper, we propose a new domain adaptation-based method that transfers the task knowledge from multiple source domains to a target domain. Our method presents a new setting for EM where the objective is to capture the task-specific knowledge from pretraining our model using multiple source domains, then testing our model on a target domain. We study the zero-shot learning case on the target domain, and demonstrate that our method learns the EM task and transfers knowledge to the target domain. We extensively study fine-tuning our model on the target dataset from multiple domains, and demonstrate that our model generalizes better than state-of-the-art methods in EM.
Mohamed Trabelsi 0003, Jeff Heflin
WSDM2
2022 StruBERT: Structure-aware BERT for Table Search and Matching
abstract
A table is composed of data values that are organized in rows and columns providing implicit structural information. A table is usually accompanied by secondary information such as the caption, page title, etc., that form the textual information. Understanding the connection between the textual and structural information is an important, yet neglected aspect in table retrieval, as previous methods treat each source of information independently. In this paper, we propose StruBERT, a structure-aware BERT model that fuses the textual and structural information of a data table to produce context-aware representations for both textual and tabular content of a data table. We introduce the concept of horizontal self-attention, which extends the idea of vertical self-attention introduced in TaBERT and allows us to treat both dimensions of a table equally. StruBERT features are integrated in a new end-to-end neural ranking model to solve three table-related downstream tasks: keyword- and content-based table retrieval, and table similarity. We evaluate our approach using three datasets, and we demonstrate substantial improvements in terms of retrieval and classification metrics over state-of-the-art methods.
Mohamed Trabelsi 0003, Zhiyu Chen 0001, Shuo Zhang 0006, Brian D. Davison 0001, Jeff Heflin
WWW5
2021 MGNETS: Multi-Graph Neural Networks for Table Search
abstract
Table search aims to retrieve a list of tables given a user's query. Previous methods only consider the textual information of tables and the structural information is rarely used. In this paper, we propose to model the complex relations in the table corpus as one or more graphs and then utilize graph neural networks to learn representations of queries and tables. We show that the text-based table retrieval methods can be further improved by graph-based predictions which fuse multiple field-level information.
Zhiyu Chen 0001, Mohamed Trabelsi 0003, Jeff Heflin, Dawei Yin 0001, Brian D. Davison 0001
CIKM3
2021 SeLaB: Semantic Labeling with BERT
abstract
Generating schema labels automatically for column values of data tables has many data science applications such as schema matching, and data discovery and linking. For example, automatically extracted tables with missing headers can be filled by the predicted schema labels which significantly reduces human effort. Furthermore, the predicted labels can reduce the impact of inconsistent names across multiple data tables. In this paper, we propose a context-aware semantic labeling method using both data values and contextual information of columns. Our proposed method is based on formulating the semantic labeling task as a structured prediction problem, where we sequentially predict labels for an input table with missing headers. We incorporate both the values and context of each data column using the pre-trained contextualized language model, BERT. To our knowledge, we are the first to successfully adapt BERT to solve the semantic labeling task. We evaluate our approach using two real-world datasets from different domains, and we demonstrate substantial improvements in terms of evaluation metrics over state-of-the-art feature-based methods.
Mohamed Trabelsi 0003, Jeff Heflin
IJCNN3
2021 Neural ranking models for document retrieval
abstract
Abstract Ranking models are the main components of information retrieval systems. Several approaches to ranking are based on traditional machine learning algorithms using a set of hand-crafted features. Recently, researchers have leveraged deep learning models in information retrieval. These models are trained end-to-end to extract features from the raw data for ranking tasks, so that they overcome the limitations of hand-crafted features. A variety of deep learning models have been proposed, and each model presents a set of neural network components to extract features that are used for ranking. In this paper, we compare the proposed models in the literature along different dimensions in order to understand the major contributions and limitations of each model. In our discussion of the literature, we analyze the promising neural components, and propose future research directions. We also show the analogy between document retrieval and other retrieval tasks where the items to be ranked are structured documents, answers, images and videos.
Mohamed Trabelsi 0003, Zhiyu Chen 0001, Brian D. Davison 0001, Jeff Heflin
Inf. Retr. J.4
2020 An Exploratory Interface for Dataset Repositories Using Cell-Centric Indexing
abstract
Large collections of datasets are being published on the Web at an increasing rate. This poses a problem to researchers and data journalists who must sift through these large quantities of data to find datasets that meet their needs. Our solution to this problem is cell-centric indexing, a novel approach which considers the individual cell of a dataset to be the fundamental unit of search, indexing the corresponding metadata to each individual cell. This facilitates a new style of user interface that allows users to explore the collection via histograms that show the distributions of various terms organized by how they are used in the dataset.
Drake Johnson, Keith Register, Brian D. Davison 0001, Jeff Heflin
IEEE BigData4
2020 A Hybrid Deep Model for Learning to Rank Data Tables
abstract
We address the problem of ad hoc table retrieval via a new neural architecture that incorporates both semantic and relevance matching. Understanding the connection between the structured form of a table and query tokens is an important yet neglected problem in information retrieval. We use a learning- to-rank approach to train a system to capture semantic and relevance signals within interactions between the structured form of candidate tables and query tokens. Convolutional filters that extract contextual features from query/table interactions are combined with a feature vector based on the distributions of term similarity between queries and tables. We propose using row and column summaries to incorporate table content into our new neural model. We evaluate our approach using two datasets, and we demonstrate substantial improvements in terms of retrieval metrics over state-of-the-art methods in table retrieval and document retrieval, and neural architectures from sentence, document, and table type classification adapted to the table retrieval task. Our ablation study supports the importance of both semantic and relevance matching in the table retrieval.
Mohamed Trabelsi 0003, Zhiyu Chen 0001, Brian D. Davison 0001, Jeff Heflin
IEEE BigData4
2020 Relational Graph Embeddings for Table Retrieval
abstract
Ad hoc table retrieval is the problem of identifying the most relevant datasets to a user's query. We present an approach to the problem that builds a knowledge graph by combining information about the collection of tables with external sources such as WordNet and pretrained Glove embeddings. We apply multi-relational graph convolutional networks to learn embeddings for the knowledge graph nodes and utilize three different methods to create vectors representing the tables and queries from these embeddings. We create a novel learning-to-rank neural architecture that incorporates the multiple embeddings in order to improve table retrieval results. We evaluate our approach using two large collections of tables from public WikiTables and Web tables data, demonstrating substantial improvements over state-of-the-art methods in table retrieval.
Mohamed Trabelsi 0003, Zhiyu Chen 0001, Brian D. Davison 0001, Jeff Heflin
IEEE BigData4
2020 Leveraging Schema Labels to Enhance Dataset Search
Zhiyu Chen 0001, Haiyan Jia, Jeff Heflin, Brian D. Davison 0001
ECIR (1)3
2020 Table Search Using a Deep Contextualized Language Model
abstract
Pretrained contextualized language models such as BERT have achieved impressive results on various natural language processing benchmarks. Benefiting from multiple pretraining tasks and large scale training corpora, pretrained models can capture complex syntactic word relations. In this paper, we use the deep contextualized language model BERT for the task of ad hoc table retrieval. We investigate how to encode table content considering the table structure and input length limit of BERT. We also propose an approach that incorporates features from prior literature on table retrieval and jointly trains them with BERT. In experiments on public datasets, we show that our best approach can outperform the previous state-of-the-art method and BERT baselines with a large margin under different evaluation metrics.
Zhiyu Chen 0001, Mohamed Trabelsi 0003, Jeff Heflin, Yinan Xu 0005, Brian D. Davison 0001
SIGIR3
2019 Improved Table Retrieval Using Multiple Context Embeddings for Attributes
abstract
Table retrieval is the task of extracting the most relevant tables to answer a user's query. Table retrieval is an important task because many domains have tables that contain useful information in a structured form. Given a user's query, the goal is to obtain a relevance ranking for query-table pairs, such that higher ranked tables should be more relevant to the query. In this paper, we present a context-aware table retrieval method that is based on a novel embedding for attribute tokens. We find that differentiated types of contexts are useful in building word embeddings. We also find that including a specialized representation of numerical cell values in our model improves table retrieval performance. We use the trained model to predict different contexts of every table. We show that the predicted contexts are useful in ranking tables against a query using a multi-field ranking approach. We evaluate our approach using public WikiTables data, and we demonstrate improvements in terms of NDCG over unsupervised baseline methods in the table retrieval task.
Mohamed Trabelsi 0003, Brian D. Davison 0001, Jeff Heflin
IEEE BigData3
2017 Linking Heterogeneous Data in the Semantic Web Using Scalable and Domain-Independent Candidate Selection
abstract
Due to the decentralized nature of the Semantic Web, the same real-world entity may be described in various data sources with different ontologies and assigned syntactically distinct identifiers. In order to facilitate data utilization and consumption in the Semantic Web, without compromising the freedom of people to publish their data, one critical problem is to appropriately interlink such heterogeneous data. This interlinking process is sometimes referred to as Entity Matching, i.e., finding which identifiers refer to the same real-world entity. In this paper, we propose two candidate selection algorithms to improve the scalability of entity matching systems. First of all, we propose HistSim that utilizes the matching histories of the instances to prune instance pairs that are not sufficiently similar to the same pool of other instances. A sigmoid function based thresholding method is proposed to automatically adjust the threshold for such commonality on-the-fly. Furthermore, we propose DisNGram that selects candidate instance pairs by computing a character-level similarity metric on discriminating literal values that are chosen using domain-independent unsupervised learning. Instances are indexed on the chosen predicates' literal values to enable efficient look-up for similar instances. Finally, in order to be able to handle heterogeneous datasets with a large number of predicates, a mechanism for automatically determining predicate comparability is proposed. We evaluate our two candidate selection algorithms against six state-of-the-art systems on three Semantic Web datasets, and demonstrate that our proposed algorithms frequently outperform state-of-the-art systems on F1-score and runtime.
Dezhao Song, Jeff Heflin
IEEE Trans. Knowl. Data Eng.3
2016 Ontology Instance Linking: Towards Interlinked Knowledge Graphs
abstract
Due to the decentralized nature of the Semantic Web, the same real-world entity may be described in various data sources with different ontologies and assigned syntactically distinct identifiers. In order to facilitate data utilization and consumption in the Semantic Web, without compromising the freedom of people to publish their data, one critical problem is to appropriately interlink such heterogeneous data. This interlinking process is sometimes referred to as Entity Coreference, i.e., finding which identifiers refer to the same real-world entity. In this paper, we first summarize state-of-the-art algorithms in detecting such coreference relationships between ontology instances. We then discuss various techniques in scaling entity coreference to large-scale datasets. Finally, we present well-adopted evaluation datasets and metrics, and compare the performance of the state-of-the-art algorithms on such datasets.
Jeff Heflin, Dezhao Song
AAAI1
2015 Multimodal Entity Coreference for Cervical Dysplasia Diagnosis
abstract
Cervical cancer is the second most common type of cancer for women. Existing screening programs for cervical cancer, such as Pap Smear, suffer from low sensitivity. Thus, many patients who are ill are not detected in the screening process. Using images of the cervix as an aid in cervical cancer screening has the potential to greatly improve sensitivity, and can be especially useful in resource-poor regions of the world. In this paper, we develop a data-driven computer algorithm for interpreting cervical images based on color and texture. We are able to obtain 74% sensitivity and 90% specificity when differentiating high-grade cervical lesions from low-grade lesions and normal tissue. On the same dataset, using Pap tests alone yields a sensitivity of 37% and specificity of 96%, and using HPV test alone gives a 57% sensitivity and 93% specificity. Furthermore, we develop a comprehensive algorithmic framework based on Multimodal Entity Coreference for combining various tests to perform disease classification and diagnosis. When integrating multiple tests, we adopt information gain and gradient-based approaches for learning the relative weights of different tests. In our evaluation, we present a novel algorithm that integrates cervical images, Pap, HPV, and patient age, which yields 83.21% sensitivity and 94.79% specificity, a statistically significant improvement over using any single source of information alone.
Dezhao Song, Sharon X. Huang, Joseph Patruno, Hector Muñoz-Avila, Jeff Heflin, L. Rodney Long, Sameer K. Antani
IEEE Trans. Medical Imaging6
2014 Exploring Linked Data with contextual tag clouds
Xingjian Zhang 0006, Dezhao Song, Sambhawa Priya, Zachary A. Daniels, Kelly Reynolds, Jeff Heflin
J. Web Semant.6
2013 Infrastructure for Efficient Exploration of Large Scale Linked Data via Contextual Tag Clouds
Xingjian Zhang 0006, Dezhao Song, Sambhawa Priya, Jeff Heflin
ISWC (1)4
2013 Using Instance Texts to Improve Keyword-Based Class Retrieval
abstract
In this paper we investigate the keyword based class retrieval problem, which we define as how to identify ontological classes that best match a keyword based query. Most previous applications use simple syntactic matching approaches on the class labels and/or comments, or expand the keyword query by using lexicons such as Word Net, but fail to retrieve relevant resources in many scenarios. Instead of relying on external sources, we investigate this problem by using the annotations of instances associated with classes in the knowledge base. We propose a general framework of this approach, which consists of two phases: the keyword query is first used to locate relevant instances, then we induce the classes given this list of weighted matched instances. If we identify sufficient text for the instances, then the first phase can be solved by a traditional information retrieval (IR) query, however the second phase might be cast in different ways: as an additive value function, as an IR problem with instance as queries, or as an instance-based ontology alignment problem. With many applicable strategies initiated from different viewpoints, we find that some of them are mathematically equivalent or very similar. In the experiments we compare our proposed framework to simple syntactic approaches and evaluate different strategies.
Xingjian Zhang 0006, Jeff Heflin
Web Intelligence2
2013 On a steady path to semantic technology evaluation
Raúl García-Castro, Stuart N. Wrigley, Jeff Heflin, Heiner Stuckenschmidt
J. Web Semant.3
2012 An ontology-based system to identify complex network attacks
abstract
Intrusion Detection Systems are tools used to detect attacks against networks. Many of these attacks are a sequence of multiple simple attacks. These complex attacks are more difficult to identify because (a) they are difficult to predict, (b) almost anything could be an attack, and (c) there are a huge number of possibilities. The problem is that the expertise of what constitutes an attack lies in the tacit knowledge of experienced network engineers. By providing an ontological representation of what constitutes a network attack human expertise to be codified and tested. The details of this representation are explained. An implementation of the representation has been developed. Lastly, the use of the representation in an Intrusion Detection System for complex attack detection has been demonstrated using use cases.
Lisa Frye, Liang Cheng 0001, Jeff Heflin
ICC3
2012 Accuracy vs. Speed: Scalable Entity Coreference on the Semantic Web with On-the-Fly Pruning
abstract
One challenge for the Semantic Web is to scalably establish high quality owl: same As links between co referent ontology instances in different data sources, traditional approaches that exhaustively compare every pair of instances do not scale well to large datasets. In this paper, we propose a pruning-based algorithm for reducing the complexity of entity co reference. First, we discard candidate pairs of instances that are not sufficiently similar to the same pool of other instances. A sigmoid function based thresholding method is proposed to automatically adjust the threshold for such commonality on-the-fly. In our prior work, each instance is associated with a context graph consisting of neighboring RDF nodes. In this paper, we speed up the comparison for a single pair of instances by pruning insignificant context in the graph, this is accomplished by evaluating its potential contribution to the final similarity measure. We evaluate our system on three Semantic Web instance categories. We verify the effectiveness of our thresholding and context pruning methods by comparing to nine state-of-the-art systems. We show that our algorithm frequently outperforms those systems with a runtime speedup factor of 18 to 24 while maintaining competitive F1-scores. For datasets of up to 1 million instances, this translates to as much as 370 hours improvement in runtime.
Dezhao Song, Jeff Heflin
Web Intelligence2
2012 Web-scale semantic information processing
Jeff Heflin, Heiner Stuckenschmidt
J. Web Semant.1
2011 Finding VIPs - A visual image persons search using a content property reasoner and web ontology
abstract
We present a semantic based search tool, VIPs, i.e. Visual Image Persons Search, on the domain of VIPs, i.e. very important people. Our tool explores the possibilities of content based image search supported by ontological reasoning. Our framework integrates information from both image processing algorithms and semantic knowledge bases to perform interesting queries that would otherwise be impossible. We describe a novel property reasoner that is able to translate low level image features into semantically relevant object properties. Finally, we demonstrate interesting searches supported by our framework on the domain of people, the majority of whom are movie celebrities, using the properties translated by our system as well as existing ontologies available on the web.
Sharon X. Huang, Jeff Heflin
ICME3
2011 Learning to detect abnormal semantic web data
abstract
No abstract available.
Yang Yu 0029, Xingjian Zhang 0006, Jeff Heflin
K-CAP3
2011 Automatically Generating Data Linkages Using a Domain-Independent Candidate Selection Approach
Dezhao Song, Jeff Heflin
ISWC (1)2
2011 Extending Functional Dependency to Detect Abnormal Data in RDF Graphs
Yang Yu 0029, Jeff Heflin
ISWC (1)2
2010 Query optimization for ontology-based information integration
abstract
In recent years, there has been an explosion of publicly avail-able RDF and OWL data sources. In order to effectively and quickly answer queries in such an environment, we present an approach to identifying the potentially relevant Semantic Web data sources using query rewritings and a term index. We demonstrate that such an approach must carefully han-dle query goals that lack constants; otherwise the algorithm may identify many sources that do not contribute to even-tual answers. This is because the term index only indicates if URIs are present in a document, and specific answers to a subgoal cannot be calculated until the source is physi-cally accessed- an expensive operation given disk/network latency. We present an algorithm that, given a set of query rewritings that accounts for ontology heterogeneity, incre-mentally selects and processes sources in order to maintain selectivity. Once sources are selected, we use an OWL rea-soner to answer queries over these sources and their corre-sponding ontologies. We present the results of experiments using both a synthetic data set and a subset of the real-world Billion Triple Challenge data.
Yingjie Li 0004, Jeff Heflin
CIKM2
2010 Domain-independent entity coreference in RDF graphs
abstract
In this paper, we present a novel entity coreference algorithm for Semantic Web instances. The key issues include how to locate context information and how to utilize the context appropriately. To collect context information, we select a neighborhood (consisting of triples) of each instance from the RDF graph. To determine the similarity between two instances, our algorithm computes the similarity between comparable property values in the neighborhood graphs. The similarity of distinct URIs and blank nodes is computed by comparing their outgoing links. To provide the best possible domain-independent matches, we examine an appropriate way to compute the discriminability of triples. To reduce the impact of distant nodes, we explore a distance-based discounting approach. We evaluated our algorithm using different instance categories in two datasets. Our experiments show that the best results are achieved by including both our triple discrimination and discounting approaches.
Dezhao Song, Jeff Heflin
CIKM2
2010 Using Reformulation Trees to Optimize Queries over Distributed Heterogeneous Sources
Yingjie Li 0004, Jeff Heflin
ISWC (1)2
2010 A Scalable Indexing Mechanism for Ontology-Based Information Integration
abstract
In recent years, there has been an explosion of publicly available RDF and OWL web pages. Typically, these pages are small, heterogeneous and prone to change frequently. In order to effectively integrate them, we propose to adapt a query reformulation algorithm and combine it with an information retrieval inspired index in order to select all sources relevant to a query. We treat each RDF document as a bag of URIs and literals and build an inverted index. Our system first reformulates the user’s query into a set of sub goals and then translates these into Boolean queries against the index in order to determine which sources are relevant. Finally, the selected data sources and the relevant ontology mappings are used in conjunction with a description logic reasoner to provide an efficient query answering solution for the Semantic Web. We have evaluated our system using ontology mappings and ten million real world data sources.
Yingjie Li 0004, Abir Qasem, Jeff Heflin
Web Intelligence3
2009 A Case Study in Integrating Multiple E-commerce Standards via Semantic Web Technology
Yang Yu 0029, Donald Hillman, Basuki Setio, Jeff Heflin
ISWC4
2008 DLDB2: A Scalable Multi-perspective Semantic Web Repository
abstract
A true semantic Web repository must scale both in terms of number of ontologies and quantity of data. It should also support reasoning using different points of view about the meanings and relationships of concepts and roles. Our DLDB2 system has these features. Our system is sound and complete on a sizable subset of description horn logic when answering extensional conjunctive queries, but more importantly also computes many entailments from OWL DL. By delegating TBox reasoning to a DL reasoner, we focus on the design of the table schema, database views, and algorithms that achieve essential ABox reasoning over an RDBMS. We evaluate the system using synthetic benchmarks as well as real-world data and queries.
Zhengxiang Pan, Xingjian Zhang 0006, Jeff Heflin
Web Intelligence3
2008 Goal Node Search for Semantic Web Source Selection
abstract
We present an efficient search approach for selecting all potentially relevant data sources for a conjunctive Semantic Web query. We use map ontologies to align heterogeneous domain ontologies. This allows us to select data sources that may be relevant to the query but generally do not describe their data directly in terms of the ontology of the query. The "goal node search" algorithm is a significant improvement on our original source selection algorithm. The new algorithm allows a more expressive knowledge representation language to describe domain ontologies and it is about three times more efficient than the original source selection algorithm when performing similar tasks.
Abir Qasem, Dimitre A. Dimitrov, Jeff Heflin
Web Intelligence3
2007 Document-Centric Query Answering for the Semantic Web
abstract
In this paper, we propose document-centric query answering, a novel form of query answering for the Semantic Web. We discuss how we have built a knowledge base system to support the new queries. In particular, we describe the key techniques used in the system in order to address scalability issues. In addition, we show encouraging experimental results.
Jeff Heflin
Web Intelligence2
2007 A Requirements Driven Framework for Benchmarking Semantic Web Knowledge Base Systems
abstract
A key challenge for the semantic Web is to acquire the capability to effectively query large knowledge bases. As there will be several competing systems, we need benchmarks that will objectively evaluate these systems. Development of effective benchmarks in an emerging domain is a challenging endeavor. In this paper, we propose a requirements driven framework for developing benchmarks for semantic Web knowledge base systems (SW KBSs). In this paper, we make two major contributions. First, we provide a list of requirements for SW KBS benchmarks. This can serve as an unbiased guide to both the benchmark developers and personnel responsible for systems acquisition and benchmarking. Second, we provide an organized collection of techniques and tools needed to develop such benchmarks. In particular, the collection contains a detailed guide for generating benchmark workload, defining performance metrics, and interpreting experimental results
Abir Qasem, Zhengxiang Pan, Jeff Heflin
IEEE Trans. Knowl. Data Eng.4
2006 Large Scale Knowledge Base Systems: An Empirical Evaluation Perspective
Abir Qasem, Jeff Heflin
AAAI3
2006 An Investigation into the Feasibility of the Semantic Web
Zhengxiang Pan, Abir Qasem, Jeff Heflin
AAAI3
2006 Information Integration Via an End-to-End Distributed Semantic Web System
Dimitre A. Dimitrov, Jeff Heflin, Abir Qasem, Nanbor Wang
ISWC2
2005 On Logical Consequence for Collections of OWL Documents
Jeff Heflin
ISWC2
2005 Rapid Benchmarking for Semantic Web Knowledge Base Systems
Sui-Yu Wang, Abir Qasem, Jeff Heflin
ISWC4
2005 LUBM: A benchmark for OWL knowledge base systems
Zhengxiang Pan, Jeff Heflin
J. Web Semant.3
2004 An Evaluation of Knowledge Base Systems for Large OWL Datasets
Zhengxiang Pan, Jeff Heflin
ISWC3
2004 A Model Theoretic Semantics for Ontology Versioning
Jeff Heflin, Zhengxiang Pan
ISWC1
2003 Benchmarking DAML+OIL Repositories
Jeff Heflin, Zhengxiang Pan
ISWC2