VLDB 2026 Research / reviewers in the wild / expert
Oktie Hassanzadeh
dblp:h/OktieHassanzadeh
· DBLP profile ↗
45ranked-venue papers
15as first author
14since 2021 · last 2026
0000-0001-5307-9857ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 28 · 12 first-author · 6 since 2021Artificial intelligence and machine learning · 16 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When Text-to-SQL Evaluation Misleads: Rethinking Benchmarking Practices
Oktie Hassanzadeh, Yotam Perlitz, Nhan Pham, Timothy Dinger, Tanvi Kaple, Long Hai Vu, Michael R. Glass, Dharmashankar Subramanian |
ICDE | 1 |
| 2025 | TabSketchFM: Sketch-Based Tabular Representation Learning for Data Discovery Over Data LakesabstractEnterprises have a growing need to identify relevant tables in data lakes; e.g. tables that are unionable, joinable, or subsets of each other. Tabular neural models can be help-ful for such data discovery tasks. In this paper, we present TabSketchFM, a neural tabular model for data discovery over data lakes. First, we propose novel pre-training: a sketch-based approach to enhance the effectiveness of data discovery in neural tabular models. Second, we finetune the pretrained model for identifying unionable, joinable, and subset table pairs and show significant improvement over previous tabular neural models. Third, we present a detailed ablation study to highlight which sketches are crucial for which tasks. Fourth, we use these finetuned models to perform table search; i.e., given a query table, find other tables in a corpus that are unionable, joinable, or that are subsets of the query. Our results demonstrate significant improvements in F1 scores for search compared to state-of-the-art techniques. Finally, we show significant transfer across datasets and tasks establishing that our model can generalize across different tasks and over different data lakes. Aamod Khatiwada, Harsha Kokel, Ibrahim Abdelaziz, Subhajit Chaudhury, Julian Dolby, Oktie Hassanzadeh, Zhenhan Huang, Tejaswini Pedapati, Horst Samulowitz, Kavitha Srinivas |
ICDE | 6 |
| 2024 | CHRONOS: A Schema-Based Event Understanding and Prediction SystemabstractChronological and Hierarchical Reasoning Over Naturally Occurring Schemas (CHRONOS) is a system that combines language model-based natural language processing with symbolic knowledge representations to analyze and make predictions about newsworthy events. CHRONOS consists of an event-centric information extraction pipeline and a complex event schema instantiation and prediction system. Resulting predictions are detailed with arguments, event types from Wikidata, schema-based justifications, and source document provenance. We evaluate our system by its ability to capture the structure of unseen events described in news articles and make plausible predictions as judged by human annotators. Maria Chang 0001, Achille Fokoue, Rosario Uceda-Sosa, Parul Awasthy, Ken Barker 0002, Sadhana Kumaravel, Oktie Hassanzadeh, Elton F. S. Soares, Debarun Bhattacharjya, Radu Florian, Salim Roukos |
AAAI | 7 |
| 2024 | Distilling Event Sequence Knowledge From Large Language Models
Somin Wadhwa, Oktie Hassanzadeh, Debarun Bhattacharjya, Ken Barker 0002, Jian Ni |
ISWC (1) | 2 |
| 2023 | Probabilistic Attention-to-Influence Neural Models for Event SequencesabstractDiscovering knowledge about which types of events influence others, using datasets of event sequences without time stamps, has several practical applications. While neural sequence models are able to capture complex and potentially long-range historical dependencies, they often lack the interpretability of simpler models for event sequence dynamics. We provide a novel neural framework in such a setting - a probabilistic attention-to-influence neural model - which not only captures complex instance-wise interactions between events but also learns influencers for each event type of interest. Given event sequence data and a prior distribution on type-wise influence, we efficiently learn an approximate posterior for type-wise influence by an attention-to-influence transformation using variational inference. Our method subsequently models the conditional likelihood of sequences by sampling the above posterior to focus attention on influencing event types. We motivate our general framework and show improved performance in experiments compared to existing baselines on synthetic data as well as real-world benchmarks, for tasks involving prediction and influencing set identification. Xiao Shou, Debarun Bhattacharjya, Dharmashankar Subramanian, Oktie Hassanzadeh, Kristin P. Bennett |
ICML | 5 |
| 2023 | Probabilistic Rule Induction from Event Sequences with Logical Summary Markov ModelsabstractEvent sequences are widely available across application domains and there is a long history of models for representing and analyzing such datasets. Summary Markov models are a recent addition to the literature that help identify the subset of event types that influence event types of interest to a user. In this paper, we introduce logical summary Markov models, which are a family of models for event sequences that enable interpretable predictions through logical rules that relate historical predicates to the probability of observing an event type at any arbitrary position in the sequence. We illustrate their connection to prior parametric summary Markov models as well as probabilistic logic programs, and propose new models from this family along with efficient greedy search algorithms for learning them from data. The proposed models outperform relevant baselines on most datasets in an empirical investigation on a probabilistic prediction task. We also compare the number of influencers that various logical summary Markov models learn on real-world datasets, and conduct a brief exploratory qualitative study to gauge the promise of such symbolic models around guiding large language models for predicting societal events. Debarun Bhattacharjya, Oktie Hassanzadeh, Ronny Luss, Keerthiram Murugesan |
IJCAI | 2 |
| 2023 | Pairwise Causality Guided Transformers for Event SequencesabstractAlthough pairwise causal relations have been extensively studied in observational longitudinal analyses across many disciplines, incorporating knowledge of causal pairs into deep learning models for temporal event sequences remains largely unexplored. In this paper, we propose a novel approach for enhancing the performance of transformer-based models in multivariate event sequences by injecting pairwise qualitative causal knowledge such as `event Z amplifies future occurrences of event Y'. We establish a new framework for causal inference in temporal event sequences using a transformer architecture, providing a theoretical justification for our approach, and show how to obtain unbiased estimates of the proposed measure. Experimental results demonstrate that our approach outperforms several state-of-the-art models in terms of prediction accuracy by effectively leveraging knowledge about causal pairs.
We also consider a unique application where we extract knowledge around sequences of societal events by generating them from a large language model, and demonstrate how a causal knowledge graph can help with event prediction in such sequences.
Overall, our framework offers a practical means of improving the performance of transformer-based models in multivariate event sequences by explicitly exploiting pairwise causal information. Xiao Shou, Debarun Bhattacharjya, Dharmashankar Subramanian, Oktie Hassanzadeh, Kristin P. Bennett |
NeurIPS | 5 |
| 2023 | Event Prediction using Case-Based Reasoning over Knowledge GraphsabstractApplying link prediction (LP) methods over knowledge graphs (KG) for tasks such as causal event prediction presents an exciting opportunity. However, typical LP models are ill-suited for this task as they are incapable of performing inductive link prediction for new, unseen event entities and they require retraining as knowledge is added or changed in the underlying KG. We introduce a case-based reasoning model, EvCBR, to predict properties about new consequent events based on similar cause-effect events present in the KG. EvCBR uses statistical measures to identify similar events and performs path-based predictions, requiring no training step. To generalize our methods beyond the domain of event prediction, we frame our task as a 2-hop LP task, where the first hop is a causal relation connecting a cause event to a new effect event and the second hop is a property about the new event which we wish to predict. The effectiveness of our method is demonstrated using a novel dataset of newsworthy events with causal relations curated from Wikidata, where EvCBR outperforms baselines including translational-distance-based, GNN-based, and rule-based LP models. Sola S. Shirai, Debarun Bhattacharjya, Oktie Hassanzadeh |
WWW | 3 |
| 2023 | An analysis of one-to-one matching algorithms for entity resolutionabstractAbstract Entity resolution (ER) is the task of finding records that refer to the same real-world entities. A common scenario, which we refer to as Clean-Clean ER, is to resolve records across two clean sources (i.e., they are duplicate-free and contain one record per entity). Matching algorithms for Clean-Clean ER yield bipartite graphs, which are further processed by clustering algorithms to produce the end result. In this paper, we perform an extensive empirical evaluation of eight bipartite graph matching algorithms that take as input a bipartite similarity graph and provide as output a set of matched records. We consider a wide range of matching algorithms, including algorithms that have not previously been applied to ER, or have been evaluated only in other ER settings. We assess the relative performance of these algorithms with respect to accuracy and time efficiency over ten established real-world data sets, from which we generated over 700 different similarity graphs. Our results provide insights into the relative performance of these algorithms and guidelines for choosing the best one, depending on the data at hand. George Papadakis 0001, Vasilis Efthymiou, Emmanouil Thanos, Oktie Hassanzadeh, Peter Christen |
VLDB J. | 4 |
| 2022 | Bipartite Graph Matching Algorithms for Clean-Clean Entity Resolution: An Empirical Evaluation
George Papadakis 0001, Vasilis Efthymiou, Emmanouil Thanos, Oktie Hassanzadeh |
EDBT | 4 |
| 2022 | Summary Markov Models for Event SequencesabstractDatasets involving sequences of different types of events without meaningful time stamps are prevalent in many applications, for instance when extracted from textual corpora. We propose a family of models for such event sequences -- summary Markov models -- where the probability of observing an event type depends only on a summary of historical occurrences of its influencing set of event types. This Markov model family is motivated by Granger causal models for time series, with the important distinction that only one event can occur in a position in an event sequence. We show that a unique minimal influencing set exists for any set of event types of interest and choice of summary function, formulate two novel models from the general family that represent specific sequence dynamics, and propose a greedy search algorithm for learning them from event sequence data. We conduct an experimental investigation comparing the proposed models with relevant baselines, and illustrate their knowledge acquisition and discovery capabilities through case studies involving sequences from text. Debarun Bhattacharjya, Saurabh Sihag, Oktie Hassanzadeh, Liza Bialik |
IJCAI | 3 |
| 2022 | Knowledge-Based News Event Analysis and Forecasting ToolkitabstractWe present a toolkit for knowledge-based news event analysis and forecasting. The toolkit is powered by a Knowledge Graph (KG) of events curated from structured and unstructured sources of event-related knowledge. The toolkit provides functions for 1) mapping ongoing news headlines to concepts in the KG, 2) retrieval, reasoning, and visualization for causal analysis and forecasting, and 3) extraction of causal knowledge from text documents to augment the KG with additional domain knowledge. Each function has a number of implementations using a wide range of state-of-the-art neuro-symbolic techniques. We show how the toolkit enables building a human-in-the-loop explainable solution for event analysis and forecasting. Oktie Hassanzadeh, Parul Awasthy, Ken Barker 0002, Onkar Bhardwaj, Debarun Bhattacharjya, Mark Feblowitz, Lee Martie, Jian Ni, Kavitha Srinivas, Lucy Yip |
IJCAI | 1 |
| 2021 | Unsupervised Causal Knowledge Extraction from Text using Natural Language Inference (Student Abstract)abstractIn this paper, we address the problem of extracting causal knowledge from text documents in a weakly supervised manner. We target use cases in decision support and risk management, where causes and effects are general phrases without any constraints. We present a method called CaKNowLI which only takes as input the text corpus and extracts a high-quality collection of cause-effect pairs in an automated way. We approach this problem using state-of-the-art natural language understanding techniques based on pre-trained neural models for Natural Language Inference (NLI). Finally, we evaluate the proposed method on existing and new benchmark data sets. Manik Bhandari, Mark Feblowitz, Oktie Hassanzadeh, Kavitha Srinivas, Shirin Sohrabi |
AAAI | 3 |
| 2021 | IBM Scenario Planning Advisor: A Neuro-Symbolic ERM SolutionabstractScenario Planning is a commonly used Enterprise Risk Management (ERM) technique to help decision makers with longterm plans by considering multiple alternative futures. It is typically a manual, highly labor intensive process involving dozens of experts and hundreds to thousands of person-hours. We previously introduced a Scenario Planning Advisor prototype (Sohrabi et al. 2018a,b) that focuses on generating scenarios quickly based on expert-developed models. We present the evolution of that prototype into a full-scale, cloud deployed ERM solution that: (i) can automatically (through NLP) create models from authoritative documents such as books, reports and articles, such that what typically took hundreds to thousands of person-hours can now be achieved in minutes to hours; (ii) can gather news and other feeds relevant to forces in the risk models and group them into storylines without any other user input; (iii) can generate scenarios at scale, starting with dozens of forces of interest from models with thousands of forces in seconds; (iv) provides interactive visualizations of scenario and force model graphs, including a full model editor in the browser. The SPA solution is deployed under a non-commercial use license at https://spa-service.draco.res.ibm.com and includes a user guide to help new users get started. A video demonstration is available at https://www.youtube.com/watch?v=IaX3d37NUl8. Mark Feblowitz, Oktie Hassanzadeh, Michael Katz 0001, Shirin Sohrabi, Kavitha Srinivas, Octavian Udrea |
AAAI | 2 |
| 2020 | Causal Knowledge Extraction through Large-Scale Text MiningabstractIn this demonstration, we present a system for mining causal knowledge from large corpuses of text documents, such as millions of news articles. Our system provides a collection of APIs for causal analysis and retrieval. These APIs enable searching for the effects of a given cause and the causes of a given effect, as well as the analysis of existence of causal relation given a pair of phrases. The analysis includes a score that indicates the likelihood of the existence of a causal relation. It also provides evidence from an input corpus supporting the existence of a causal relation between input phrases. Our system uses generic unsupervised and weakly supervised methods of causal relation extraction that do not impose semantic constraints on causes and effects. We show example use cases developed for a commercial application in enterprise risk management. Oktie Hassanzadeh, Debarun Bhattacharjya, Mark Feblowitz, Kavitha Srinivas, Michael Perrone, Shirin Sohrabi, Michael Katz 0001 |
AAAI | 1 |
| 2020 | SemTab 2019: Resources to Benchmark Tabular Data to Knowledge Graph Matching Systems
Ernesto Jiménez-Ruiz, Oktie Hassanzadeh, Vasilis Efthymiou, Jiaoyan Chen 0001, Kavitha Srinivas |
ESWC | 2 |
| 2019 | Answering Binary Causal Questions Through Large-Scale Text Mining: An Evaluation Using Cause-Effect Pairs from Human ExpertsabstractIn this paper, we study the problem of answering questions of type "Could X cause Y?" where X and Y are general phrases without any constraints. Answering such questions will assist with various decision analysis tasks such as verifying and extending presumed causal associations used for decision making. Our goal is to analyze the ability of an AI agent built using state-of-the-art unsupervised methods in answering causal questions derived from collections of cause-effect pairs from human experts. We focus only on unsupervised and weakly supervised methods due to the difficulty of creating a large enough training set with a reasonable quality and coverage. The methods we examine rely on a large corpus of text derived from news articles, and include methods ranging from large-scale application of classic NLP techniques and statistical analysis to the use of neural network based phrase embeddings and state-of-the-art neural language models. Oktie Hassanzadeh, Debarun Bhattacharjya, Mark Feblowitz, Kavitha Srinivas, Michael Perrone, Shirin Sohrabi, Michael Katz 0001 |
IJCAI | 1 |
| 2018 | Semantic Concept Discovery over Event Databases
Oktie Hassanzadeh, Shari Trewin, Alfio Massimiliano Gliozzo |
ESWC | 1 |
| 2018 | IBM Scenario Planning Advisor: Plan Recognition as AI Planning in PracticeabstractWe present the IBM Research Scenario Planning Advisor (SPA), a decision support system that allows users to generate diverse alternate scenarios of the future and enhance their ability to imagine the different possible outcomes, including unlikely but potentially impactful futures. The system includes tooling for experts to intuitively encode their domain knowledge, and uses AI Planning to reason about this knowledge and the current state of the world, including news and social media, when generating scenarios. Shirin Sohrabi, Michael Katz 0001, Oktie Hassanzadeh, Octavian Udrea, Mark Feblowitz |
IJCAI | 3 |
| 2018 | Inducing Implicit Relations from Text Using Distantly Supervised Deep Nets
Michael R. Glass, Alfio Massimiliano Gliozzo, Oktie Hassanzadeh, Nandana Mihindukulasooriya, Gaetano Rossiello |
ISWC (1) | 3 |
| 2017 | Matching Web Tables with Knowledge Base Entities: From Entity Lookups to Entity Embeddings
Vasilis Efthymiou, Oktie Hassanzadeh, Mariano Rodriguez-Muro, Vassilis Christophides |
ISWC (1) | 2 |
| 2017 | Large-scale structural and textual similarity-based mining of knowledge graph to predict drug-drug interactions
Ibrahim Abdelaziz, Achille Fokoue, Oktie Hassanzadeh, Ping Zhang 0016, Mohammad Sadoghi |
J. Web Semant. | 3 |
| 2016 | Tiresias: Knowledge Engineering and Large-Scale Machine Learning for Interpretable Drug-Drug Interaction Prediction
Achille Fokoue, Oktie Hassanzadeh, Mohammad Sadoghi, Ping Zhang 0016 |
AMIA | 2 |
| 2016 | Towards Large-Scale Predictive Drug Safety: A Computational Framework for Inferring Drug Interactions Through Similarity-Based Link Prediction
Achille Fokoue, Ping Zhang 0016, Oktie Hassanzadeh, Mohammad Sadoghi |
AMIA | 3 |
| 2016 | Joint Learning of Local and Global Features for Entity Linking via Neural NetworksabstractPrevious studies have highlighted the necessity for entity linking systems to capture the local entity-mention similarities and the global topical coherence. We introduce a novel framework based on convolutional neural networks and recurrent neural networks to simultaneously model the local and global features for entity linking. The proposed model benefits from the capacity of convolutional neural networks to induce the underlying representations for local contexts and the advantage of recurrent neural networks to adaptively compress variable length sequences of predictions for global constraints. Our evaluation on multiple datasets demonstrates the effectiveness of the model and yields the state-of-the-art performance on such datasets. In addition, we examine the entity linking systems on the domain adaptation setting that further demonstrates the cross-domain robustness of the proposed model. Thien Huu Nguyen, Nicolas R. Fauceglia, Mariano Rodriguez-Muro, Oktie Hassanzadeh, Alfio Massimiliano Gliozzo, Mohammad Sadoghi |
COLING | 4 |
| 2016 | Finding Diverse High-Quality Plans for Hypothesis GenerationabstractIn this paper, we address the problem of finding diverse high-quality plans motivated by the hypothesis generation problem. To this end, we present a planner called TK*that first efficiently solves the “top-k” cost-optimal planning problem to find k best plans, followed by clustering to produce diverse plans as cluster representatives. Shirin Sohrabi, Anton Riabov, Octavian Udrea, Oktie Hassanzadeh |
ECAI | 4 |
| 2016 | Self-Curating Databases
Mohammad Sadoghi, Kavitha Srinivas, Oktie Hassanzadeh, Yuan-Chi Chang, Mustafa Canim, Achille Fokoue, Yishai A. Feldman |
EDBT | 3 |
| 2016 | Predicting Drug-Drug Interactions Through Large-Scale Similarity-Based Link Prediction
Achille Fokoue, Mohammad Sadoghi, Oktie Hassanzadeh, Ping Zhang 0016 |
ESWC | 3 |
| 2016 | Interactive Planning-Based Hypothesis Generation with LTS++
Shirin Sohrabi, Octavian Udrea, Anton Riabov, Oktie Hassanzadeh |
IJCAI | 4 |
| 2015 | Automatic Curation of Clinical Trials Data in LinkedCT
Oktie Hassanzadeh, Renée J. Miller |
ISWC (2) | 1 |
| 2015 | Toward a complete dataset of drug-drug interaction information from publicly available sourcesabstractAlthough potential drug-drug interactions (PDDIs) are a significant source of preventable drug-related harm, there is currently no single complete source of PDDI information. In the current study, all publically available sources of PDDI information that could be identified using a comprehensive and broad search were combined into a single dataset. The combined dataset merged fourteen different sources including 5 clinically-oriented information sources, 4 Natural Language Processing (NLP) Corpora, and 5 Bioinformatics/Pharmacovigilance information sources. As a comprehensive PDDI source, the merged dataset might benefit the pharmacovigilance text mining community by making it possible to compare the representativeness of NLP corpora for PDDI text extraction tasks, and specifying elements that can be useful for future PDDI extraction purposes. An analysis of the overlap between and across the data sources showed that there was little overlap. Even comprehensive PDDI lists such as DrugBank, KEGG, and the NDF-RT had less than 50% overlap with each other. Moreover, all of the comprehensive lists had incomplete coverage of two data sources that focus on PDDIs of interest in most clinical settings. Based on this information, we think that systems that provide access to the comprehensive lists, such as APIs into RxNorm, should be careful to inform users that the lists may be incomplete with respect to PDDIs that drug experts suggest clinicians be aware of. In spite of the low degree of overlap, several dozen cases were identified where PDDI information provided in drug product labeling might be augmented by the merged dataset. Moreover, the combined dataset was also shown to improve the performance of an existing PDDI NLP pipeline and a recently published PDDI pharmacovigilance protocol. Future work will focus on improvement of the methods for mapping between PDDI information sources, identifying methods to improve the use of the merged dataset in PDDI NLP algorithms, integrating high-quality PDDI information from the merged dataset into Wikidata, and making the combined dataset accessible as Semantic Web Linked Data. Serkan Ayvaz, John R. Horn, Oktie Hassanzadeh, Qian Zhu 0003, Johann Stan, Nicholas P. Tatonetti, Santiago Vilar, Mathias Brochhausen, Matthias Samwald, Majid Rastegar-Mojarad, Michel Dumontier, Richard D. Boyce |
J. Biomed. Informatics | 3 |
| 2015 | Schema Management for Document StoresabstractDocument stores that provide the efficiency of a schema-less interface are widely used by developers in mobile and cloud applications. However, the simplicity developers achieved controversially leads to complexity for data management due to lack of a schema. In this paper, we present a schema management framework for document stores. This framework discovers and persists schemas of JSON records in a repository, and also supports queries and schema summarization. The major technical challenge comes from varied structures of records caused by the schema-less data model and schema evolution. In the discovery phase, we apply a canonical form based method and propose an algorithm based on equivalent sub-trees to group equivalent schemas efficiently. Together with the algorithm, we propose a new data structure, eSiBu-Tree, to store schemas and support queries. In order to present a single summarized representation for heterogenous schemas in records, we introduce the concept of "skeleton", and propose to use it as a relaxed form of the schema, which captures a small set of core attributes. Finally, extensive experiments based on real data sets demonstrate the efficiency of our proposed schema discovery algorithms, and practical use cases in real-world data exploration and integration scenarios are presented to illustrate the effectiveness of using skeletons in these applications. Lanjun Wang, Oktie Hassanzadeh, Juwei Shi, Limei Jiao, Jia Zou 0001, Chen Wang 0018 |
Proc. VLDB Endow. | 2 |
| 2013 | Next Generation Data Analytics at IBM ResearchabstractNo abstract available. Oktie Hassanzadeh, Anastasios Kementsietsidis, Benny Kimelfeld, Rajasekar Krishnamurthy, Fatma Özcan 0001, Ippokratis Pandis |
Proc. VLDB Endow. | 1 |
| 2013 | Discovering Linkage Points over Web DataabstractA basic step in integration is the identification of linkage points, i.e., finding attributes that are shared (or related) between data sources, and that can be used to match records or entities across sources. This is usually performed using a match operator, that associates attributes of one database to another. However, the massive growth in the amount and variety of unstructured and semi-structured data on the Web has created new challenges for this task. Such data sources often do not have a fixed pre-defined schema and contain large numbers of diverse attributes. Furthermore, the end goal is not schema alignment as these schemas may be too heterogeneous (and dynamic) to meaningfully align. Rather, the goal is to align any overlapping data shared by these sources. We will show that even attributes with different meanings (that would not qualify as schema matches) can sometimes be useful in aligning data. The solution we propose in this paper replaces the basic schema-matching step with a more complex instance-based schema analysis and linkage discovery. We present a framework consisting of a library of efficient lexical analyzers and similarity functions, and a set of search algorithms for effective and efficient identification of linkage points over Web data. We experimentally evaluate the effectiveness of our proposed algorithms in real-world integration scenarios in several domains. Oktie Hassanzadeh, Ken Q. Pu, Soheil Hassas Yeganeh, Renée J. Miller, Lucian Popa 0001, Mauricio A. Hernández, C. T. Howard Ho |
Proc. VLDB Endow. | 1 |
| 2012 | Data Management Issues on the Semantic WebabstractWe provide an overview of the current data management research issues in the context of the Semantic Web. The objective is to introduce the audience into the area of the Semantic Web, and to highlight the fact that the area provides many interesting research opportunities for the data management community. A new model, the Resource Description Framework (RDF), coupled with a new query language, called SPARQL, lead us to revisit some classical data management problems, including efficient storage, query optimization, and data integration. These are problems that the Semantic Web community has only recently started to explore, and therefore the experience and long tradition of the database community can prove valuable. We target both experienced and novice researchers that are looking for a thorough presentation of the area and its key research topics. Oktie Hassanzadeh, Anastasios Kementsietsidis, Yannis Velegrakis |
ICDE | 1 |
| 2012 | Instance-Based Matching of Large Ontologies Using Locality-Sensitive Hashing
Songyun Duan, Achille Fokoue, Oktie Hassanzadeh, Anastasios Kementsietsidis, Kavitha Srinivas, Michael Jeffrey Ward |
ISWC (1) | 3 |
| 2011 | Linking Semistructured Data on the Web
Oktie Hassanzadeh, Soheil Hassas Yeganeh, Renée J. Miller |
WebDB | 1 |
| 2010 | Online annotation of text streams with structured entitiesabstractWe propose a framework and algorithm for annotating unbounded text streams with entities of a structured database. The algorithm allows one to correlate unstructured and dirty text streams from sources such as emails, chats and blogs, to entities stored in structured databases. In contrast to previous work on entity extraction, our emphasis is on performing entity annotation in a completely online fashion. The algorithm continuously extracts important phrases and assigns to them top-k relevant entities. Our algorithm does so with a guarantee of constant time and space complexity for each additional word in the text stream, thus infinite text streams can be annotated. Our framework allows the online annotation algorithm to adapt to changing stream rate by self-adjusting multiple run-time parameters to reduce or improve the quality of annotation for fast or slow streams, respectively. The framework also allows the online annotation algorithm to incorporate query feedback to learn the user preference and personalize the annotation for individual users. Ken Q. Pu, Oktie Hassanzadeh, Richard Drake, Renée J. Miller |
CIKM | 2 |
| 2009 | A framework for semantic link discovery over relational dataabstractDiscovering links between different data items in a single data source or across different data sources is a challenging problem faced by many information systems today. In particular, the recent Linking Open Data (LOD) community project has highlighted the paramount importance of establishing semantic links among web data sources. Currently, LOD sources provide billions of RDF triples, but only millions of links between data sources. Many of these data sources are published using tools that operate over relational data stored in a standard RDBMS. In this paper, we present a framework for discovery of semantic links from relational data. Our framework is based on declarative specification of linkage requirements by a user. We illustrate the use of our framework using several link discovery algorithms on a real world scenario. Our framework allows data publishers to easily find and publish high-quality links to other data sources, and therefore could significantly enhance the value of the data in the next generation of web. Oktie Hassanzadeh, Anastasios Kementsietsidis, Lipyeow Lim, Renée J. Miller, Min Wang 0001 |
CIKM | 1 |
| 2009 | A declarative framework for semantic link discovery over relational dataabstractIn this paper, we present a framework for online discovery of semantic links from relational data. Our framework is based on declarative specification of the linkage requirements by the user, that allows matching data items in many real-world scenarios. These requirements are translated to queries that can run over the relational data source, potentially using the semantic knowledge to enhance the accuracy of link discovery. Our framework lets data publishers to easily find and publish high-quality links to other data sources, and therefore could significantly enhance the value of the data in the next generation of web. Oktie Hassanzadeh, Lipyeow Lim, Anastasios Kementsietsidis, Min Wang 0001 |
WWW | 1 |
| 2009 | Framework for Evaluating Clustering Algorithms in Duplicate DetectionabstractThe presence of duplicate records is a major data quality concern in large databases. To detect duplicates, entity resolution also known as duplication detection or record linkage is used as a part of the data cleaning process to identify records that potentially refer to the same real-world entity. We present the Stringer system that provides an evaluation framework for understanding what barriers remain towards the goal of truly scalable and general purpose duplication detection algorithms. In this paper, we use Stringer to evaluate the quality of the clusters (groups of potential duplicates) obtained from several unconstrained clustering algorithms used in concert with approximate join techniques. Our work is motivated by the recent significant advancements that have made approximate join algorithms highly scalable. Our extensive evaluation reveals that some clustering algorithms that have never been considered for duplicate detection, perform extremely well in terms of both accuracy and scalability. Oktie Hassanzadeh, Fei Chiang, Renée J. Miller, Hyun Chul Lee |
Proc. VLDB Endow. | 1 |
| 2009 | Linkage Query WriterabstractWe present Linkage Query Writer (LinQuer), a system for generating SQL queries for semantic link discovery over relational data. The LinQuer framework consists of (a) LinQL, a language for specification of linkage requirements; (b) a web interface and an API for translating LinQL queries to standard SQL queries; (c) an interface that assists users in writing LinQL queries. We discuss the challenges involved in the design and implementation of a declarative and easy to use framework for discovering links between different data items in a single data source or across different data sources. We demonstrate different steps of the linkage requirements specification and discovery process in several real world scenarios and show how the LinQuer system can be used to create high-quality linked data sources. Oktie Hassanzadeh, Reynold Xin, Renée J. Miller, Anastasios Kementsietsidis, Lipyeow Lim, Min Wang 0001 |
Proc. VLDB Endow. | 1 |
| 2009 | Creating probabilistic databases from duplicated data
Oktie Hassanzadeh, Renée J. Miller |
VLDB J. | 1 |
| 2007 | Benchmarking declarative approximate selection predicatesabstractDeclarative data quality has been an active research topic. The fundamental principle behind a declarative approach to data quality is the use of declarative statements to realize data quality primitives on top of any relational data source. A primary advantage of such an approach is the ease of use and integration with existing applications. Over the last few years several similarity predicates have been proposed for common quality primitives (approximate selections, joins, etc) and have been fully expressed using declarative SQL statements. In this paper we propose new similarity predicates along with their declarative realization, based on notions of probabilistic information retrieval. In particular we show how language models and hidden Markov models can be utilized as similarity predicates for data quality and present their full declarative instantiation. We also show how other scoring methods from information retrieval, can be utilized in a similar setting. We then present full declarative specifications of previously proposed similarity predicates in the literature, grouping them into classes according to their primary characteristics. Finally, we present a thorough performance and accuracy study comparing a large number of similarity predicates for data cleaning operations. We quantify both their runtime performance as well as their accuracy for several types of common quality problems encountered in operational databases. Amit Chandel, Oktie Hassanzadeh, Nick Koudas, Mohammad Sadoghi, Divesh Srivastava |
SIGMOD Conference | 2 |
| 2005 | A Hybrid Approach for Refreshing Web Page Repositories
Mohammad Ghodsi, Oktie Hassanzadeh, Shahab Kamali, Morteza Monemizadeh |
DASFAA | 2 |