Fatma Özcan 0001

dblp:o/FatmaOzcan · also Fatma Ozcan 0001 · DBLP profile ↗
in reviewer pool ← Back
75ranked-venue papers in the field
18as first author
23since 2021 · last 2026
0000-0002-4418-4724ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 72 (17 first)Information Retrieval & Web Search · 2 (1 first)Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2026 High-Fidelity and Complex Test Data Generation for Google SQL Code Generation Services
Shivasankari Kannan, Yeounoh Chung, Amita Gondi, Tristan Swadell, Fatma Özcan 0001
ICDE5
2026 SemBench: A Benchmark for Semantic Query Processing Engines
Jiale Lao, Andreas Zimmerer, Olga Ovcharenko, Tianji Cong, Matthew Russo, Gerardo Vitagliano, Michael Cochez, Fatma Özcan 0001, Gautam Gupta, Thibaud Hottelier, H. V. Jagadish, Kris Kissel, Sebastian Schelter, Andreas Kipf, Immanuel Trummer
Proc. VLDB Endow.8
2025 Towards Foundation Database Models
Johannes Wehrstein, Carsten Binnig, Fatma Özcan 0001, Shobha Vasudevan
CIDR3
2025 Filtered Vector Search: State-of-the-art and Research Challenges
abstract
This tutorial provides a comprehensive overview of filtered vector search (fvs). Fvs queries combine vector search with relational operators. The tutorial explores the challenges of integrating vector search into database engines and emphasizes the need for new optimization techniques. It explains the three primary filtered search methods for fvs queries over generic tree-based and graph-based indices and examines the factors influencing the selection of the most efficient method. A key objective is to highlight the importance of achieving stable recall, ideally in a declarative manner, ensuring consistent recall across queries. The tutorial then discusses recent filter-optimized vector indices and concludes by identifying open research challenges in the field of fvs, aiming to inspire further research and development.
Helena Caminal, Yannis Chronis, Yannis Papakonstantinou, Fatma Özcan 0001, Anastasia Ailamaki
Proc. VLDB Endow.4
2025 Is Long Context All You Need? Leveraging LLM's Extended Context for NL2SQL
abstract
Large Language Models (LLMs) have demonstrated impressive capabilities across a range of natural language processing tasks. In particular, improvements in reasoning abilities and the expansion of context windows have opened new avenues for leveraging these powerful models. NL2SQL is challenging in that the natural language question is inherently ambiguous, while the SQL generation requires a precise understanding of complex data schema and semantics. One approach to this semantic ambiguous problem is to provide more and sufficient contextual information. In this work, we explore the performance and the latency tradeoffs of the extended context window (a.k.a., long context) offered by Google's state-of-the-art LLM ( gemini-1.5-pro ). We study the impact of various contextual information, including column example values, question and SQL query pairs, user-provided hints, SQL documentation, and schema. To the best of our knowledge, this is the first work to study how the extended context window and extra contextual information can help NL2SQL generation with respect to both accuracy and latency cost. We show that long context LLMs are robust and do not get lost in the extended contextual information. Additionally, our long-context NL2SQL pipeline based on Google's Gemini-pro-1.5 achieves strong performance across multiple benchmark datasets without fine-tuning or expensive self-consistency based techniques.
Yeounoh Chung, Gaurav Tarlok Kakkar, Brenton Milne, Fatma Özcan 0001
Proc. VLDB Endow.5
2025 Editorial for Special Issue: VLDB 2022
Juliana Freire, Fatma Özcan 0001, Xuemin Lin 0001
VLDB J.2
2023 HERMES: data placement and schema optimization for enterprise knowledge bases
Chuan Lei, Abdul Quamar, Vasilis Efthymiou, Fatma Özcan 0001, Rana Alotaibi
VLDB J.4
2022 Reflections On My Data Management Research Journey (VLDB Women in Database Research Award Talk)
abstract
Data-driven decision making is critical for all kinds of enterprises, public and private. It has been my mission to find more efficient, and effective ways to store, manage, query and analyze data to drive actionable insights. Throughout my career, I worked on many different technologies and systems, including semi-structured query processing, and large-scale data analytics. In this talk, I will talk about lessons learned, both technical and non-technical, using two of these systems as examples.
Fatma Özcan 0001
Proc. VLDB Endow.1
2022 Front Matter
Fatma Özcan 0001, Juliana Freire, Xuemin Lin 0001
Proc. VLDB Endow.1
2022 Front Matter
Fatma Özcan 0001, Juliana Freire, Xuemin Lin 0001
Proc. VLDB Endow.1
2022 Front Matter
Fatma Özcan 0001, Juliana Freire, Xuemin Lin 0001
Proc. VLDB Endow.1
2022 Front Matter
Fatma Özcan 0001, Juliana Freire, Xuemin Lin 0001
Proc. VLDB Endow.1
2022 Front Matter
Fatma Özcan 0001, Juliana Freire, Xuemin Lin 0001
Proc. VLDB Endow.1
2022 Front Matter
Fatma Özcan 0001, Juliana Freire, Xuemin Lin 0001
Proc. VLDB Endow.1
2022 Front Matter
Fatma Özcan 0001, Juliana Freire, Xuemin Lin 0001
Proc. VLDB Endow.1
2022 Front Matter
Fatma Özcan 0001, Juliana Freire, Xuemin Lin 0001
Proc. VLDB Endow.1
2022 Front Matter
Fatma Özcan 0001, Juliana Freire, Xuemin Lin 0001
Proc. VLDB Endow.1
2021 Property Graph Schema Optimization for Domain-Specific Knowledge Graphs
abstract
Enterprises are creating domain-specific knowledge graphs by curating and integrating their business data from multiple sources. Ontologies provide a semantic abstraction for such knowledge graphs to describe their data in terms of the entities involved and their relationships. There has been a lot of effort to build systems that enable efficient querying over knowledge graphs, represented as property graphs. However the problem of schema optimization in the property graph setting has been largely ignored. In this work, we show that graph schema design has significant impact on query performance, and propose two algorithms to generate an optimized property graph schema from the domain ontology. To the best of our knowledge, we are the first to present an ontology-driven approach for property graph schema optimization. The rich semantic relationships in an ontology contain a variety of opportunities to reduce edge traversals and consequently improve the graph query performance. Our experimental study with two real-world knowledge graphs shows that our algorithms produce high-quality schemas, achieving up to 2 orders of magnitude speed-up compared to alternative schema designs.
Rana Alotaibi, Chuan Lei, Abdul Quamar, Vasilis Efthymiou, Fatma Özcan 0001
ICDE5
2021 MEDTO: Medical Data to Ontology Matching Using Hybrid Graph Neural Networks
abstract
Medical ontologies are widely used to describe and organize medical terminologies and to support many critical applications on healthcare databases. These ontologies are often manually curated (e.g., UMLS, SNOMED CT, and MeSH) by medical experts. Medical databases, on the other hand, are often created by database administrators, using different terminology and structures. The discrepancies between medical ontologies and databases compromise interoperability between them. Data to ontology matching is the process of finding semantic correspondences between tables in databases to standard ontologies. Existing solutions such as ontology matching have mostly focused on engineering features from terminological, structural, and semantic model information extracted from the ontologies. However, this is often labor intensive and the accuracy varies greatly across different ontologies. Worse yet, the ontology capturing a medical database is often not given in practice. In this paper, we propose MEDTO, a novel end-to-end framework that consists of three innovative techniques: (1) a lightweight yet effective method that bootstrap a semantically rich ontology from a given medical database, (2) a hyperbolic graph convolution layer that encodes hierarchical concepts in the hyperbolic space, and (3) a heterogeneous graph layer that encodes both local and global context information of a concept. Experiments on two real-world medical datasets matching against SNOMED CT show significant improvements compared to the state-of-the-art methods. MEDTO also consistently achieves competitive results on a benchmark from the Ontology Alignment Evaluation Initiative.
Junheng Hao, Chuan Lei, Vasilis Efthymiou, Abdul Quamar, Fatma Özcan 0001, Yizhou Sun, Wei Wang 0010
KDD5
2021 Boomerang: Proactive Insight-Based Recommendations for Guiding Conversational Data Analysis
abstract
Natural-language interfaces are gaining popularity due to their potential to democratize access to data and insights by making the interaction with data more natural and accessible for a wide range of business users. To fully embrace the goal of democratization, it is also necessary to provide effective and continuous guidance support for data exploration. Conversational interfaces enable exploration of the data and insights search space in small incremental steps as the conversation with the data progresses. In this demo, we describe Boomerang, a system that recommends data-driven insights to guide exploration of datasets through a conversational interface. Boomerang aggregates recommendations from a variety of statistical, collaborative, and content-based recommenders, and selects insights that match closely to the user's current state of data exploration, represented as the \em conversational context. Boomerang combines various metrics, such as \em relevance, \em interestingness and \em timeliness, to rank the insights and recommends the insights based on current conversational context. In the demo, we will show how Boomerang enables guided data exploration on a sales dataset, containing information about products, retailers, sales, orders, inventory levels and regions.
Doris Jung Lin Lee, Abdul Quamar, Eser Kandogan, Fatma Özcan 0001
SIGMOD Conference4
2021 Medical Entity Disambiguation Using Graph Neural Networks
abstract
Medical knowledge bases (KBs), distilled from biomedical literature and regulatory actions, are expected to provide high-quality information to facilitate clinical decision making. Entity disambiguation (also referred to as entity linking) is considered as an essential task in unlocking the wealth of such medical KBs. However, existing medical entity disambiguation methods are not adequate due to word discrepancies between the entities in the KB and the text snippets in the source documents. Recently, graph neural networks (GNNs) have proven to be very effective and provide state-of-the-art results for many real-world applications with graph-structured data. In this paper, we introduce ED-GNN based on three representative GNNs (GraphSAGE, R-GCN, and MAGNN) for medical entity disambiguation. We develop two optimization techniques to fine-tune and improve ED-GNN. First, we introduce a novel strategy to represent entities that are mentioned in text snippets as a query graph. Second, we design an effective negative sampling strategy that identifies hard negative samples to improve the model's disambiguation capability. Compared to the best performing state-of-the-art solutions, our ED-GNN offers an average improvement of 7.3% in terms of F1 score on five real-world datasets.
Alina Vretinaris, Chuan Lei, Vasilis Efthymiou, Xiao Qin 0003, Fatma Özcan 0001
SIGMOD Conference5
2021 Front Matter
Fatma Özcan 0001, Juliana Freire, Xuemin Lin 0001
Proc. VLDB Endow.1
2021 Guest Editorial: Special issue on VLDB 2019
Fatma Özcan 0001, Lei Chen 0002
VLDB J.1
2020 Expanding Query Answers on Medical Knowledge Bases
Chuan Lei, Vasilis Efthymiou, Rebecca Geis, Fatma Özcan 0001
EDBT4
2020 State of the Art and Open Challenges in Natural Language Interfaces to Data
abstract
Recent advances in natural language understanding and processing resulted in renewed interest in natural language based interfaces to data, which provide an easy mechanism for non-technical users to access and query the data. While early systems only allowed simple selection queries over a single table, some recent work supports complex BI queries, with many joins and aggregation, and even nested queries. There are various approaches in the literature for interpreting user's natural language query. Rule-based systems try to identify the entities in the query, and understand the intended relationships between those entities. Recent years have seen the emergence and popularity of neural network based approaches which try to interpret the query holistically, by learning the patterns. In this tutorial, we will review these natural language interface solutions in terms of their interpretation approach, as well as the complexity of the queries they can generate. We will also discuss open research challenges.
Fatma Özcan 0001, Abdul Quamar, Jaydeep Sen, Chuan Lei, Vasilis Efthymiou
SIGMOD Conference1
2020 An Ontology-Based Conversation System for Knowledge Bases
abstract
Domain-specific knowledge bases (KB), carefully curated from various data sources, provide an invaluable reference for professionals. Conversation systems make these KBs easily accessible to professionals and are gaining popularity due to recent advances in natural language understanding and AI. Despite the increasing use of various conversation systems in open-domain applications, the requirements of a domain-specific conversation system are quite different and challenging. In this paper, we propose an ontology-based conversation system for domain-specific KBs. In particular, we exploit the domain knowledge inherent in the domain ontology to identify user intents, and the corresponding entities to bootstrap the conversation space. We incorporate the feedback from domain experts to further refine these patterns, and use them to generate training samples for the conversation model, lifting the heavy burden from the conversation designers. We have incorporated our innovations into a conversation agent focused on healthcare as a feature of the IBM Micromedex product.
Abdul Quamar, Chuan Lei, Dorian Miller, Fatma Özcan 0001, Jeffrey T. Kreulen, Robert J. Moore, Vasilis Efthymiou
SIGMOD Conference4
2020 Db2 Event Store: A Purpose-Built IoT Database Engine
abstract
The requirements of Internet of Things (IoT) workloads are unique in the database space. While significant effort has been spent over the last decade rearchitecting OLTP and Analytics workloads for the public cloud, little has been done to rearchitect IoT workloads for the cloud. In this paper we present IBM Db2 Event Store ™ , a cloud-native database system designed specifically for IoT workloads, which require extremely high-speed ingest, efficient and open data storage, and near real-time analytics. Additionally, by leveraging the Db2 SQL compiler, optimizer and runtime, developed and refined over the last 30 years, we demonstrate that rearchitecting for the public cloud doesn't require rewriting all components. Reusing components that have been built out and optimized for decades dramatically reduced the development effort and immediately provided rich SQL support and excellent run-time query performance.
Christian Garcia-Arellano, Adam J. Storm, David Kalmuk, Hamdi Roumani, Ron Barber, Yuanyuan Tian 0001, Richard Sidle, Fatma Özcan 0001, Matt Spilchen, Josh Tiefenbach, Daniel C. Zilio, Lan Pham, Kostas Rakopoulos, Alexander Cheung, Darren Pepper, Imran Sayyid, Gidon Gershinsky, Gal Lushi, Hamid Pirahesh
Proc. VLDB Endow.8
2020 Conversational BI: An Ontology-Driven ConversationSystem for Business Intelligence Applications
abstract
Business intelligence (BI) applications play an important role in the enterprise to make critical business decisions. Conversational interfaces enable non-technical enterprise users to explore their data, democratizing access to data significantly. In this paper, we describe an ontology-based framework for creating a conversation system for BI applications termed as Conversational BI. We create an ontology from a business model underlying the BI application, and use this ontology to automatically generate various artifacts of the conversation system. These include the intents, entities, as well as the training samples for each intent. Our approach builds upon our earlier work, and exploits common BI access patterns to generate intents, their training examples and adapt the dialog structure to support typical BI operations. We have implemented our techniques in Health Insights (HI) , an IBM Watson Healthcare offering, providing analysis over insurance data on claims. Our user study demonstrates that our system is quite intuitive for gaining business insights from data. We also show that our approach not only captures the analysis available in the fixed application dashboards, but also enables new queries and explorations.
Abdul Quamar, Fatma Özcan 0001, Dorian Miller, Robert J. Moore, Rebecca Niehus, Jeffrey T. Kreulen
Proc. VLDB Endow.2
2020 ATHENA++: Natural Language Querying for Complex Nested SQL Queries
Jaydeep Sen, Chuan Lei, Abdul Quamar, Fatma Özcan 0001, Vasilis Efthymiou, Ayushi Dalmia, Greg Stager, Ashish R. Mittal, Diptikalyan Saha, Karthik Sankaranarayanan
Proc. VLDB Endow.4
2019 Natural Language Querying of Complex Business Intelligence Queries
abstract
Natural Language Interface to Database (NLIDB) eliminates the need for an end user to use complex query languages like SQL by translating the input natural language statements to SQL automatically. Although NLIDB systems have seen rapid growth of interest recently, the current state-of-the-art systems can at best handle point queries to retrieve certain column values satisfying some filters, or aggregation queries involving basic SQL aggregation functions. In this demo, we showcase our NLIDB system with extended capabilities for business applications that require complex nested SQL queries without prior training or feedback from human in-the-loop. In particular, our system uses novel algorithms that combine linguistic analysis with deep domain reasoning for solving core challenges in handling nested queries. To demonstrate the capabilities, we propose a new benchmark dataset containing realistic business intelligence queries, conforming to an ontology derived from FIBO and FRO financial ontologies. In this demo, we will showcase a wide range of complex business intelligence queries against our benchmark dataset, with increasing level of complexity. The users will be able to examine the SQL queries generated, and also will be provided with an English description of the interpretation.
Jaydeep Sen, Fatma Özcan 0001, Abdul Quamar, Greg Stager, Ashish R. Mittal, Manasa Jammi, Chuan Lei, Diptikalyan Saha, Karthik Sankaranarayanan
SIGMOD Conference2
2019 Front Matter
Lei Chen 0002, Fatma Özcan 0001
Proc. VLDB Endow.2
2019 Front Matter
Lei Chen 0002, Fatma Özcan 0001
Proc. VLDB Endow.2
2019 Front Matter
Lei Chen 0002, Fatma Özcan 0001
Proc. VLDB Endow.2
2019 Front Matter
Lei Chen 0002, Fatma Özcan 0001
Proc. VLDB Endow.2
2019 Front Matter
Lei Chen 0002, Fatma Özcan 0001
Proc. VLDB Endow.2
2019 Front Matter
Lei Chen 0002, Fatma Özcan 0001
Proc. VLDB Endow.2
2019 Front Matter
Lei Chen 0002, Fatma Özcan 0001
Proc. VLDB Endow.2
2019 Front Matter
Lei Chen 0002, Fatma Özcan 0001
Proc. VLDB Endow.2
2019 Front Matter
Lei Chen 0002, Fatma Özcan 0001
Proc. VLDB Endow.2
2018 Front Matter
Lei Chen 0002, Fatma Özcan 0001
Proc. VLDB Endow.2
2018 Front Matter
Lei Chen 0002, Fatma Özcan 0001
Proc. VLDB Endow.2
2018 Front Matter
Lei Chen 0002, Fatma Özcan 0001
Proc. VLDB Endow.2
2017 Evolving Databases for New-Gen Big Data Applications
Ron Barber, Christian Garcia-Arellano, Ronen Grosman, René Müller 0001, Vijayshankar Raman, Richard Sidle, Matt Spilchen, Adam J. Storm, Yuanyuan Tian 0001, Pinar Tözün, Daniel C. Zilio, Matt Huras, Guy M. Lohman, C. Mohan 0001, Fatma Özcan 0001, Hamid Pirahesh
CIDR15
2017 Hybrid Transactional/Analytical Processing: A Survey
abstract
The popularity of large-scale real-time analytics applications (real-time inventory/pricing, recommendations from mobile apps, fraud detection, risk analysis, IoT, etc.) keeps rising. These applications require distributed data management systems that can handle fast concurrent transactions (OLTP) and analytics on the recent data. Some of them even need running analytical queries (OLAP) as part of transactions. Efficient processing of individual transactional and analytical requests, however, leads to different optimizations and architectural decisions while building a data management system.
Fatma Özcan 0001, Yuanyuan Tian 0001, Pinar Tözün
SIGMOD Conference1
2017 Creation and Interaction with Large-scale Domain-Specific Knowledge Bases
abstract
The ability to create and interact with large-scale domain-specific knowledge bases from unstructured/semi-structured data is the foundation for many industry-focused cognitive systems. We will demonstrate the Content Services system that provides cloud services for creating and querying high-quality domain-specific knowledge bases by analyzing and integrating multiple (un/semi)structured content sources. We will showcase an instantiation of the system for a financial domain. We will also demonstrate both cross-lingual natural language queries and programmatic API calls for interacting with this knowledge base.
Shreyas Bharadwaj, Laura Chiticariu, Marina Danilevsky, Samarth Dhingra, Samved Divekar, Arnaldo Carreno-Fuentes, Nitin Gupta 0005, Sang-Don Han, Mauricio A. Hernández, C. T. Howard Ho, Parag Jain, Salil Joshi 0001, Hima P. Karanam, Saravanan Krishnan, Rajasekar Krishnamurthy, Yunyao Li 0001, Satishkumaar Manivannan, Ashish R. Mittal, Fatma Özcan 0001, Abdul Quamar, Poornima Chozhiyath Raman, Diptikalyan Saha, Karthik Sankaranarayanan, Jaydeep Sen, Prithviraj Sen, Shivakumar Vaithyanathan, Mitesh Vasa, Huaiyu Zhu 0001
Proc. VLDB Endow.20
2016 Wildfire: Concurrent Blazing Data Ingest and Analytics
abstract
We demonstrate Hybrid Transactional and Analytics Processing (HTAP) on the Spark platform by the Wildfire prototype, which can ingest up to ~6 million inserts per second per node and simultaneously perform complex SQL analytics queries. Here, a simplified mobile application uses Wildfire to recommend advertising to mobile customers based upon their distance from stores and their interest in products sold by these stores, while continuously graphing analytics results as those customers move and respond to the ads with purchases.
Ron Barber, Matt Huras, Guy M. Lohman, C. Mohan 0001, René Müller 0001, Fatma Özcan 0001, Hamid Pirahesh, Vijayshankar Raman, Richard Sidle, Oleg Sidorkin, Adam J. Storm, Yuanyuan Tian 0001, Pinar Tözün
SIGMOD Conference6
2016 ATHENA: An Ontology-Driven System for Natural Language Querying over Relational Data Stores
abstract
In this paper, we present ATHENA, an ontology-driven system for natural language querying of complex relational databases. Natural language interfaces to databases enable users easy access to data, without the need to learn a complex query language, such as SQL. ATHENA uses domain specific ontologies, which describe the semantic entities, and their relationships in a domain. We propose a unique two-stage approach, where the input natural language query (NLQ) is first translated into an intermediate query language over the ontology, called OQL, and subsequently translated into SQL. Our two-stage approach allows us to decouple the physical layout of the data in the relational store from the semantics of the query, providing physical independence. Moreover, ontologies provide richer semantic information, such as inheritance and membership relations, that are lost in a relational schema. By reasoning over the ontologies, our NLQ engine is able to accurately capture the user intent. We study the effectiveness of our approach using three different workloads on top of geographical (GEO), academic (MAS) and financial (FIN) data. ATHENA achieves 100% precision on the GEO and MAS workloads, and 99% precision on the FIN workload which operates on a complex financial ontology. Moreover, ATHENA attains 87.2%, 88.3%, and 88.9% recall on the GEO, MAS, and FIN workloads, respectively.
Diptikalyan Saha, Avrilia Floratou, Karthik Sankaranarayanan, Umar Farooq Minhas, Ashish R. Mittal, Fatma Özcan 0001
Proc. VLDB Endow.6
2016 Building a Hybrid Warehouse: Efficient Joins between Data Stored in HDFS and Enterprise Warehouse
abstract
The Hadoop Distributed File System (HDFS) has become an important data repository in the enterprise as the center for all business analytics, from SQL queries and machine learning to reporting. At the same time, enterprise data warehouses (EDWs) continue to support critical business analytics. This has created the need for a new generation of a special federation between Hadoop-like big data platforms and EDWs, which we call the hybrid warehouse . There are many applications that require correlating data stored in HDFS with EDW data, such as the analysis that associates click logs stored in HDFS with the sales data stored in the database. All existing solutions reach out to HDFS and read the data into the EDW to perform the joins, assuming that the Hadoop side does not have efficient SQL support. In this article, we show that it is actually better to do most data processing on the HDFS side, provided that we can leverage a sophisticated execution engine for joins on the Hadoop side. We identify the best hybrid warehouse architecture by studying various algorithms to join database and HDFS tables. We utilize Bloom filters to minimize the data movement and exploit the massive parallelism in both systems to the fullest extent possible. We describe a new zigzag join algorithm and show that it is a robust join algorithm for hybrid warehouses that performs well in almost all cases. We further develop a sophisticated cost model for the various join algorithms and show that it can facilitate query optimization in the hybrid warehouse to correctly choose the right algorithm under different predicate and join selectivities.
Yuanyuan Tian 0001, Fatma Özcan 0001, Romulo Goncalves, Hamid Pirahesh
ACM Trans. Database Syst.2
2015 A Generic Solution to Integrate SQL and Analytics for Big Data
abstract
There is a need to integrate SQL processing with more advanced machine learning (ML) analytics to drive actionable insights from large volumes of data. As a first step towards this integration, we study how to efficiently connect big SQL systems (either MPP databases or new-generation SQL-on-Hadoop systems) with distributed big ML systems. We identify two important challenges to address in the integrated data analytics pipeline: data transformation, how to efficiently transform SQL data into a form suitable for ML, and data transfer, how to efficiently handover SQL data to ML systems. For the data transformation problem, we propose an In-SQL approach to incorporate common data transformations for ML inside SQL systems through extended user-defined functions (UDFs), by exploiting the massive parallelism of the big SQL systems. We propose and study a general method for transferring data between big SQL and big ML systems in a parallel streaming fashion. Furthermore, we explore caching intermediate or final results of data transformation to improve the performance. Our techniques are generic: they apply to any big SQL system that supports UDFs and any big ML system that uses Hadoop InputFormats to ingest input data.
Nikos R. Katsipoulakis, Yuanyuan Tian 0001, Fatma Özcan 0001, Hamid Pirahesh, Berthold Reinwald
EDBT3
2015 Joins for Hybrid Warehouses: Exploiting Massive Parallelism in Hadoop and Enterprise Data Warehouses
abstract
HDFS has become an important data repository in the enterprise as the center for all business analytics, from SQL queries, machine learning to reporting. At the same time, enterprise data warehouses (EDWs) continue to support critical business analytics. This has created the need for a new generation of special federation between Hadoop-like big data platforms and EDWs, which we call the hybrid warehouse. There are many applications that require correlating data stored in HDFS with EDW data, such as the analysis that associates click logs stored in HDFS with the sales data stored in the database. All existing solutions reach out to HDFS and read the data into the EDW to perform the joins, assuming that the Hadoop side does not have the efficient SQL support. In this paper, we show that it is actually better to do most data processing on the HDFS side, provided that we can leverage a sophisticated execution engine for joins on the Hadoop side. We identify the best hybrid warehouse architecture by studying various algorithms to join database and HDFS tables. We utilize Bloom filters to minimize the data movement, and exploit the massive parallelism in both systems to the fullest extent possible. We describe a new zigzag join algorithm, and show that it is a robust join algorithm for hybrid warehouses which performs well in almost all cases.
Yuanyuan Tian 0001, Fatma Özcan 0001, Romulo Goncalves, Hamid Pirahesh
EDBT3
2015 Tutorial: SQL-on-Hadoop Systems
abstract
Enterprises are increasingly using Apache Hadoop, more specifically HDFS, as a central repository for all their data; data coming from various sources, including operational systems, social media and the web, sensors and smart devices, as well as their applications. At the same time many enterprise data management tools (e.g. from SAP ERP and SAS to Tableau) rely on SQL and many enterprise users are familiar and comfortable with SQL. As a result, SQL processing over Hadoop data has gained significant traction over the recent years, and the number of systems that provide such capability has increased significantly. In this tutorial we use the term SQL-on-Hadoop to refer to systems that provide some level of declarative SQL(-like) processing over HDFS and noSQL data sources, using architectures that include computational or storage engines compatible with Apache Hadoop.
Daniel J. Abadi, Shivnath Babu, Fatma Özcan 0001, Ippokratis Pandis
Proc. VLDB Endow.3
2015 Front Matter
Fatma Özcan 0001
Proc. VLDB Endow.1
2015 Clash of the Titans: MapReduce vs. Spark for Large Scale Data Analytics
abstract
MapReduce and Spark are two very popular open source cluster computing frameworks for large scale data analytics. These frameworks hide the complexity of task parallelism and fault-tolerance, by exposing a simple programming API to users. In this paper, we evaluate the major architectural components in MapReduce and Spark frameworks including: shuffle, execution model, and caching, by using a set of important analytic workloads. To conduct a detailed analysis, we developed two profiling tools: (1) We correlate the task execution plan with the resource utilization for both MapReduce and Spark, and visually present this correlation; (2) We provide a break-down of the task execution time for in-depth analysis. Through detailed experiments, we quantify the performance differences between MapReduce and Spark. Furthermore, we attribute these performance differences to different components which are architected differently in the two frameworks. We further expose the source of these performance differences by using a set of micro-benchmark experiments. Overall, our experiments show that Spark is about 2.5x, 5x, and 5x faster than MapReduce, for Word Count, k-means, and PageRank, respectively. The main causes of these speedups are the efficiency of the hash-based aggregation component for combine, as well as reduced CPU and disk overheads due to RDD caching in Spark. An exception to this is the Sort workload, for which MapReduce is 2x faster than Spark. We show that MapReduce's execution model is more efficient for shuffling data than Spark, thus making Sort run faster on MapReduce.
Juwei Shi, Yunjie Qiu, Umar Farooq Minhas, Limei Jiao, Berthold Reinwald, Fatma Özcan 0001
Proc. VLDB Endow.7
2014 Dynamically optimizing queries over large scale data platforms
abstract
Enterprises are adapting large-scale data processing platforms, such as Hadoop, to gain actionable insights from their "big data". Query optimization is still an open challenge in this environment due to the volume and heterogeneity of data, comprising both structured and un/semi-structured datasets. Moreover, it has become common practice to push business logic close to the data via user-defined functions (UDFs), which are usually opaque to the optimizer, further complicating cost-based optimization. As a result, classical relational query optimization techniques do not fit well in this setting, while at the same time, suboptimal query plans can be disastrous with large datasets. In this paper, we propose new techniques that take into account UDFs and correlations between relations for optimizing queries running on large scale clusters. We introduce "pilot runs", which execute part of the query over a sample of the data to estimate selectivities, and employ a cost-based optimizer that uses these selectivities to choose an initial query plan. Then, we follow a dynamic optimization approach, in which plans evolve as parts of the queries get executed. Our experimental results show that our techniques produce plans that are at least as good as, and up to 2x (4x) better for Jaql (Hive) than, the best hand-written left-deep query plans.
Konstantinos Karanasos, Andrey Balmin, Marcel Kutsch, Fatma Özcan 0001, Vuk Ercegovac, Chunyang Xia, Jesse Jackson
SIGMOD Conference4
2014 Are we experiencing a big data bubble?
abstract
No abstract available.
Fatma Özcan 0001, Nesime Tatbul, Daniel J. Abadi, Marcel Kornacker, C. Mohan 0001, Karthikeyan Ramasamy, Janet L. Wiener
SIGMOD Conference1
2014 SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures
abstract
SQL query processing for analytics over Hadoop data has recently gained significant traction. Among many systems providing some SQL support over Hadoop, Hive is the first native Hadoop system that uses an underlying framework such as MapReduce or Tez to process SQL-like statements. Impala, on the other hand, represents the new emerging class of SQL-on-Hadoop systems that exploit a shared-nothing parallel database architecture over Hadoop. Both systems optimize their data ingestion via columnar storage, and promote different file formats: ORC and Parquet. In this paper, we compare the performance of these two systems by conducting a set of cluster experiments using a TPC-H like benchmark and two TPC-DS inspired workloads. We also closely study the I/O efficiency of their columnar formats using a set of micro-benchmarks. Our results show that Impala is 3.3 X to 4.4 X faster than Hive on MapReduce and 2.1 X to 2.8 X than Hive on Tez for the overall TPC-H experiments. Impala is also 8.2 X to 10 X faster than Hive on MapReduce and about 4.3 X faster than Hive on Tez for the TPC-DS inspired experiments. Through detailed analysis of experimental results, we identify the reasons for this performance gap and examine the strengths and limitations of each system.
Avrilia Floratou, Umar Farooq Minhas, Fatma Özcan 0001
Proc. VLDB Endow.3
2013 Eagle-eyed elephant: split-oriented indexing in Hadoop
abstract
An increasingly important analytics scenario for Hadoop involves multiple (often ad hoc) grouping and aggregation queries with selection predicates over a slowly changing dataset. These queries are typically expressed via high-level query languages such as Jaql, Pig, and Hive, and are used either directly for business-intelligence applications or to prepare the data for statistical model building and machine learning. In such scenarios it has been increasingly recognized that, as in classical databases, techniques for avoiding access to irrelevant data can dramatically improve query performance. Prior work on Hadoop, however, has simply ported classical techniques to the MapReduce setting, focusing on record-level indexing and key-based partition elimination. Unfortunately, record-level indexing only slightly improves overall query performance, because it does not minimize the number of mapper "waves", which is determined by the number of processed splits. Moreover, key-based partitioning requires data reorganization, which is usually impractical in Hadoop settings. We therefore need to re-envision how data access mechanisms are defined and implemented. To this end, we introduce the Eagle-Eyed Elephant (E3) framework for boosting the efficiency of query processing in Hadoop by avoiding accesses of data splits that are irrelevant to the query at hand. Using novel techniques involving inverted indexes over splits, domain segmentation, materialized views, and adaptive caching, E3 avoids accessing irrelevant splits even in the face of evolving workloads and data. Our experiments show that E3 can achieve up to 20x cost savings with small to moderate storage overheads.
Mohamed Y. Eltabakh, Fatma Özcan 0001, Yannis Sismanis, Peter J. Haas, Hamid Pirahesh, Jan Vondrák
EDBT2
2013 Next Generation Data Analytics at IBM Research
abstract
No abstract available.
Oktie Hassanzadeh, Anastasios Kementsietsidis, Benny Kimelfeld, Rajasekar Krishnamurthy, Fatma Özcan 0001, Ippokratis Pandis
Proc. VLDB Endow.5
2012 CDMW 2012 - city data management workshop: workshop summary
abstract
Cities today have become highly dense, dynamic living areas for the majority of planet's population and also focal points of innovation, commerce, and growth in a highly modernized world. Due to its intensifying importance, cities need to transform into sustainable, smarter and credible places to enable a tenantable and comfortable life for their citizens. City data, which is the source of our digitized knowledge about the cities, is a highly important element to achive this goal as it is the main input to build complex city ecosystems and to solve particular problems that are encountered in the cities today. In this respect, this workshop will provide a major forum to identify the challenges and opportunities in terms of better managing city data and to reveal its discriminating importance in various applications in a city ecosystem. As city data becomes more widespread and prevailing, it poses novel research problems which importantly are open to the investigation of a broad community of researchers in various fields.
Veli Bicer, Thanh Tran 0001, Fatma Özcan 0001, Opher Etzion
CIKM3
2011 Emerging trends in the enterprise data analytics: connecting Hadoop and DB2 warehouse
abstract
Enterprises are dealing with ever increasing volumes of data, reaching into the petabyte scale. With many of our customer engagements, we are observing an emerging trend: They are using Hadoop-based solutions in conjunction with their data warehouses. They are using Hadoop to deal with the data volume, as well as the lack of strict structure in their data to conduct various analyses, including but not limited to Web log analysis, sophisticated data mining, machine learning and model building. This first stage of the analysis is off-line and suitable for Hadoop. But, once their data is summarized or cleansed enough, and their models are built, they are loading the results into a warehouse for interactive querying and report generation. At this later stage, they leverage the wealth of business intelligence tools, which they are accustomed to, that exist for warehouses. In this paper, we outline this use case and discuss the bidirectional connectors we developed between IBM DB2 and IBM InfoSphere BigInsights.
Fatma Özcan 0001, David Hoa, Kevin S. Beyer, Andrey Balmin, Chuan Jie Liu
SIGMOD Conference1
2011 Jaql: A Scripting Language for Large Scale Semistructured Data Analysis
Kevin S. Beyer, Vuk Ercegovac, Rainer Gemulla, Andrey Balmin, Mohamed Y. Eltabakh, Carl-Christian Kanne, Fatma Özcan 0001, Eugene J. Shekita
Proc. VLDB Endow.7
2011 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop
abstract
Hadoop has become an attractive platform for large-scale data analytics. In this paper, we identify a major performance bottleneck of Hadoop: its lack of ability to colocate related data on the same set of nodes. To overcome this bottleneck, we introduce CoHadoop, a lightweight extension of Hadoop that allows applications to control where data are stored. In contrast to previous approaches, CoHadoop retains the flexibility of Hadoop in that it does not require users to convert their data to a certain format (e.g., a relational database or a specific file format). Instead, applications give hints to CoHadoop that some set of files are related and may be processed jointly; CoHadoop then tries to colocate these files for improved efficiency. Our approach is designed such that the strong fault tolerance properties of Hadoop are retained. Colocation can be used to improve the efficiency of many operations, including indexing, grouping, aggregation, columnar storage, joins, and sessionization. We conducted a detailed study of joins and sessionization in the context of log processing---a common use case for Hadoop---, and propose efficient map-only algorithms that exploit colocated data partitions. In our experiments, we observed that CoHadoop outperforms both plain Hadoop and previous work. In particular, our approach not only performs better than repartition-based algorithms, but also outperforms map-only algorithms that do exploit data partitioning but not colocation. 8.
Mohamed Y. Eltabakh, Yuanyuan Tian 0001, Fatma Özcan 0001, Rainer Gemulla, Aljoscha Krettek, John McPherson
Proc. VLDB Endow.3
2009 Search Driven Analysis of Heterogenous XML Data
Andrey Balmin, Latha S. Colby, Emiran Curtmola, Quanzhong Li 0002, Fatma Özcan 0001
CIDR5
2008 Grouping and Optimization of XPath Expressions in System RX
abstract
Several XML DBMS support XQuery and/or SQL/XML languages, which are based on navigational primitives in the form of XPath expressions. Typically, these systems either model each XPath step as a separate query plan operator, or employ holistic approaches that can evaluate multiple steps of a single XPath expression. There have also been proposals to execute as many XPath expressions as possible within a single FLWOR block simultaneously in a data streaming context. We observe in our System-RX prototype that blindly combining all possible XPath expressions for concurrent execution can result in significant performance degradation. We identify two main problems. First, the simple strategy of grouping all XPath expressions on a single document does not always work if the query involves more than one data source or has nested query blocks. Second, merging XPath expressions may result in unnecessary execution of branches that can be filtered by predicates in other branches or elsewhere in the query. To rectify these problems, we develop a combination of heuristic-based rewrite transformations, to decide which XPath expressions should be grouped for concurrent evaluation, and cost-based optimization to globally order the groups within the query execution plan, and locally order the branches within individual groups. Experimental evaluation confirms that selectively grouping multiple XPath expressions allows for better query evaluation performance and reduces the query optimization complexity.
Andrey Balmin, Fatma Özcan 0001, Edison Ting
ICDE2
2008 Grouping and optimization of XPath expressions in DB2 pureXML
abstract
Several XML DBMSs support XQuery and/or SQL/XML languages, which are based on navigational primitives in the form of XPath expressions. Typically, these systems either model each XPath step as a separate query plan operator, or employ holistic approaches that can evaluate multiple steps of a single XPath expression. There have also been proposals to execute as many XPath expressions as possible within a single FLWOR block simultaneously in a data streaming context.
Andrey Balmin, Fatma Özcan 0001, Edison Ting
SIGMOD Conference2
2008 SEDA: a system for search, exploration, discovery, and analysis of XML Data
abstract
Keyword search in XML repositories is a powerful tool for interactive data exploration. Much work has recently been done on making XML search aware of relationship information embedded in XML document structure, but without a clear winner in all data and query scenarios. Furthermore, due to its imprecise nature, search results cannot easily be analyzed and summarized to gain more insights into the data. We address these shortcomings with SEDA: a system for Search, Exploration, Discovery, and Analysis of XML Data. SEDA is based on a paradigm of search and user interaction to help users start with simple keyword-style querying and perform rich analysis of XML data by leveraging both the content and structure of the data. SEDA is an interactive system that allows the user to refine her query iteratively to explore the XML data and discover interesting relationships. SEDA first employs a top-k algorithm to compute the most relevant top-k answers fast, and returns tuples of nodes ranked by relevance. SEDA provides several novel data structures and techniques for efficient top-k computation over graph-structured XML data. SEDA also computes all the contexts in which the query terms are found and all the connection paths that connect the query terms in the XML data. These two summaries enable the user to refine her query by disambiguating the contexts and connections relevant to her query. With the user feedback, the system has enough information to compute all query results, not just the top-k. From the complete results, SEDA automatically deduces a star schema, which is then instantiated with the query results and augmented with additional values required for a well-defined data cube. The tables computed at this step are input into an OLAP engine for further analysis.
Andrey Balmin, Latha S. Colby, Emiran Curtmola, Quanzhong Li 0002, Fatma Özcan 0001, Sharath Srinivas, Zografoula Vagena
Proc. VLDB Endow.5
2006 On the Path to Efficient XML Queries
Andrey Balmin, Kevin S. Beyer, Fatma Özcan 0001, Matthias Nicola
VLDB3
2005 Extending XQuery for Analytics
abstract
XQuery is a query language under development by the W3C XML Query Working Group. The language contains constructs for navigating, searching, and restructuring XML data. With XML gaining importance as the standard for representing business data, XQuery must support the types of queries that are common in business analytics. One such class of queries is OLAP-style aggregation queries. Although these queries are expressible in XQuery Version 1, the lack of explicit grouping constructs makes the construction of these queries non-intuitive and places a burden on the XQuery engine to recognize and optimize the implicit grouping constructs. Furthermore, although the flexibility of the XML data model provides an opportunity for advanced forms of grouping that are not easily represented in relational systems, these queries are difficult to express using the current XQuery syntax. In this paper, we provide a proposal for extending the XQuery FLWOR expression with explicit syntax for grouping and for numbering of results. We show that these new XQuery constructs not only simplify the construction and evaluation of queries requiring grouping and ranking but also enable complex analytic queries such as moving-window aggregation and rollups along dynamic hierarchies to be expressed without additional language extensions.
Kevin S. Beyer, Donald D. Chamberlin, Latha S. Colby, Fatma Özcan 0001, Hamid Pirahesh
SIGMOD Conference4
2005 DB2/XML: designing for evolution
abstract
DB2 provides native XML storage, indexing, navigation and query processing through both SQL/XML and XQuery using the XML data type introduced by SQL/XML. In this tutorial we focus on DB2's XML support for schema evolution, especially DB2's schema repository and document-level validation.
Kevin S. Beyer, Fatma Özcan 0001, Sundar Saiprasad, Bert Van der Linden
SIGMOD Conference2
2005 System RX: One Part Relational, One Part XML
abstract
This paper describes the overall architecture and design aspects of a hybrid relational and XML database system called System RX. We believe that such a system is fundamental in the evolution of enterprise data management solutions: XML and relational data will co-exist and complement each other in enterprise solutions. Furthermore, a successful XML repository requires much of the same infrastructure that already exists in a relational database management system. Finally, XML query languages have considerable conceptual and functional overlap with relational dataflow engines. System RX is the first truly hybrid system that comingles XML and relational data, giving them equal footing. The new support for XML includes native support for storage and indexing as well as query compilation and evaluation support for the latest industry-standard query languages, SQL/XML and XQuery. By building a hybrid system, we leverage more than 20 years of data management research to advance XML technology to the same standards expected from mature relational systems.
Kevin S. Beyer, Roberta Cochrane, Vanja Josifovski, Jim Kleewein, George Lapis, Guy M. Lohman, Robert Lyle, Fatma Özcan 0001, Hamid Pirahesh, Normen Seemann, Tuong C. Truong, Bert Van der Linden, Brian Vickery
SIGMOD Conference8
2004 A Framework for Using Materialized XPath Views in XML Query Processing
Andrey Balmin, Fatma Özcan 0001, Kevin S. Beyer, Roberta Cochrane, Hamid Pirahesh
VLDB2
1999 Cost Models DO Matter: Providing Cost Information for Diverse Data Sources in a Federated System
Mary Roth, Fatma Özcan 0001, Laura M. Haas
VLDB2
1997 Multidatabase Query Optimization
Cem Evrendilek, Asuman Dogac, Sena Nural Arpinar, Fatma Özcan 0001
Distributed Parallel Databases4
1996 Dynamic Query Optimization on a Distributed Object Management Platform
abstract
A DistributedObject Management (DOM) architecture, when used as the infrastructure of a multidatabase system, not only enables easy and flexible interoperation of D13MSS, but also facilitates interoperation of the multidatabase system with other repositories that do not have DIBMS capabilities.Thk is an important advantage, since most of data still resides on repositories that do not have DIBMS capabilities.In thk paper, we describe a dynamic query optimization technique for a multidatabaae system, namely MIND, implement ed on a DOM environment.Dynamic query optimization, which schedules intersite operations at runtime, fits better to such an environment since it benefits from location transparency provided by the DOM framework.In thk way, the dynamic changes in the configuration of system resources such as a relocated DBMS or a new mirror to an existing DBMS, do not affect the optimized query execution in the system.Furthermore, the uncertainty in estimating the appearance times (i.e., the execution time of the global @ 199(j ACM @89791 +73+/9fj/l 1 ..$3 .5~117 experiments indicate that the dynamic query optimization technique presented in thk paper has a better performance.
Fatma Özcan 0001, Sena Nural Arpinar, Pinar Koksal, Cem Evrendilek, Asuman Dogac
CIKM1
1996 METU Interoperable Database System
Asuman Dogac, Ugur Halici, Ebru Kilic, Gökhan Özhan, Fatma Özcan 0001, Sena Nural Arpinar, Cevdet Dengi, Sema Mancuhan, Ismailcem Budak Arpinar, Pinar Koksal, Cem Evrendilek
SIGMOD Conference5