Vinayak R. Borkar

dblp:43/3643 · DBLP profile ↗
← Back
29ranked-venue papers
10as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 26 · 9 first-authorArtificial intelligence and machine learning · 3Software engineering, systems software and programming languages · 2Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
18 papers
Data integration and cleaning · 21% Data models and query languages · 16% Database system architecture and tuning · 14%
Software engineering, system software, and programming languages
4 papers
Services computing and microservices · 80% Compilers and program optimization · 20%

Topics — the 26 heaviest of 36, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning
data services
0.452010
Graphical XQuery in the aqualogic data services platform · SIGMOD Conference 2010
Access control in the aqualogic data services platform · SIGMOD Conference 2009
Updates in the AquaLogic Data Services Platform · ICDE 2009
Indexing and storage engines
LSM-tree
0.422014
Storage Management in AsterixDB · Proc. VLDB Endow. 2014
AsterixDB: A Scalable, Open Source BDMS · Proc. VLDB Endow. 2014
Data integration and cleaning
enterprise information integration
0.342010
Graphical XQuery in the aqualogic data services platform · SIGMOD Conference 2010
Access control in the aqualogic data services platform · SIGMOD Conference 2009
The BEA AquaLogic data services platform (Demo) · SIGMOD Conference 2006
Transaction processing and concurrency control
ACID transactions
0.212014
Storage Management in AsterixDB · Proc. VLDB Endow. 2014
Graph data management
distributed graph processing
0.212014
Pregelix: Big(ger) Graph Analytics on a Dataflow Engine · Proc. VLDB Endow. 2014
Data models and query languages › query language
semistructured query language
0.212014
AsterixDB: A Scalable, Open Source BDMS · Proc. VLDB Endow. 2014
Indexing and storage engines
storage management
0.212014
Storage Management in AsterixDB · Proc. VLDB Endow. 2014
Data models and query languages › XML query languages
XQuery
0.222009
XQuery Reloaded · Proc. VLDB Endow. 2009
Updates in the AquaLogic Data Services Platform · ICDE 2009
Database system architecture and tuning › main-memory database
distributed in-memory database
0.212013
Scuba: Diving into Data at Facebook · Proc. VLDB Endow. 2013
Data models and query languages › query language
declarative query language
0.112012
ASTERIX: An Open Source System for "Big Data" Management and Analysis · Proc. VLDB Endow. 2012
Database system architecture and tuning
parallel database system
0.112012
ASTERIX: An Open Source System for "Big Data" Management and Analysis · Proc. VLDB Endow. 2012
Query processing and optimization › aggregation
online aggregation
0.112011
Online Aggregation for Large MapReduce Jobs · Proc. VLDB Endow. 2011
Query processing and optimization › XML query processing › XML query optimization
XQuery optimization
0.112009
XQuery Reloaded · Proc. VLDB Endow. 2009
Authentication and access control › access control
fine-grained access control
0.112009
Access control in the aqualogic data services platform · SIGMOD Conference 2009
Data models and query languages
XML query languages
0.112008
XQSE: An XQuery Scripting Extension for the AquaLogic Data Services Platform · ICDE 2008
Information retrieval › cross-language information retrieval
query translation
0.112006
SQL to XQuery Translation in the AquaLogic Data Services Platform · ICDE 2006
Performance modeling and evaluation
performance monitoring
0.012013
Scuba: Diving into Data at Facebook · Proc. VLDB Endow. 2013
Data models and query languages › conceptual modeling
XML conceptual modeling
0.012004
Liquid Data for WebLogic: Integrating Enterprise Data and Services · SIGMOD Conference 2004
Spatial and temporal data management › spatial query processing
geo-spatial query
0.012012
ASTERIX: An Open Source System for "Big Data" Management and Analysis · Proc. VLDB Endow. 2012
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.012011
Hyracks: A flexible and extensible foundation for data-intensive computing · ICDE 2011
Parallel and multicore computing › data-parallel programming
mapreduce
0.012011
Online Aggregation for Large MapReduce Jobs · Proc. VLDB Endow. 2011
Data integration and cleaning › data preprocessing
data cleaning
0.012001
Automatic Segmentation of Text into Structured Records · SIGMOD Conference 2001
Data integration and cleaning › enterprise information integration
enterprise data integration
0.012009
Updates in the AquaLogic Data Services Platform · ICDE 2009
Data integration and cleaning
heterogeneous data source integration
0.012008
XQSE: An XQuery Scripting Extension for the AquaLogic Data Services Platform · ICDE 2008
Data integration and cleaning
heterogeneous data sources
0.012006
SQL to XQuery Translation in the AquaLogic Data Services Platform · ICDE 2006
Services computing and microservices
service-oriented architecture
0.012006
The BEA AquaLogic data services platform (Demo) · SIGMOD Conference 2006

Methods — techniques the papers use, named apart from their topics

in-memory aggregation · 0.3distributed query processing · 0.3partitioned-parallel execution · 0.2DAG-based dataflow · 0.2message passing · 0.2iterative dataflow · 0.2procedural XQuery · 0.2declarative integration logic · 0.2update virtual machine · 0.1update maps · 0.1
YearPublicationVenuePosition
2020 Similarity query support in big data management systems
Taewoo Kim 0001, Wenhai Li, Alexander Behm, Inci Cetindil, Rares Vernica, Vinayak R. Borkar, Michael J. Carey 0001, Chen Li 0001
Inf. Syst.6
2020 Robust and efficient memory management in Apache AsterixDB
abstract
Summary Traditional relational database systems handle data by dividing their memory into sections such as a buffer cache and working memory, assigning a memory budget to each section to efficiently manage a limited amount of overall memory. They also assign memory budgets to memory‐intensive operators such as sorts and joins and control the allocation of memory to these operators; each memory‐intensive operator attempts to maximize its memory usage to reduce disk I/O cost. Implementing such memory‐intensive operators requires a careful design and application of appropriate algorithms that properly utilize memory. Today's Big Data management systems need the ability to handle large amounts of data similarly, as it is unrealistic to assume that truly big data will fit into memory. In this article, we share our memory management experiences in Apache AsterixDB, an open‐source Big Data management software platform that scales out horizontally on shared‐nothing commodity computing clusters. We describe the implementation of AsterixDB's memory‐intensive operators and their designs related to memory management. We also discuss memory management at the global (cluster) level. We conducted an experimental study using several synthetic and real datasets to explore the impact of this work. We believe that future Big Data management system builders can benefit from these experiences.
Taewoo Kim 0001, Alexander Behm, Michael Blow, Vinayak R. Borkar, Yingyi Bu, Michael J. Carey 0001, Murtadha Al Hubail, Shiva Jahangiri, Jianfeng Jia, Chen Li 0001, Chen Luo 0002, Ian Maxon, Pouria Pirzadeh
Softw. Pract. Exp.4
2018 Supporting Similarity Queries in Apache AsterixDB
Taewoo Kim 0001, Wenhai Li, Alexander Behm, Inci Cetindil, Rares Vernica, Vinayak R. Borkar, Michael J. Carey 0001, Chen Li 0001
EDBT6
2015 A scalable parallel XQuery processor
abstract
The wide use of XML for document management and data exchange has created the need to query large repositories of XML data. To efficiently query such large data and take advantage of parallelism, we have implemented Apache VXQuery, an open-source scalable XQuery processor. The system builds upon two other open-source frameworks: Hyracks, a parallel execution engine, and Algebricks, a language agnostic compiler toolbox. Apache VXQuery extends these frameworks and provides an implementation of the XQuery specifics (data model, data-model dependent functions and optimizations, and a parser). We describe the architecture of Apache VXQuery, its integration with Hyracks and Algebricks, and the XQuery optimization rules applied to the query plan to improve path expression efficiency and to enable query parallelism. An experimental evaluation using a real 500GB dataset with various selection, aggregation and join XML queries shows that Apache VXQuery performs well both in terms of scale-up and speed-up. Our experiments show that it is about 3.5x faster than Saxon (an open-source and commercial XQuery processor) on a 4-core, single node implementation, and around 2.5x faster than Apache MRQL (a MapReduce-based parallel query processor) on an eight (4-core) node cluster.
E. Preston Carman Jr., Till Westmann, Vinayak R. Borkar, Michael J. Carey 0001, Vassilis J. Tsotras
IEEE BigData3
2015 External Data Access And Indexing In AsterixDB
abstract
Traditional database systems offer rich query interfaces (SQL) and efficient query execution for data that they store. Recent years have seen the rise of Big Data analytics platforms offering query-based access to "raw" external data, e.g., file-resident data (often in HDFS). In this paper, we describe techniques to achieve the qualities offered by DBMSs when accessing external data. This work has been built into Apache AsterixDB, an open source Big Data Management System. We describe how we build distributed indexes over external data, partition external indexes, provide query consistency across access paths, and manage external indexes amidst concurrent activities. We compare the performance of this new AsterixDB capability to an external-only solution (Hive) and to its internally managed data and indexes.
Abdullah Abdulrahman Alamoudi, Raman Grover, Michael J. Carey 0001, Vinayak R. Borkar
CIKM4
2015 Algebricks: a data model-agnostic compiler backend for big data languages
abstract
A number of high-level query languages, such as Hive, Pig, Flume, and Jaql, have been developed in recent years to increase analyst productivity when processing and analyzing very large datasets. The implementation of each of these languages includes a complete, data model-dependent query compiler, yet each involves a number of similar optimizations. In this work, we describe a new query compiler architecture that separates language-specific and data model-dependent aspects from a more general query compiler backend that can generate executable data-parallel programs for shared-nothing clusters and can be used to develop multiple languages with different data models. We have built such a data model-agnostic query compiler substrate, called Algebricks, and have used it to implement three different query languages --- HiveQL, AQL, and XQuery --- to validate the efficacy of this approach. Experiments show that all three query languages benefit from the parallelization and optimization that Algebricks provides and thus have good parallel speedup and scaleup characteristics for large datasets.
Vinayak R. Borkar, Yingyi Bu, E. Preston Carman Jr., Nicola Onose, Till Westmann, Pouria Pirzadeh, Michael J. Carey 0001, Vassilis J. Tsotras
SoCC1
2014 AsterixDB: A Scalable, Open Source BDMS
abstract
AsterixDB is a new, full-function BDMS (Big Data Management System) with a feature set that distinguishes it from other platforms in today's open source Big Data ecosystem. Its features make it well-suited to applications like web data warehousing, social data storage and analysis, and other use cases related to Big Data. AsterixDB has a flexible NoSQL style data model; a query language that supports a wide range of queries; a scalable runtime; partitioned, LSM-based data storage and indexing (including B + -tree, R-tree, and text indexes); support for external as well as natively stored data; a rich set of built-in types; support for fuzzy, spatial, and temporal types and queries; a built-in notion of data feeds for ingestion of data; and transaction support akin to that of a NoSQL store. Development of AsterixDB began in 2009 and led to a mid-2013 initial open source release. This paper is the first complete description of the resulting open source AsterixDB system. Covered herein are the system's data model, its query language, and its software architecture. Also included are a summary of the current status of the project and a first glimpse into how AsterixDB performs when compared to alternative technologies, including a parallel relational DBMS, a popular NoSQL store, and a popular Hadoop-based SQL data analytics platform, for things that both technologies can do. Also included is a brief description of some initial trials that the system has undergone and the lessons learned (and plans laid) based on those early "customer" engagements.
Sattam Alsubaiee, Yasser Altowim, Hotham Altwaijry, Alexander Behm, Vinayak R. Borkar, Yingyi Bu, Michael J. Carey 0001, Inci Cetindil, Madhusudan Cheelangi, Khurram Faraaz, Eugenia Gabrielova, Raman Grover, Zachary Heilbron, Young-Seok Kim, Chen Li 0001, Guangqiang Li, Ji Mahn Ok, Nicola Onose, Pouria Pirzadeh, Vassilis J. Tsotras, Rares Vernica, Till Westmann
Proc. VLDB Endow.5
2014 Storage Management in AsterixDB
abstract
Social networks, online communities, mobile devices, and instant messaging applications generate complex, unstructured data at a high rate, resulting in large volumes of data. This poses new challenges for data management systems that aim to ingest, store, index, and analyze such data efficiently. In response, we released the first public version of AsterixDB, an open-source Big Data Management System (BDMS), in June of 2013. This paper describes the storage management layer of AsterixDB, providing a detailed description of its ingestion-oriented approach to local storage and a set of initial measurements of its ingestion-related performance characteristics. In order to support high frequency insertions, AsterixDB has wholly adopted Log-Structured Merge-trees as the storage technology for all of its index structures. We describe how the AsterixDB software framework enables "LSM-ification" (conversion from an in-place update, disk-based data structure to a deferred-update, append-only data structure) of any kind of index structure that supports certain primitive operations, enabling the index to ingest data efficiently. We also describe how AsterixDB ensures the ACID properties for operations involving multiple heterogeneous LSM-based indexes. Lastly, we highlight the challenges related to managing the resources of a system when many LSM indexes are used concurrently and present AsterixDB's initial solution.
Sattam Alsubaiee, Alexander Behm, Vinayak R. Borkar, Zachary Heilbron, Young-Seok Kim, Michael J. Carey 0001, Markus Dreseler, Chen Li 0001
Proc. VLDB Endow.3
2014 Pregelix: Big(ger) Graph Analytics on a Dataflow Engine
abstract
There is a growing need for distributed graph processing systems that are capable of gracefully scaling to very large graph datasets. Unfortunately, this challenge has not been easily met due to the intense memory pressure imposed by process-centric, message passing designs that many graph processing systems follow. Pregelix is a new open source distributed graph processing system that is based on an iterative dataflow design that is better tuned to handle both in-memory and out-of-core workloads. As such, Pregelix offers improved performance characteristics and scaling properties over current open source systems (e.g., we have seen up to 15X speedup compared to Apache Giraph and up to 35X speedup compared to distributed GraphLab), and more effective use of available machine resources to support Big(ger) Graph Analytics.
Yingyi Bu, Vinayak R. Borkar, Jianfeng Jia, Michael J. Carey 0001, Tyson Condie
Proc. VLDB Endow.2
2013 A bloat-aware design for big data applications
abstract
Over the past decade, the increasing demands on data-driven business intelligence have led to the proliferation of large-scale, data-intensive applications that often have huge amounts of data (often at terabyte or petabyte scale) to process. An object-oriented programming language such as Java is often the developer's choice for implementing such applications, primarily due to its quick development cycle and rich community resource. While the use of such languages makes programming easier, significant performance problems can often be seen --- the combination of the inefficiencies inherent in a managed run-time system and the impact of the huge amount of data to be processed in the limited memory space often leads to memory bloat and performance degradation at a surprisingly early stage.
Yingyi Bu, Vinayak R. Borkar, Guoqing Harry Xu, Michael J. Carey 0001
ISMM2
2013 Scuba: Diving into Data at Facebook
abstract
Facebook takes performance monitoring seriously. Performance issues can impact over one billion users so we track thousands of servers, hundreds of PB of daily network traffic, hundreds of daily code changes, and many other metrics. We require latencies of under a minute from events occuring (a client request on a phone, a bug report filed, a code change checked in) to graphs showing those events on developers' monitors. Scuba is the data management system Facebook uses for most real-time analysis. Scuba is a fast, scalable, distributed, in-memory database built at Facebook. It currently ingests millions of rows (events) per second and expires data at the same rate. Scuba stores data completely in memory on hundreds of servers each with 144 GB RAM. To process each query, Scuba aggregates data from all servers. Scuba processes almost a million queries per day. Scuba is used extensively for interactive, ad hoc, analysis queries that run in under a second over live data. In addition, Scuba is the workhorse behind Facebook's code regression analysis, bug report monitoring, ads revenue monitoring, and performance debugging.
Lior Abraham, John Allen, Oleksandr Barykin, Vinayak R. Borkar, Bhuwan Chopra, Ciprian Gerea, Daniel Merl, Josh Metzler, David Reiss, Subbu Subramanian, Janet L. Wiener, Okay Zed
Proc. VLDB Endow.4
2012 Inside "Big Data management": ogres, onions, or parfaits?
abstract
In this paper we review the history of systems for managing "Big Data" as well as today's activities and architectures from the (perhaps biased) perspective of three "database guys" who have been watching this space for a number of years and are currently working together on "Big Data" problems. Our focus is on architectural issues, and particularly on the components and layers that have been developed recently (in open source and elsewhere) and on how they are being used (or abused) to tackle challenges posed by today's notion of "Big Data". Also covered is the approach we are taking in the ASTERIX project at UC Irvine, where we are developing our own set of answers to the questions of the "right" components and the "right" set of layers for taming the "Big Data" beast. We close by sharing our opinions on what some of the important open questions are in this area as well as our thoughts on how the dataintensive computing community might best seek out answers.
Vinayak R. Borkar, Michael J. Carey 0001, Chen Li 0001
EDBT1
2012 ASTERIX: An Open Source System for "Big Data" Management and Analysis
abstract
At UC Irvine, we are building a next generation parallel database system, called ASTERIX, as our approach to addressing today's "Big Data" management challenges. ASTERIX aims to combine time-tested principles from parallel database systems with those of the Web-scale computing community, such as fault tolerance for long running jobs. In this demo, we present a whirlwind tour of ASTERIX, highlighting a few of its key features. We will demonstrate examples of our data definition language to model semi-structured data, and examples of interesting queries using our declarative query language. In particular, we will show the capabilities of ASTERIX for answering geo-spatial queries and fuzzy queries, as well as ASTERIX' data feed construct for continuously ingesting data.
Sattam Alsubaiee, Yasser Altowim, Hotham Altwaijry, Alexander Behm, Vinayak R. Borkar, Yingyi Bu, Michael J. Carey 0001, Raman Grover, Zachary Heilbron, Young-Seok Kim, Chen Li 0001, Nicola Onose, Pouria Pirzadeh, Rares Vernica
Proc. VLDB Endow.5
2011 Map-reduce extensions and recursive queries
abstract
We survey the recent wave of extensions to the popular map-reduce systems, including those that have begun to address the implementation of recursive queries using the same computing environment as map-reduce. A central problem is that recursive tasks cannot deliver their output only at the end, which makes recovery from failures much more complicated than in map-reduce and its nonrecursive extensions. We propose several algorithmic ideas for efficient implementation of recursions in the map-reduce environment and discuss several alternatives for supporting recovery from failures without restarting the entire job.
Foto N. Afrati, Vinayak R. Borkar, Michael J. Carey 0001, Neoklis Polyzotis, Jeffrey D. Ullman
EDBT2
2011 Hyracks: A flexible and extensible foundation for data-intensive computing
abstract
Hyracks is a new partitioned-parallel software platform designed to run data-intensive computations on large shared-nothing clusters of computers. Hyracks allows users to express a computation as a DAG of data operators and connectors. Operators operate on partitions of input data and produce partitions of output data, while connectors repartition operators' outputs to make the newly produced partitions available at the consuming operators. We describe the Hyracks end user model, for authors of dataflow jobs, and the extension model for users who wish to augment Hyracks' built-in library with new operator and/or connector types. We also describe our initial Hyracks implementation. Since Hyracks is in roughly the same space as the open source Hadoop platform, we compare Hyracks with Hadoop experimentally for several different kinds of use cases. The initial results demonstrate that Hyracks has significant promise as a next-generation platform for data-intensive applications.
Vinayak R. Borkar, Michael J. Carey 0001, Raman Grover, Nicola Onose, Rares Vernica
ICDE1
2011 ASTERIX: towards a scalable, semistructured data platform for evolving-world models
Alexander Behm, Vinayak R. Borkar, Michael J. Carey 0001, Raman Grover, Chen Li 0001, Nicola Onose, Rares Vernica, Alin Deutsch, Yannis Papakonstantinou, Vassilis J. Tsotras
Distributed Parallel Databases2
2011 Online Aggregation for Large MapReduce Jobs
Niketan Pansare, Vinayak R. Borkar, Chris Jermaine, Tyson Condie
Proc. VLDB Endow.2
2010 Graphical XQuery in the aqualogic data services platform
abstract
The AquaLogic Data Services Platform (ALDSP) is a middleware platform developed at BEA Systems for building services, referred to as data services, that integrate, access, and manipulate information coming from multiple heterogeneous sources of data (including databases, files, and other services). ALDSP uses functions that produce and consume XML to model heterogeneous information sources, so the integration logic for data access services in ALDSP is specified declaratively using the XQuery language. A challenge that we faced in developing ALDSP was providing effective graphical tooling to help data service developers to develop information integration queries. In this paper, we describe the graphical XQuery Editor (XQE) that resulted from our attempt to tackle this challenge. XQE handles the full XQuery language and provides a robust two-way editing experience involving both graphical and source views of each query. XQE is novel in being the first commercial graphical XQuery editor to support both of these features.
Vinayak R. Borkar, Michael J. Carey 0001, Sebu Koleth, Alexander Kotopoulis, Kautul Mehta, Joshua Spiegel, Sachin Thatte, Till Westmann
SIGMOD Conference1
2009 Updates in the AquaLogic Data Services Platform
abstract
The BEA aqualogic data services platform (ALDSP) is a middleware platform for creating services that integrate and manipulate information from disparate enterprise data sources. This paper provides a technical overview of the all-new update support in ALDSP 3.0, released in January 2008. It describes the update side of data services, our unique model for making update automation transparent and flexible, and the use of the XQuery Scripting Extension (XQSE) for further customizing the system's default handling of updates. It also gives an overview of the ALDSP update processing machinery, including the automatic generation of update maps from read functions, translation of update maps into update virtual machine (UVM) programs, the UVM instruction interpreter, and SQL generation for updates to data drawn from relational data sources.
Michael Blow, Vinayak R. Borkar, Michael J. Carey 0001, Chris Hillery, Alexander Kotopoulis, Dmitry Lychagin, Radu Preotiuc-Pietro, Panagiotis Reveliotis, Joshua Spiegel, Till Westmann
ICDE2
2009 Access control in the aqualogic data services platform
abstract
The AquaLogic Data Services Platform (ALDSP) is a middleware platform for building data services that integrate and provide operations over data drawn from spanning multiple heterogeneous information sources. A data service consists of an XML Schema instance, describing its information content, and a collection of XQuery functions and procedures that comprise its set of operations. This paper describes access control in ALDSP. We describe ALDSP's securable resource hierarchy, its fine-grained access control capabilities for securing portions of data service schemas, how XQuery can be used to specify data-driven security policies, and how user identity mapping is supported. We then provide an in-depth overview of how ALDSP works, including implementation techniques to keep access control checking from interacting badly with view rewriting, query optimization, and caching.
Vinayak R. Borkar, Michael J. Carey 0001, Daniel Engovatov, Dmitry Lychagin, Panagiotis Reveliotis, Joshua Spiegel, Sachin Thatte, Till Westmann
SIGMOD Conference1
2009 XQuery Reloaded
abstract
This paper describes a number of XQuery-related projects. Its goal is to show that XQuery is a useful tool for many different application scenarios. In particular, this paper tries to correct a common myth that XQuery is merely a query language and that SQL is the better query language. Instead, XQuery is a full-fledged programming language for Web applications and services. Furthermore, this paper tries to correct a second myth that XQuery is slow. This paper gives an overview of the state-of-the-art in XQuery implementation and optimization techniques and discusses one particular open-source XQuery processor, Zorba, in more detail. Among others, this paper presents an XQuery Benchmark Service which helps practitioners and XQuery processor vendors to find performance problems in an XQuery processor.
Roger Bamford, Vinayak R. Borkar, Matthias Brantner, Peter M. Fischer 0001, Daniela Florescu, David A. Graf, Donald Kossmann, Tim Kraska, Dan Muresan, Sorin Nasoi, Markos Zacharioudaki
Proc. VLDB Endow.2
2008 XQSE: An XQuery Scripting Extension for the AquaLogic Data Services Platform
abstract
The AquaLogic Data Services Platform (ALDSP) is a BEA middleware platform for creating services that access and manipulate information drawn from multiple heterogeneous sources of data. The integration logic for read services is specified declaratively using the XQuery language. ALDSP 3.0, available in December 2007, includes a new XQuery-based Scripting Extension - XQSE - that enables developers to write procedural as well as declarative logic without leaving the XQuery world. In this paper, we describe the XQSE extensions to XQuery and show how they help to support important new classes of data services in ALDSP 3.0.
Vinayak R. Borkar, Michael J. Carey 0001, Daniel Engovatov, Dmitry Lychagin, Till Westmann, Warren Wong
ICDE1
2007 Inverse Functions in the AquaLogic Data Services Platform
Nicola Onose, Vinayak R. Borkar, Michael J. Carey 0001
VLDB2
2006 SQL to XQuery Translation in the AquaLogic Data Services Platform
abstract
SQL has long been the standard language for retrieving and manipulating data in relational database systems. XML has become the standard format for data exchange, and XQuery is on its way to becoming the standard language for querying XML data. The BEA AquaLogic Data Services Platform provides a service-oriented, XML-based view of heterogeneous enterprise data sources and allows this view to be queried using XQuery. AquaLogic DSP includes a JDBC driver that connects the old (SQL) world with the new (XML) world via a SQL-to-XQuery translator. This paper outlines the issues related to creating such a driver and details the approach used to translate SQL queries into XQuery expressions. The paper also touches on performance considerations related to handling XML query results in a context where JDBC result sets are the desired output format.
Sunil Jigyasu, Sujeet Banerjee, Vinayak R. Borkar, Michael J. Carey 0001, Kanad Dixit, Anil Malkani, Sachin Thatte
ICDE3
2006 The BEA AquaLogic data services platform (Demo)
abstract
We showcase the BEA AquaLogic Data Services Platform (ALDSP), a middleware infrastructure product that enables the declarative development of data services for service-oriented architectures (SOA). ALDSP includes support for modeling networks of interrelated data services, for realizing data services using either graphical or source-based XQuery editors, for testing data services as they are developed, and for identifying and incorporating changes in the structure of the underlying sources of data. Physical data sources supported include relational tables and views, Web services, packaged applications, stored procedures, XML files, delimited files, and custom Java applications. Data service definitions can be layered; as with relational views, such layering is virtual, and is rewritten away at query compilation time. ALDSP supports both read and update data service functions, and the ALDSP XML query runtime includes a number of interesting query operators and distributed query optimizations. In addition, ALDSP supports function caching, fine-grained security, and SQL-based data access as well as providing service-based and XQuery access to SOA data. We plan to demonstrate as much of this as time permits.
Vinayak R. Borkar, Michael J. Carey 0001, Dmitry Lychagin, Till Westmann
SIGMOD Conference1
2006 Query Processing in the AquaLogic Data Services Platform
Vinayak R. Borkar, Michael J. Carey 0001, Dmitry Lychagin, Till Westmann, Daniel Engovatov, Nicola Onose
VLDB1
2004 Liquid Data for WebLogic: Integrating Enterprise Data and Services
abstract
Information in today's enterprises commonly resides in a variety of heterogeneous data sources, including relational databases, web services, files, packaged applications, and custom data repositories. BEA's enterprise information integration product, Liquid Data for WebLogic, takes an XML-based approach to providing integrated access to such heterogeneous information. This demonstration highlights the XML technologies involved - including web services, XQuery, and XML Schema - and shows how they can be brought to bear on the enterprise information integration problem. The demonstration uses a simple end-to-end example, one that involves integrating data from relational databases and web services, to walk the audience through the overall architecture, XML-based data modeling approach, programming model, declarative query and view facilities, and distributed processing features of Liquid Data.
Vinayak R. Borkar
SIGMOD Conference1
2003 XML queries and algebra in the Enosys integration platform
Yannis Papakonstantinou, Vinayak R. Borkar, Maxim Orgiyan, Konstantinos Stathatos, Lucian Suta, Vasilis Vassalos, Pavel E. Velikhov
Data Knowl. Eng.2
2001 Automatic Segmentation of Text into Structured Records
abstract
In this paper we present a method for automatically segmenting unformatted text records into structured elements. Several useful data sources today are human-generated as continuous text whereas convenient usage requires the data to be organized as structured records. A prime motivation is the warehouse address cleaning problem of transforming dirty addresses stored in large corporate databases as a single text field into subfields like “City” and “Street”. Existing tools rely on hand-tuned, domain-specific rule-based systems.
Vinayak R. Borkar, Kaustubh Deshmukh, Sunita Sarawagi
SIGMOD Conference1