EDBT 2026 Demo / reviewers in the wild / expert
Vinayak R. Borkar
dblp:43/3643
· DBLP profile ↗
29ranked-venue papers
10as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 26 · 9 first-authorArtificial intelligence and machine learning · 3Software engineering, systems software and programming languages · 2Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
18 papers |
Data integration and cleaning · 21% Data models and query languages · 16% Database system architecture and tuning · 14% | |
| Software engineering, system software, and programming languages
4 papers |
Services computing and microservices · 80% Compilers and program optimization · 20% |
Topics — the 26 heaviest of 36, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data integration and cleaning
data services |
0.4 | 5 | 2010 | Graphical XQuery in the aqualogic data services platform · SIGMOD Conference 2010 Access control in the aqualogic data services platform · SIGMOD Conference 2009 Updates in the AquaLogic Data Services Platform · ICDE 2009 |
Indexing and storage engines
LSM-tree |
0.4 | 2 | 2014 | Storage Management in AsterixDB · Proc. VLDB Endow. 2014 AsterixDB: A Scalable, Open Source BDMS · Proc. VLDB Endow. 2014 |
Data integration and cleaning
enterprise information integration |
0.3 | 4 | 2010 | Graphical XQuery in the aqualogic data services platform · SIGMOD Conference 2010 Access control in the aqualogic data services platform · SIGMOD Conference 2009 The BEA AquaLogic data services platform (Demo) · SIGMOD Conference 2006 |
Transaction processing and concurrency control
ACID transactions |
0.2 | 1 | 2014 | Storage Management in AsterixDB · Proc. VLDB Endow. 2014 |
Graph data management
distributed graph processing |
0.2 | 1 | 2014 | Pregelix: Big(ger) Graph Analytics on a Dataflow Engine · Proc. VLDB Endow. 2014 |
Data models and query languages › query language
semistructured query language |
0.2 | 1 | 2014 | AsterixDB: A Scalable, Open Source BDMS · Proc. VLDB Endow. 2014 |
Indexing and storage engines
storage management |
0.2 | 1 | 2014 | Storage Management in AsterixDB · Proc. VLDB Endow. 2014 |
Data models and query languages › XML query languages
XQuery |
0.2 | 2 | 2009 | XQuery Reloaded · Proc. VLDB Endow. 2009 Updates in the AquaLogic Data Services Platform · ICDE 2009 |
Database system architecture and tuning › main-memory database
distributed in-memory database |
0.2 | 1 | 2013 | Scuba: Diving into Data at Facebook · Proc. VLDB Endow. 2013 |
Data models and query languages › query language
declarative query language |
0.1 | 1 | 2012 | ASTERIX: An Open Source System for "Big Data" Management and Analysis · Proc. VLDB Endow. 2012 |
Database system architecture and tuning
parallel database system |
0.1 | 1 | 2012 | ASTERIX: An Open Source System for "Big Data" Management and Analysis · Proc. VLDB Endow. 2012 |
Query processing and optimization › aggregation
online aggregation |
0.1 | 1 | 2011 | Online Aggregation for Large MapReduce Jobs · Proc. VLDB Endow. 2011 |
Query processing and optimization › XML query processing › XML query optimization
XQuery optimization |
0.1 | 1 | 2009 | XQuery Reloaded · Proc. VLDB Endow. 2009 |
Authentication and access control › access control
fine-grained access control |
0.1 | 1 | 2009 | Access control in the aqualogic data services platform · SIGMOD Conference 2009 |
Data models and query languages
XML query languages |
0.1 | 1 | 2008 | XQSE: An XQuery Scripting Extension for the AquaLogic Data Services Platform · ICDE 2008 |
Information retrieval › cross-language information retrieval
query translation |
0.1 | 1 | 2006 | SQL to XQuery Translation in the AquaLogic Data Services Platform · ICDE 2006 |
Performance modeling and evaluation
performance monitoring |
0.0 | 1 | 2013 | Scuba: Diving into Data at Facebook · Proc. VLDB Endow. 2013 |
Data models and query languages › conceptual modeling
XML conceptual modeling |
0.0 | 1 | 2004 | Liquid Data for WebLogic: Integrating Enterprise Data and Services · SIGMOD Conference 2004 |
Spatial and temporal data management › spatial query processing
geo-spatial query |
0.0 | 1 | 2012 | ASTERIX: An Open Source System for "Big Data" Management and Analysis · Proc. VLDB Endow. 2012 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.0 | 1 | 2011 | Hyracks: A flexible and extensible foundation for data-intensive computing · ICDE 2011 |
Parallel and multicore computing › data-parallel programming
mapreduce |
0.0 | 1 | 2011 | Online Aggregation for Large MapReduce Jobs · Proc. VLDB Endow. 2011 |
Data integration and cleaning › data preprocessing
data cleaning |
0.0 | 1 | 2001 | Automatic Segmentation of Text into Structured Records · SIGMOD Conference 2001 |
Data integration and cleaning › enterprise information integration
enterprise data integration |
0.0 | 1 | 2009 | Updates in the AquaLogic Data Services Platform · ICDE 2009 |
Data integration and cleaning
heterogeneous data source integration |
0.0 | 1 | 2008 | XQSE: An XQuery Scripting Extension for the AquaLogic Data Services Platform · ICDE 2008 |
Data integration and cleaning
heterogeneous data sources |
0.0 | 1 | 2006 | SQL to XQuery Translation in the AquaLogic Data Services Platform · ICDE 2006 |
Services computing and microservices
service-oriented architecture |
0.0 | 1 | 2006 | The BEA AquaLogic data services platform (Demo) · SIGMOD Conference 2006 |
Methods — techniques the papers use, named apart from their topics
in-memory aggregation · 0.3distributed query processing · 0.3partitioned-parallel execution · 0.2DAG-based dataflow · 0.2message passing · 0.2iterative dataflow · 0.2procedural XQuery · 0.2declarative integration logic · 0.2update virtual machine · 0.1update maps · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Similarity query support in big data management systems
Taewoo Kim 0001, Wenhai Li, Alexander Behm, Inci Cetindil, Rares Vernica, Vinayak R. Borkar, Michael J. Carey 0001, Chen Li 0001 |
Inf. Syst. | 6 |
| 2020 | Robust and efficient memory management in Apache AsterixDBabstractSummary Traditional relational database systems handle data by dividing their memory into sections such as a buffer cache and working memory, assigning a memory budget to each section to efficiently manage a limited amount of overall memory. They also assign memory budgets to memory‐intensive operators such as sorts and joins and control the allocation of memory to these operators; each memory‐intensive operator attempts to maximize its memory usage to reduce disk I/O cost. Implementing such memory‐intensive operators requires a careful design and application of appropriate algorithms that properly utilize memory. Today's Big Data management systems need the ability to handle large amounts of data similarly, as it is unrealistic to assume that truly big data will fit into memory. In this article, we share our memory management experiences in Apache AsterixDB, an open‐source Big Data management software platform that scales out horizontally on shared‐nothing commodity computing clusters. We describe the implementation of AsterixDB's memory‐intensive operators and their designs related to memory management. We also discuss memory management at the global (cluster) level. We conducted an experimental study using several synthetic and real datasets to explore the impact of this work. We believe that future Big Data management system builders can benefit from these experiences. Taewoo Kim 0001, Alexander Behm, Michael Blow, Vinayak R. Borkar, Yingyi Bu, Michael J. Carey 0001, Murtadha Al Hubail, Shiva Jahangiri, Jianfeng Jia, Chen Li 0001, Chen Luo 0002, Ian Maxon, Pouria Pirzadeh |
Softw. Pract. Exp. | 4 |
| 2018 | Supporting Similarity Queries in Apache AsterixDB
Taewoo Kim 0001, Wenhai Li, Alexander Behm, Inci Cetindil, Rares Vernica, Vinayak R. Borkar, Michael J. Carey 0001, Chen Li 0001 |
EDBT | 6 |
| 2015 | A scalable parallel XQuery processorabstractThe wide use of XML for document management and data exchange has created the need to query large repositories of XML data. To efficiently query such large data and take advantage of parallelism, we have implemented Apache VXQuery, an open-source scalable XQuery processor. The system builds upon two other open-source frameworks: Hyracks, a parallel execution engine, and Algebricks, a language agnostic compiler toolbox. Apache VXQuery extends these frameworks and provides an implementation of the XQuery specifics (data model, data-model dependent functions and optimizations, and a parser). We describe the architecture of Apache VXQuery, its integration with Hyracks and Algebricks, and the XQuery optimization rules applied to the query plan to improve path expression efficiency and to enable query parallelism. An experimental evaluation using a real 500GB dataset with various selection, aggregation and join XML queries shows that Apache VXQuery performs well both in terms of scale-up and speed-up. Our experiments show that it is about 3.5x faster than Saxon (an open-source and commercial XQuery processor) on a 4-core, single node implementation, and around 2.5x faster than Apache MRQL (a MapReduce-based parallel query processor) on an eight (4-core) node cluster. E. Preston Carman Jr., Till Westmann, Vinayak R. Borkar, Michael J. Carey 0001, Vassilis J. Tsotras |
IEEE BigData | 3 |
| 2015 | External Data Access And Indexing In AsterixDBabstractTraditional database systems offer rich query interfaces (SQL) and efficient query execution for data that they store. Recent years have seen the rise of Big Data analytics platforms offering query-based access to "raw" external data, e.g., file-resident data (often in HDFS). In this paper, we describe techniques to achieve the qualities offered by DBMSs when accessing external data. This work has been built into Apache AsterixDB, an open source Big Data Management System. We describe how we build distributed indexes over external data, partition external indexes, provide query consistency across access paths, and manage external indexes amidst concurrent activities. We compare the performance of this new AsterixDB capability to an external-only solution (Hive) and to its internally managed data and indexes. Abdullah Abdulrahman Alamoudi, Raman Grover, Michael J. Carey 0001, Vinayak R. Borkar |
CIKM | 4 |
| 2015 | Algebricks: a data model-agnostic compiler backend for big data languagesabstractA number of high-level query languages, such as Hive, Pig, Flume, and Jaql, have been developed in recent years to increase analyst productivity when processing and analyzing very large datasets. The implementation of each of these languages includes a complete, data model-dependent query compiler, yet each involves a number of similar optimizations. In this work, we describe a new query compiler architecture that separates language-specific and data model-dependent aspects from a more general query compiler backend that can generate executable data-parallel programs for shared-nothing clusters and can be used to develop multiple languages with different data models. We have built such a data model-agnostic query compiler substrate, called Algebricks, and have used it to implement three different query languages --- HiveQL, AQL, and XQuery --- to validate the efficacy of this approach. Experiments show that all three query languages benefit from the parallelization and optimization that Algebricks provides and thus have good parallel speedup and scaleup characteristics for large datasets. Vinayak R. Borkar, Yingyi Bu, E. Preston Carman Jr., Nicola Onose, Till Westmann, Pouria Pirzadeh, Michael J. Carey 0001, Vassilis J. Tsotras |
SoCC | 1 |
| 2014 | AsterixDB: A Scalable, Open Source BDMSabstractAsterixDB is a new, full-function BDMS (Big Data Management System) with a feature set that distinguishes it from other platforms in today's open source Big Data ecosystem. Its features make it well-suited to applications like web data warehousing, social data storage and analysis, and other use cases related to Big Data. AsterixDB has a flexible NoSQL style data model; a query language that supports a wide range of queries; a scalable runtime; partitioned, LSM-based data storage and indexing (including B + -tree, R-tree, and text indexes); support for external as well as natively stored data; a rich set of built-in types; support for fuzzy, spatial, and temporal types and queries; a built-in notion of data feeds for ingestion of data; and transaction support akin to that of a NoSQL store. Development of AsterixDB began in 2009 and led to a mid-2013 initial open source release. This paper is the first complete description of the resulting open source AsterixDB system. Covered herein are the system's data model, its query language, and its software architecture. Also included are a summary of the current status of the project and a first glimpse into how AsterixDB performs when compared to alternative technologies, including a parallel relational DBMS, a popular NoSQL store, and a popular Hadoop-based SQL data analytics platform, for things that both technologies can do. Also included is a brief description of some initial trials that the system has undergone and the lessons learned (and plans laid) based on those early "customer" engagements. Sattam Alsubaiee, Yasser Altowim, Hotham Altwaijry, Alexander Behm, Vinayak R. Borkar, Yingyi Bu, Michael J. Carey 0001, Inci Cetindil, Madhusudan Cheelangi, Khurram Faraaz, Eugenia Gabrielova, Raman Grover, Zachary Heilbron, Young-Seok Kim, Chen Li 0001, Guangqiang Li, Ji Mahn Ok, Nicola Onose, Pouria Pirzadeh, Vassilis J. Tsotras, Rares Vernica, Till Westmann |
Proc. VLDB Endow. | 5 |
| 2014 | Storage Management in AsterixDBabstractSocial networks, online communities, mobile devices, and instant messaging applications generate complex, unstructured data at a high rate, resulting in large volumes of data. This poses new challenges for data management systems that aim to ingest, store, index, and analyze such data efficiently. In response, we released the first public version of AsterixDB, an open-source Big Data Management System (BDMS), in June of 2013. This paper describes the storage management layer of AsterixDB, providing a detailed description of its ingestion-oriented approach to local storage and a set of initial measurements of its ingestion-related performance characteristics. In order to support high frequency insertions, AsterixDB has wholly adopted Log-Structured Merge-trees as the storage technology for all of its index structures. We describe how the AsterixDB software framework enables "LSM-ification" (conversion from an in-place update, disk-based data structure to a deferred-update, append-only data structure) of any kind of index structure that supports certain primitive operations, enabling the index to ingest data efficiently. We also describe how AsterixDB ensures the ACID properties for operations involving multiple heterogeneous LSM-based indexes. Lastly, we highlight the challenges related to managing the resources of a system when many LSM indexes are used concurrently and present AsterixDB's initial solution. Sattam Alsubaiee, Alexander Behm, Vinayak R. Borkar, Zachary Heilbron, Young-Seok Kim, Michael J. Carey 0001, Markus Dreseler, Chen Li 0001 |
Proc. VLDB Endow. | 3 |
| 2014 | Pregelix: Big(ger) Graph Analytics on a Dataflow EngineabstractThere is a growing need for distributed graph processing systems that are capable of gracefully scaling to very large graph datasets. Unfortunately, this challenge has not been easily met due to the intense memory pressure imposed by process-centric, message passing designs that many graph processing systems follow. Pregelix is a new open source distributed graph processing system that is based on an iterative dataflow design that is better tuned to handle both in-memory and out-of-core workloads. As such, Pregelix offers improved performance characteristics and scaling properties over current open source systems (e.g., we have seen up to 15X speedup compared to Apache Giraph and up to 35X speedup compared to distributed GraphLab), and more effective use of available machine resources to support Big(ger) Graph Analytics. Yingyi Bu, Vinayak R. Borkar, Jianfeng Jia, Michael J. Carey 0001, Tyson Condie |
Proc. VLDB Endow. | 2 |
| 2013 | A bloat-aware design for big data applicationsabstractOver the past decade, the increasing demands on data-driven business intelligence have led to the proliferation of large-scale, data-intensive applications that often have huge amounts of data (often at terabyte or petabyte scale) to process. An object-oriented programming language such as Java is often the developer's choice for implementing such applications, primarily due to its quick development cycle and rich community resource. While the use of such languages makes programming easier, significant performance problems can often be seen --- the combination of the inefficiencies inherent in a managed run-time system and the impact of the huge amount of data to be processed in the limited memory space often leads to memory bloat and performance degradation at a surprisingly early stage. Yingyi Bu, Vinayak R. Borkar, Guoqing Harry Xu, Michael J. Carey 0001 |
ISMM | 2 |
| 2013 | Scuba: Diving into Data at FacebookabstractFacebook takes performance monitoring seriously. Performance issues can impact over one billion users so we track thousands of servers, hundreds of PB of daily network traffic, hundreds of daily code changes, and many other metrics. We require latencies of under a minute from events occuring (a client request on a phone, a bug report filed, a code change checked in) to graphs showing those events on developers' monitors. Scuba is the data management system Facebook uses for most real-time analysis. Scuba is a fast, scalable, distributed, in-memory database built at Facebook. It currently ingests millions of rows (events) per second and expires data at the same rate. Scuba stores data completely in memory on hundreds of servers each with 144 GB RAM. To process each query, Scuba aggregates data from all servers. Scuba processes almost a million queries per day. Scuba is used extensively for interactive, ad hoc, analysis queries that run in under a second over live data. In addition, Scuba is the workhorse behind Facebook's code regression analysis, bug report monitoring, ads revenue monitoring, and performance debugging. Lior Abraham, John Allen, Oleksandr Barykin, Vinayak R. Borkar, Bhuwan Chopra, Ciprian Gerea, Daniel Merl, Josh Metzler, David Reiss, Subbu Subramanian, Janet L. Wiener, Okay Zed |
Proc. VLDB Endow. | 4 |
| 2012 | Inside "Big Data management": ogres, onions, or parfaits?abstractIn this paper we review the history of systems for managing "Big Data" as well as today's activities and architectures from the (perhaps biased) perspective of three "database guys" who have been watching this space for a number of years and are currently working together on "Big Data" problems. Our focus is on architectural issues, and particularly on the components and layers that have been developed recently (in open source and elsewhere) and on how they are being used (or abused) to tackle challenges posed by today's notion of "Big Data". Also covered is the approach we are taking in the ASTERIX project at UC Irvine, where we are developing our own set of answers to the questions of the "right" components and the "right" set of layers for taming the "Big Data" beast. We close by sharing our opinions on what some of the important open questions are in this area as well as our thoughts on how the dataintensive computing community might best seek out answers. Vinayak R. Borkar, Michael J. Carey 0001, Chen Li 0001 |
EDBT | 1 |
| 2012 | ASTERIX: An Open Source System for "Big Data" Management and AnalysisabstractAt UC Irvine, we are building a next generation parallel database system, called ASTERIX, as our approach to addressing today's "Big Data" management challenges. ASTERIX aims to combine time-tested principles from parallel database systems with those of the Web-scale computing community, such as fault tolerance for long running jobs. In this demo, we present a whirlwind tour of ASTERIX, highlighting a few of its key features. We will demonstrate examples of our data definition language to model semi-structured data, and examples of interesting queries using our declarative query language. In particular, we will show the capabilities of ASTERIX for answering geo-spatial queries and fuzzy queries, as well as ASTERIX' data feed construct for continuously ingesting data. Sattam Alsubaiee, Yasser Altowim, Hotham Altwaijry, Alexander Behm, Vinayak R. Borkar, Yingyi Bu, Michael J. Carey 0001, Raman Grover, Zachary Heilbron, Young-Seok Kim, Chen Li 0001, Nicola Onose, Pouria Pirzadeh, Rares Vernica |
Proc. VLDB Endow. | 5 |
| 2011 | Map-reduce extensions and recursive queriesabstractWe survey the recent wave of extensions to the popular map-reduce systems, including those that have begun to address the implementation of recursive queries using the same computing environment as map-reduce. A central problem is that recursive tasks cannot deliver their output only at the end, which makes recovery from failures much more complicated than in map-reduce and its nonrecursive extensions. We propose several algorithmic ideas for efficient implementation of recursions in the map-reduce environment and discuss several alternatives for supporting recovery from failures without restarting the entire job. Foto N. Afrati, Vinayak R. Borkar, Michael J. Carey 0001, Neoklis Polyzotis, Jeffrey D. Ullman |
EDBT | 2 |
| 2011 | Hyracks: A flexible and extensible foundation for data-intensive computingabstractHyracks is a new partitioned-parallel software platform designed to run data-intensive computations on large shared-nothing clusters of computers. Hyracks allows users to express a computation as a DAG of data operators and connectors. Operators operate on partitions of input data and produce partitions of output data, while connectors repartition operators' outputs to make the newly produced partitions available at the consuming operators. We describe the Hyracks end user model, for authors of dataflow jobs, and the extension model for users who wish to augment Hyracks' built-in library with new operator and/or connector types. We also describe our initial Hyracks implementation. Since Hyracks is in roughly the same space as the open source Hadoop platform, we compare Hyracks with Hadoop experimentally for several different kinds of use cases. The initial results demonstrate that Hyracks has significant promise as a next-generation platform for data-intensive applications. Vinayak R. Borkar, Michael J. Carey 0001, Raman Grover, Nicola Onose, Rares Vernica |
ICDE | 1 |
| 2011 | ASTERIX: towards a scalable, semistructured data platform for evolving-world models
Alexander Behm, Vinayak R. Borkar, Michael J. Carey 0001, Raman Grover, Chen Li 0001, Nicola Onose, Rares Vernica, Alin Deutsch, Yannis Papakonstantinou, Vassilis J. Tsotras |
Distributed Parallel Databases | 2 |
| 2011 | Online Aggregation for Large MapReduce Jobs
Niketan Pansare, Vinayak R. Borkar, Chris Jermaine, Tyson Condie |
Proc. VLDB Endow. | 2 |
| 2010 | Graphical XQuery in the aqualogic data services platformabstractThe AquaLogic Data Services Platform (ALDSP) is a middleware platform developed at BEA Systems for building services, referred to as data services, that integrate, access, and manipulate information coming from multiple heterogeneous sources of data (including databases, files, and other services). ALDSP uses functions that produce and consume XML to model heterogeneous information sources, so the integration logic for data access services in ALDSP is specified declaratively using the XQuery language. A challenge that we faced in developing ALDSP was providing effective graphical tooling to help data service developers to develop information integration queries. In this paper, we describe the graphical XQuery Editor (XQE) that resulted from our attempt to tackle this challenge. XQE handles the full XQuery language and provides a robust two-way editing experience involving both graphical and source views of each query. XQE is novel in being the first commercial graphical XQuery editor to support both of these features. Vinayak R. Borkar, Michael J. Carey 0001, Sebu Koleth, Alexander Kotopoulis, Kautul Mehta, Joshua Spiegel, Sachin Thatte, Till Westmann |
SIGMOD Conference | 1 |
| 2009 | Updates in the AquaLogic Data Services PlatformabstractThe BEA aqualogic data services platform (ALDSP) is a middleware platform for creating services that integrate and manipulate information from disparate enterprise data sources. This paper provides a technical overview of the all-new update support in ALDSP 3.0, released in January 2008. It describes the update side of data services, our unique model for making update automation transparent and flexible, and the use of the XQuery Scripting Extension (XQSE) for further customizing the system's default handling of updates. It also gives an overview of the ALDSP update processing machinery, including the automatic generation of update maps from read functions, translation of update maps into update virtual machine (UVM) programs, the UVM instruction interpreter, and SQL generation for updates to data drawn from relational data sources. Michael Blow, Vinayak R. Borkar, Michael J. Carey 0001, Chris Hillery, Alexander Kotopoulis, Dmitry Lychagin, Radu Preotiuc-Pietro, Panagiotis Reveliotis, Joshua Spiegel, Till Westmann |
ICDE | 2 |
| 2009 | Access control in the aqualogic data services platformabstractThe AquaLogic Data Services Platform (ALDSP) is a middleware platform for building data services that integrate and provide operations over data drawn from spanning multiple heterogeneous information sources. A data service consists of an XML Schema instance, describing its information content, and a collection of XQuery functions and procedures that comprise its set of operations. This paper describes access control in ALDSP. We describe ALDSP's securable resource hierarchy, its fine-grained access control capabilities for securing portions of data service schemas, how XQuery can be used to specify data-driven security policies, and how user identity mapping is supported. We then provide an in-depth overview of how ALDSP works, including implementation techniques to keep access control checking from interacting badly with view rewriting, query optimization, and caching. Vinayak R. Borkar, Michael J. Carey 0001, Daniel Engovatov, Dmitry Lychagin, Panagiotis Reveliotis, Joshua Spiegel, Sachin Thatte, Till Westmann |
SIGMOD Conference | 1 |
| 2009 | XQuery ReloadedabstractThis paper describes a number of XQuery-related projects. Its goal is to show that XQuery is a useful tool for many different application scenarios. In particular, this paper tries to correct a common myth that XQuery is merely a query language and that SQL is the better query language. Instead, XQuery is a full-fledged programming language for Web applications and services. Furthermore, this paper tries to correct a second myth that XQuery is slow. This paper gives an overview of the state-of-the-art in XQuery implementation and optimization techniques and discusses one particular open-source XQuery processor, Zorba, in more detail. Among others, this paper presents an XQuery Benchmark Service which helps practitioners and XQuery processor vendors to find performance problems in an XQuery processor. Roger Bamford, Vinayak R. Borkar, Matthias Brantner, Peter M. Fischer 0001, Daniela Florescu, David A. Graf, Donald Kossmann, Tim Kraska, Dan Muresan, Sorin Nasoi, Markos Zacharioudaki |
Proc. VLDB Endow. | 2 |
| 2008 | XQSE: An XQuery Scripting Extension for the AquaLogic Data Services PlatformabstractThe AquaLogic Data Services Platform (ALDSP) is a BEA middleware platform for creating services that access and manipulate information drawn from multiple heterogeneous sources of data. The integration logic for read services is specified declaratively using the XQuery language. ALDSP 3.0, available in December 2007, includes a new XQuery-based Scripting Extension - XQSE - that enables developers to write procedural as well as declarative logic without leaving the XQuery world. In this paper, we describe the XQSE extensions to XQuery and show how they help to support important new classes of data services in ALDSP 3.0. Vinayak R. Borkar, Michael J. Carey 0001, Daniel Engovatov, Dmitry Lychagin, Till Westmann, Warren Wong |
ICDE | 1 |
| 2007 | Inverse Functions in the AquaLogic Data Services Platform
Nicola Onose, Vinayak R. Borkar, Michael J. Carey 0001 |
VLDB | 2 |
| 2006 | SQL to XQuery Translation in the AquaLogic Data Services PlatformabstractSQL has long been the standard language for retrieving and manipulating data in relational database systems. XML has become the standard format for data exchange, and XQuery is on its way to becoming the standard language for querying XML data. The BEA AquaLogic Data Services Platform provides a service-oriented, XML-based view of heterogeneous enterprise data sources and allows this view to be queried using XQuery. AquaLogic DSP includes a JDBC driver that connects the old (SQL) world with the new (XML) world via a SQL-to-XQuery translator. This paper outlines the issues related to creating such a driver and details the approach used to translate SQL queries into XQuery expressions. The paper also touches on performance considerations related to handling XML query results in a context where JDBC result sets are the desired output format. Sunil Jigyasu, Sujeet Banerjee, Vinayak R. Borkar, Michael J. Carey 0001, Kanad Dixit, Anil Malkani, Sachin Thatte |
ICDE | 3 |
| 2006 | The BEA AquaLogic data services platform (Demo)abstractWe showcase the BEA AquaLogic Data Services Platform (ALDSP), a middleware infrastructure product that enables the declarative development of data services for service-oriented architectures (SOA). ALDSP includes support for modeling networks of interrelated data services, for realizing data services using either graphical or source-based XQuery editors, for testing data services as they are developed, and for identifying and incorporating changes in the structure of the underlying sources of data. Physical data sources supported include relational tables and views, Web services, packaged applications, stored procedures, XML files, delimited files, and custom Java applications. Data service definitions can be layered; as with relational views, such layering is virtual, and is rewritten away at query compilation time. ALDSP supports both read and update data service functions, and the ALDSP XML query runtime includes a number of interesting query operators and distributed query optimizations. In addition, ALDSP supports function caching, fine-grained security, and SQL-based data access as well as providing service-based and XQuery access to SOA data. We plan to demonstrate as much of this as time permits. Vinayak R. Borkar, Michael J. Carey 0001, Dmitry Lychagin, Till Westmann |
SIGMOD Conference | 1 |
| 2006 | Query Processing in the AquaLogic Data Services Platform
Vinayak R. Borkar, Michael J. Carey 0001, Dmitry Lychagin, Till Westmann, Daniel Engovatov, Nicola Onose |
VLDB | 1 |
| 2004 | Liquid Data for WebLogic: Integrating Enterprise Data and ServicesabstractInformation in today's enterprises commonly resides in a variety of heterogeneous data sources, including relational databases, web services, files, packaged applications, and custom data repositories. BEA's enterprise information integration product, Liquid Data for WebLogic, takes an XML-based approach to providing integrated access to such heterogeneous information. This demonstration highlights the XML technologies involved - including web services, XQuery, and XML Schema - and shows how they can be brought to bear on the enterprise information integration problem. The demonstration uses a simple end-to-end example, one that involves integrating data from relational databases and web services, to walk the audience through the overall architecture, XML-based data modeling approach, programming model, declarative query and view facilities, and distributed processing features of Liquid Data. Vinayak R. Borkar |
SIGMOD Conference | 1 |
| 2003 | XML queries and algebra in the Enosys integration platform
Yannis Papakonstantinou, Vinayak R. Borkar, Maxim Orgiyan, Konstantinos Stathatos, Lucian Suta, Vasilis Vassalos, Pavel E. Velikhov |
Data Knowl. Eng. | 2 |
| 2001 | Automatic Segmentation of Text into Structured RecordsabstractIn this paper we present a method for automatically segmenting unformatted text records into structured elements. Several useful data sources today are human-generated as continuous text whereas convenient usage requires the data to be organized as structured records. A prime motivation is the warehouse address cleaning problem of transforming dirty addresses stored in large corporate databases as a single text field into subfields like “City” and “Street”. Existing tools rely on hand-tuned, domain-specific rule-based systems. Vinayak R. Borkar, Kaustubh Deshmukh, Sunita Sarawagi |
SIGMOD Conference | 1 |