Sanjay Agrawal 0001

dblp:a/SanjayAgrawal · DBLP profile ↗
← Back
14ranked-venue papers
12as first author
1since 2021 · last 2021
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 13 · 12 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
11 papers
Information retrieval · 49% Database system architecture and tuning · 39% Query processing and optimization · 10%
Artificial intelligence
2 papers
Graph learning · 46% Representation and self-supervised learning · 46% Information extraction and text analysis · 8%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph neural network
0.512021
GraphFormers: GNN-nested Transformers for Representation Learning on Textual Graph · NeurIPS 2021
Machine learning › Representation and self-supervised learning › text embedding
text representation learning
0.512021
GraphFormers: GNN-nested Transformers for Representation Learning on Textual Graph · NeurIPS 2021
Information retrieval › distributed information retrieval
federated search
0.222010
Query portals: dynamically generating portals for entity-oriented web queries · SIGMOD Conference 2010
Exploiting web search engines to search structured databases · WWW 2009
Database system architecture and tuning › database design
physical database design
0.242006
Automatic physical design tuning: workload as a sequence · SIGMOD Conference 2006
Database tuning advisor for microsoft SQL server 2005: demo · SIGMOD Conference 2005
Materialized View and Index Selection Tool for Microsoft SQL Server 2000 · SIGMOD Conference 2001
Information retrieval
search engines
0.112010
Query portals: dynamically generating portals for entity-oriented web queries · SIGMOD Conference 2010
Database system architecture and tuning › database design › physical database design
index selection
0.132005
Database tuning advisor for microsoft SQL server 2005: demo · SIGMOD Conference 2005
Materialized View and Index Selection Tool for Microsoft SQL Server 2000 · SIGMOD Conference 2001
Database Tuning Advisor for Microsoft SQL Server 2005 · VLDB 2004
Information retrieval › distributed information retrieval
integrated search
0.112009
Exploiting web search engines to search structured databases · WWW 2009
Information retrieval › search engines
structured data search
0.112009
Exploiting web search engines to search structured databases · WWW 2009
Query processing and optimization › materialized view
materialized view selection
0.132005
Database tuning advisor for microsoft SQL server 2005: demo · SIGMOD Conference 2005
Materialized View and Index Selection Tool for Microsoft SQL Server 2000 · SIGMOD Conference 2001
Automated Selection of Materialized Views and Indexes in SQL Databases · VLDB 2000
Natural language and speech › Information extraction and text analysis
named entity recognition
0.112008
Scalable ad-hoc entity extraction from text collections · Proc. VLDB Endow. 2008
Information retrieval › indexing
inverted index
0.112008
Scalable ad-hoc entity extraction from text collections · Proc. VLDB Endow. 2008
Information retrieval
keyword search
0.122002
DBXplorer: enabling keyword search over relational databases · SIGMOD Conference 2002
DBXplorer: A System for Keyword-Based Search over Relational Databases · ICDE 2002
Information retrieval › keyword search
keyword search over relational databases
0.122002
DBXplorer: enabling keyword search over relational databases · SIGMOD Conference 2002
DBXplorer: A System for Keyword-Based Search over Relational Databases · ICDE 2002
Database system architecture and tuning › database tuning
automatic database tuning
0.112006
Automatic physical design tuning: workload as a sequence · SIGMOD Conference 2006
Database system architecture and tuning › database design › physical database design
automated physical design
0.012004
Integrating Vertical and Horizontal Partitioning Into Automated Physical Database Design · SIGMOD Conference 2004
Query processing and optimization
keyword query processing
0.022002
DBXplorer: enabling keyword search over relational databases · SIGMOD Conference 2002
DBXplorer: A System for Keyword-Based Search over Relational Databases · ICDE 2002
Data integration and cleaning
entity resolution
0.012010
Query portals: dynamically generating portals for entity-oriented web queries · SIGMOD Conference 2010
Database system architecture and tuning › database design › physical database design
index and materialized view selection
0.012001
Materialized View and Index Selection Tool for Microsoft SQL Server 2000 · SIGMOD Conference 2001
Database system architecture and tuning
horizontal partitioning
0.022005
Database tuning advisor for microsoft SQL server 2005: demo · SIGMOD Conference 2005
Integrating Vertical and Horizontal Partitioning Into Automated Physical Database Design · SIGMOD Conference 2004
Database system architecture and tuning › database design › physical database design
vertical partitioning
0.012004
Integrating Vertical and Horizontal Partitioning Into Automated Physical Database Design · SIGMOD Conference 2004

Methods — techniques the papers use, named apart from their topics

transformer · 0.5progressive learning · 0.5graph neural network · 0.5inverted index lookup · 0.2workload analysis · 0.1query-time processing · 0.1preprocessing · 0.1query federation · 0.1experimental evaluation · 0.0
YearPublicationVenuePosition
2021 GraphFormers: GNN-nested Transformers for Representation Learning on Textual Graph
abstract
The representation learning on textual graph is to generate low-dimensional embeddings for the nodes based on the individual textual features and the neighbourhood information. Recent breakthroughs on pretrained language models and graph neural networks push forward the development of corresponding techniques. The existing works mainly rely on the cascaded model architecture: the textual features of nodes are independently encoded by language models at first; the textual embeddings are aggregated by graph neural networks afterwards. However, the above architecture is limited due to the independent modeling of textual features. In this work, we propose GraphFormers, where layerwise GNN components are nested alongside the transformer blocks of language models. With the proposed architecture, the text encoding and the graph aggregation are fused into an iterative workflow, making each node's semantic accurately comprehended from the global perspective. In addition, a progressive learning strategy is introduced, where the model is successively trained on manipulated data and original data to reinforce its capability of integrating information on graph. Extensive evaluations are conducted on three large-scale benchmark datasets, where GraphFormers outperform the SOTA baselines with comparable running efficiency. The source code is released at https://github.com/microsoft/GraphFormers .
Junhan Yang, Zheng Liu 0011, Shitao Xiao, Chaozhuo Li, Defu Lian, Sanjay Agrawal 0001, Guangzhong Sun, Xing Xie 0001
NeurIPS6
2010 Query portals: dynamically generating portals for entity-oriented web queries
abstract
Many web queries seek information about named entities (such as products or people). Web search engines federate such entity-oriented queries to relevant structured databases; the results of those searches are then returned to the user along with web search results. Current federated approaches have two limitations: (i) they often fail to return important results for a broad class of such entity-oriented queries and (ii) the information they return per entity is often inadequate. In this paper, we present the Query Portals system that addresses these limitations. The Query Portals system dynamically generates a portal for an entity-oriented query. It first provides an overview of the relevant entities and further allows users to drill down to gather additional information on these entities. Our architecture uses a judicious combination of pre-processing and query time techniques so that the query portal can be generated efficiently.
Sanjay Agrawal 0001, Kaushik Chakrabarti, Surajit Chaudhuri, Venkatesh Ganti, Arnd Christian König, Dong Xin
SIGMOD Conference1
2009 Exploiting web search engines to search structured databases
abstract
Web search engines often federate many user queries to relevant structured databases. For example, a product related query might be federated to a product database containing their descriptions and specifications. The relevant structured data items are then returned to the user along with web search results. However, each structured database is searched in isolation. Hence, the search often produces empty or incomplete results as the database may not contain the required information to answer the query. In this paper, we propose a novel integrated search architecture. We establish and exploit the relationships between web search results and the items in structured databases to identify the relevant structured data items for a much wider range of queries.Our architecture leverages existing search engine components to implement this functionality at very low overhead. We demonstrate the quality and efficiency of our techniques through an extensive experimental study.
Sanjay Agrawal 0001, Kaushik Chakrabarti, Surajit Chaudhuri, Venkatesh Ganti, Arnd Christian König, Dong Xin
WWW1
2008 Scalable ad-hoc entity extraction from text collections
abstract
Supporting entity extraction from large document collections is important for enabling a variety of important data analysis tasks. In this paper, we introduce the "ad-hoc" entity extraction task where entities of interest are constrained to be from a list of entities that is specific to the task. In such scenarios, traditional entity extraction techniques that process all the documents for each ad-hoc entity extraction task can be significantly expensive. In this paper, we propose an efficient approach that leverages the inverted index on the documents to identify the subset of documents relevant to the task and processes only those documents. We demonstrate the efficiency of our techniques on real datasets.
Sanjay Agrawal 0001, Kaushik Chakrabarti, Surajit Chaudhuri, Venkatesh Ganti
Proc. VLDB Endow.1
2006 Automatic physical design tuning: workload as a sequence
abstract
The area of automatic selection of physical database design to optimize the performance of a relational database system based on a workload of SQL queries and updates has gained prominence in recent years. Major database vendors have released automated physical database design tools with the goal of reducing the total cost of ownership. An important assumption underlying these tools is that the workload is a set of SQL statements. In this paper, we show that being able to treat the workload as a sequence, i.e., exploiting the ordering of statements can significantly broaden the usage of such tools. We present scenarios where exploiting sequence information in the workload is crucial for performance tuning. We also propose techniques for addressing the technical challenges arising from treating the workload as a sequence. We evaluate the effectiveness of our techniques through experiments on Microsoft SQL Server.
Sanjay Agrawal 0001, Eric Chu, Vivek R. Narasayya
SIGMOD Conference1
2005 Database tuning advisor for microsoft SQL server 2005: demo
abstract
Database Tuning Advisor (DTA) is a physical database design tool that is part of Microsoft's SQL Server 2005 relational database management system. Previously known as "Index Tuning Wizard" in SQL Server 7.0 and SQL Server 2000, DTA adds new functionality that is not available in other contemporary physical design tuning tools. Novel aspects of DTA that will be demonstrated include: (a) Ability to take into account both performance and manageability requirements of DBAs (b) Fully integrated recommendations for indexes, materialized views and horizontal partitioning (c) Transparently leverage a test server to offload tuning load from production server and (d) Easy programmability and scriptability.
Sanjay Agrawal 0001, Surajit Chaudhuri, Lubor Kollár, Arunprasad P. Marathe, Vivek R. Narasayya, Manoj Syamala
SIGMOD Conference1
2004 Integrating Vertical and Horizontal Partitioning Into Automated Physical Database Design
abstract
In addition to indexes and materialized views, horizontal and vertical partitioning are important aspects of physical design in a relational database system that significantly impact performance. Horizontal partitioning also provides manageability; database administrators often require indexes and their underlying tables partitioned identically so as to make common operations such as backup/restore easier. While partitioning is important, incorporating partitioning makes the problem of automating physical design much harder since: (a) The choices of partitioning can strongly interact with choices of indexes and materialized views. (b) A large new space of physical design alternatives must be considered. (c) Manageability requirements impose a new constraint on the problem. In this paper, we present novel techniques for designing a scalable solution to this integrated physical design problem that takes both performance and manageability into account. We have implemented our techniques and evaluated it on Microsoft SQL Server. Our experiments highlight: (a) the importance of taking an integrated approach to automated physical design and (b) the scalability of our techniques. 1.
Sanjay Agrawal 0001, Vivek R. Narasayya, Beverly Yang
SIGMOD Conference1
2004 Database Tuning Advisor for Microsoft SQL Server 2005
Sanjay Agrawal 0001, Surajit Chaudhuri, Lubor Kollár, Arunprasad P. Marathe, Vivek R. Narasayya, Manoj Syamala
VLDB1
2003 Automated Ranking of Database Query Results
Sanjay Agrawal 0001, Surajit Chaudhuri, Gautam Das 0001, Aristides Gionis
CIDR1
2002 DBXplorer: A System for Keyword-Based Search over Relational Databases
abstract
Internet search engines have popularized the keyword-based search paradigm. While traditional database management systems offer powerful query languages, they do not allow keyword-based search. In this paper, we discuss DBXplorer, a system that enables keyword-based searches in relational databases. DBXplorer has been implemented using a commercial relational database and Web server and allows users to interact via a browser front-end. We outline the challenges and discuss the implementation of our system, including results of extensive experimental evaluation.
Sanjay Agrawal 0001, Surajit Chaudhuri, Gautam Das 0001
ICDE1
2002 DBXplorer: enabling keyword search over relational databases
abstract
No abstract available.
Sanjay Agrawal 0001, Surajit Chaudhuri, Gautam Das 0001
SIGMOD Conference1
2001 Materialized View and Index Selection Tool for Microsoft SQL Server 2000
abstract
No abstract available.
Sanjay Agrawal 0001, Surajit Chaudhuri, Vivek R. Narasayya
SIGMOD Conference1
2000 Persistent Client-Server Database Sessions
Roger S. Barga, David B. Lomet, Thomas Baby, Sanjay Agrawal 0001
EDBT4
2000 Automated Selection of Materialized Views and Indexes in SQL Databases
Sanjay Agrawal 0001, Surajit Chaudhuri, Vivek R. Narasayya
VLDB1