Gaurav Bhalotia

dblp:14/3775 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
0since 2021 · last 2004
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Information retrieval · 51% Data mining · 29% Query processing and optimization · 17%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › ranking
graph-based ranking
0.012002
Keyword Searching and Browsing in Databases using BANKS · ICDE 2002
Information retrieval
keyword search
0.012002
BANKS: Browsing and Keyword Searching in Relational Databases · VLDB 2002
Query processing and optimization › keyword query processing
keyword search over databases
0.012002
Keyword Searching and Browsing in Databases using BANKS · ICDE 2002
Information retrieval
ranking
0.012002
Keyword Searching and Browsing in Databases using BANKS · ICDE 2002
Information retrieval
retrieval models
0.012002
Keyword Searching and Browsing in Databases using BANKS · ICDE 2002
Data mining › pattern mining
association rule mining
0.012000
Turbo-charging Vertical Mining of Large Databases · SIGMOD Conference 2000
Data mining
pattern mining
0.012000
Turbo-charging Vertical Mining of Large Databases · SIGMOD Conference 2000
Data mining › pattern mining
vertical mining
0.012000
Turbo-charging Vertical Mining of Large Databases · SIGMOD Conference 2000
Query processing and optimization › query execution
relational query processing
0.012002
BANKS: Browsing and Keyword Searching in Relational Databases · VLDB 2002

Methods — techniques the papers use, named apart from their topics

snake intersection · 0.0bit-vector compression · 0.0
YearPublicationVenuePosition
2004 Tools for loading MEDLINE into a local relational database
abstract
BACKGROUND: Researchers who use MEDLINE for text mining, information extraction, or natural language processing may benefit from having a copy of MEDLINE that they can manage locally. The National Library of Medicine (NLM) distributes MEDLINE in eXtensible Markup Language (XML)-formatted text files, but it is difficult to query MEDLINE in that format. We have developed software tools to parse the MEDLINE data files and load their contents into a relational database. Although the task is conceptually straightforward, the size and scope of MEDLINE make the task nontrivial. Given the increasing importance of text analysis in biology and medicine, we believe a local installation of MEDLINE will provide helpful computing infrastructure for researchers. RESULTS: We developed three software packages that parse and load MEDLINE, and ran each package to install separate instances of the MEDLINE database. For each installation, we collected data on loading time and disk-space utilization to provide examples of the process in different settings. Settings differed in terms of commercial database-management system (IBM DB2 or Oracle 9i), processor (Intel or Sun), programming language of installation software (Java or Perl), and methods employed in different versions of the software. The loading times for the three installations were 76 hours, 196 hours, and 132 hours, and disk-space utilization was 46.3 GB, 37.7 GB, and 31.6 GB, respectively. Loading times varied due to a variety of differences among the systems. Loading time also depended on whether data were written to intermediate files or not, and on whether input files were processed in sequence or in parallel. Disk-space utilization depended on the number of MEDLINE files processed, amount of indexing, and whether abstracts were stored as character large objects or truncated. CONCLUSIONS: Relational database (RDBMS) technology supports indexing and querying of very large datasets, and can accommodate a locally stored version of MEDLINE. RDBMS systems support a wide range of queries and facilitate certain tasks that are not directly supported by the application programming interface to PubMed. Because there is variation in hardware, software, and network infrastructures across sites, we cannot predict the exact time required for a user to load MEDLINE, but our results suggest that performance of the software is reasonable. Our database schemas and conversion software are publicly available at http://biotext.berkeley.edu.
Diane E. Oliver, Gaurav Bhalotia, Ariel S. Schwartz, Russ B. Altman, Marti A. Hearst
BMC Bioinform.2
2002 Keyword Searching and Browsing in Databases using BANKS
abstract
With the growth of the Web, there has been a rapid increase in the number of users who need to access online databases without having a detailed knowledge of the schema or of query languages; even relatively simple query languages designed for non-experts are too complicated for them. We describe BANKS, a system which enables keyword-based search on relational databases, together with data and schema browsing. BANKS enables users to extract information in a simple manner without any knowledge of the schema or any need for writing complex queries. A user can get information by typing a few keywords, following hyperlinks, and interacting with controls on the displayed results. BANKS models tuples as nodes in a graph, connected by links induced by foreign key and other relationships. Answers to a query are modeled as rooted trees connecting tuples that match individual keywords in the query. Answers are ranked using a notion of proximity coupled with a notion of prestige of nodes based on inlinks, similar to techniques developed for Web search. We present an efficient heuristic algorithm for finding and ranking query results.
Gaurav Bhalotia, Arvind Hulgeri, Charuta Nakhe, Soumen Chakrabarti, S. Sudarshan 0001
ICDE1
2002 BANKS: Browsing and Keyword Searching in Relational Databases
B. Aditya, Gaurav Bhalotia, Soumen Chakrabarti, Arvind Hulgeri, Charuta Nakhe, Parag, S. Sudarshan 0001
VLDB2
2000 Turbo-charging Vertical Mining of Large Databases
abstract
In a vertical representation of a market-basket database, each item is associated with a column of values representing the transactions in which it is present. The association-rule mining algorithms that have been recently proposed for this representation show performance improvements over their classical horizontal counterparts, but are either efficient only for certain database sizes, or assume particular characteristics of the database contents, or are applicable only to specific kinds of database schemas. We present here a new vertical mining algorithm called VIPER, which is general-purpose, making no special requirements of the underlying database. VIPER stores data in compressed bit-vectors called “snakes” and integrates a number of novel optimizations for efficient snake generation, intersection, counting and storage. We analyze the performance of VIPER for a range of synthetic database workloads. Our experimental results indicate significant performance gains, especially for large databases, over previously proposed vertical and horizontal mining algorithms. In fact, there are even workload regions where VIPER outperforms an optimal, but practically infeasible, horizontal mining algorithm.
Pradeep Shenoy, Jayant R. Haritsa, S. Sudarshan 0001, Gaurav Bhalotia, Mayank Bawa, Devavrat Shah
SIGMOD Conference4