EDBT 2026 Demo / reviewers in the wild / expert
Gaurav Bhalotia
dblp:14/3775
· DBLP profile ↗
4ranked-venue papers
1as first author
0since 2021 · last 2004
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 51% Data mining · 29% Query processing and optimization · 17% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › ranking
graph-based ranking |
0.0 | 1 | 2002 | Keyword Searching and Browsing in Databases using BANKS · ICDE 2002 |
Information retrieval
keyword search |
0.0 | 1 | 2002 | BANKS: Browsing and Keyword Searching in Relational Databases · VLDB 2002 |
Query processing and optimization › keyword query processing
keyword search over databases |
0.0 | 1 | 2002 | Keyword Searching and Browsing in Databases using BANKS · ICDE 2002 |
Information retrieval
ranking |
0.0 | 1 | 2002 | Keyword Searching and Browsing in Databases using BANKS · ICDE 2002 |
Information retrieval
retrieval models |
0.0 | 1 | 2002 | Keyword Searching and Browsing in Databases using BANKS · ICDE 2002 |
Data mining › pattern mining
association rule mining |
0.0 | 1 | 2000 | Turbo-charging Vertical Mining of Large Databases · SIGMOD Conference 2000 |
Data mining
pattern mining |
0.0 | 1 | 2000 | Turbo-charging Vertical Mining of Large Databases · SIGMOD Conference 2000 |
Data mining › pattern mining
vertical mining |
0.0 | 1 | 2000 | Turbo-charging Vertical Mining of Large Databases · SIGMOD Conference 2000 |
Query processing and optimization › query execution
relational query processing |
0.0 | 1 | 2002 | BANKS: Browsing and Keyword Searching in Relational Databases · VLDB 2002 |
Methods — techniques the papers use, named apart from their topics
snake intersection · 0.0bit-vector compression · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2004 | Tools for loading MEDLINE into a local relational databaseabstractBACKGROUND: Researchers who use MEDLINE for text mining, information extraction, or natural language processing may benefit from having a copy of MEDLINE that they can manage locally. The National Library of Medicine (NLM) distributes MEDLINE in eXtensible Markup Language (XML)-formatted text files, but it is difficult to query MEDLINE in that format. We have developed software tools to parse the MEDLINE data files and load their contents into a relational database. Although the task is conceptually straightforward, the size and scope of MEDLINE make the task nontrivial. Given the increasing importance of text analysis in biology and medicine, we believe a local installation of MEDLINE will provide helpful computing infrastructure for researchers. RESULTS: We developed three software packages that parse and load MEDLINE, and ran each package to install separate instances of the MEDLINE database. For each installation, we collected data on loading time and disk-space utilization to provide examples of the process in different settings. Settings differed in terms of commercial database-management system (IBM DB2 or Oracle 9i), processor (Intel or Sun), programming language of installation software (Java or Perl), and methods employed in different versions of the software. The loading times for the three installations were 76 hours, 196 hours, and 132 hours, and disk-space utilization was 46.3 GB, 37.7 GB, and 31.6 GB, respectively. Loading times varied due to a variety of differences among the systems. Loading time also depended on whether data were written to intermediate files or not, and on whether input files were processed in sequence or in parallel. Disk-space utilization depended on the number of MEDLINE files processed, amount of indexing, and whether abstracts were stored as character large objects or truncated. CONCLUSIONS: Relational database (RDBMS) technology supports indexing and querying of very large datasets, and can accommodate a locally stored version of MEDLINE. RDBMS systems support a wide range of queries and facilitate certain tasks that are not directly supported by the application programming interface to PubMed. Because there is variation in hardware, software, and network infrastructures across sites, we cannot predict the exact time required for a user to load MEDLINE, but our results suggest that performance of the software is reasonable. Our database schemas and conversion software are publicly available at http://biotext.berkeley.edu. Diane E. Oliver, Gaurav Bhalotia, Ariel S. Schwartz, Russ B. Altman, Marti A. Hearst |
BMC Bioinform. | 2 |
| 2002 | Keyword Searching and Browsing in Databases using BANKSabstractWith the growth of the Web, there has been a rapid increase in the number of users who need to access online databases without having a detailed knowledge of the schema or of query languages; even relatively simple query languages designed for non-experts are too complicated for them. We describe BANKS, a system which enables keyword-based search on relational databases, together with data and schema browsing. BANKS enables users to extract information in a simple manner without any knowledge of the schema or any need for writing complex queries. A user can get information by typing a few keywords, following hyperlinks, and interacting with controls on the displayed results. BANKS models tuples as nodes in a graph, connected by links induced by foreign key and other relationships. Answers to a query are modeled as rooted trees connecting tuples that match individual keywords in the query. Answers are ranked using a notion of proximity coupled with a notion of prestige of nodes based on inlinks, similar to techniques developed for Web search. We present an efficient heuristic algorithm for finding and ranking query results. Gaurav Bhalotia, Arvind Hulgeri, Charuta Nakhe, Soumen Chakrabarti, S. Sudarshan 0001 |
ICDE | 1 |
| 2002 | BANKS: Browsing and Keyword Searching in Relational Databases
B. Aditya, Gaurav Bhalotia, Soumen Chakrabarti, Arvind Hulgeri, Charuta Nakhe, Parag, S. Sudarshan 0001 |
VLDB | 2 |
| 2000 | Turbo-charging Vertical Mining of Large DatabasesabstractIn a vertical representation of a market-basket database, each item is associated with a column of values representing the transactions in which it is present. The association-rule mining algorithms that have been recently proposed for this representation show performance improvements over their classical horizontal counterparts, but are either efficient only for certain database sizes, or assume particular characteristics of the database contents, or are applicable only to specific kinds of database schemas. We present here a new vertical mining algorithm called VIPER, which is general-purpose, making no special requirements of the underlying database. VIPER stores data in compressed bit-vectors called “snakes” and integrates a number of novel optimizations for efficient snake generation, intersection, counting and storage. We analyze the performance of VIPER for a range of synthetic database workloads. Our experimental results indicate significant performance gains, especially for large databases, over previously proposed vertical and horizontal mining algorithms. In fact, there are even workload regions where VIPER outperforms an optimal, but practically infeasible, horizontal mining algorithm. Pradeep Shenoy, Jayant R. Haritsa, S. Sudarshan 0001, Gaurav Bhalotia, Mayank Bawa, Devavrat Shah |
SIGMOD Conference | 4 |