Mayank Bawa

dblp:66/1099 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
0since 2021 · last 2009
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-authorSystems, architecture and hardware · 1Theory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
8 papers
Information retrieval · 47% Indexing and storage engines · 16% Distributed and cloud data management · 13%
Network and information security
3 papers
Privacy and data protection · 87% Network security · 13%
Theoretical computer science
1 paper
Distributed computing theory · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%

Topics — the 20 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › distributed information retrieval
peer-to-peer search
0.132009
Make it fresh, make it quick: searching a network of personal webservers · WWW 2003
SETS: search enhanced by topic segmentation · SIGIR 2003
Privacy-preserving indexing of documents on the network · VLDB J. 2009
Privacy and data protection
privacy-preserving search
0.112009
Privacy-preserving indexing of documents on the network · VLDB J. 2009
Information retrieval › distributed information retrieval
distributed search
0.132009
Make it fresh, make it quick: searching a network of personal webservers · WWW 2003
Privacy-preserving indexing of documents on the network · VLDB J. 2009
Privacy-Preserving Indexing of Documents on the Network · VLDB 2003
Indexing and storage engines › multidimensional indexing
high-dimensional indexing
0.112005
LSH forest: self-tuning indexes for similarity search · WWW 2005
Information retrieval › hashing › hashing for nearest neighbor search
locality-sensitive hashing
0.112005
LSH forest: self-tuning indexes for similarity search · WWW 2005
Database system architecture and tuning › index tuning
self-tuning index
0.112005
LSH forest: self-tuning indexes for similarity search · WWW 2005
Information retrieval
similarity search
0.112005
LSH forest: self-tuning indexes for similarity search · WWW 2005
Indexing and storage engines
vector index
0.112005
LSH forest: self-tuning indexes for similarity search · WWW 2005
Query processing and optimization
aggregate query processing
0.012004
The Price of Validity in Dynamic Networks · SIGMOD Conference 2004
Distributed and cloud data management
distributed query processing
0.012004
The Price of Validity in Dynamic Networks · SIGMOD Conference 2004
Privacy and data protection
privacy-preserving data analysis
0.012004
Vision Paper: Enabling Privacy for the Paranoids · VLDB 2004
Distributed computing theory
dynamic networks
0.012004
The Price of Validity in Dynamic Networks · SIGMOD Conference 2004
Information retrieval › text analysis › text segmentation
topic segmentation
0.012003
SETS: search enhanced by topic segmentation · SIGIR 2003
Privacy and data protection › privacy-preserving search
privacy-preserving index
0.012003
Privacy-Preserving Indexing of Documents on the Network · VLDB 2003
Network security
anonymity networks
0.012009
Privacy-preserving indexing of documents on the network · VLDB J. 2009
Data mining › pattern mining
association rule mining
0.012000
Turbo-charging Vertical Mining of Large Databases · SIGMOD Conference 2000
Data mining
pattern mining
0.012000
Turbo-charging Vertical Mining of Large Databases · SIGMOD Conference 2000
Data mining › pattern mining
vertical mining
0.012000
Turbo-charging Vertical Mining of Large Databases · SIGMOD Conference 2000
Privacy and data protection
anonymization
0.012004
Vision Paper: Enabling Privacy for the Paranoids · VLDB 2004
Information retrieval
distributed information retrieval
0.012003
SETS: search enhanced by topic segmentation · SIGIR 2003

Methods — techniques the papers use, named apart from their topics

cryptographic indexing · 0.2distributed indexing · 0.1centralized network state summary · 0.1locality-sensitive hashing · 0.1approximate similarity search · 0.1social network theory · 0.0machine learning · 0.0snake intersection · 0.0bit-vector compression · 0.0
YearPublicationVenuePosition
2009 Privacy-preserving indexing of documents on the network
Mayank Bawa, Roberto J. Bayardo, Rakesh Agrawal 0001, Jaideep Vaidya
VLDB J.1
2007 The price of validity in dynamic networks
Mayank Bawa, Aristides Gionis, Hector Garcia-Molina, Rajeev Motwani 0001
J. Comput. Syst. Sci.1
2005 Two Can Keep A Secret: A Distributed Architecture for Secure Database Services
Gagan Aggarwal, Mayank Bawa, Prasanna Ganesan, Hector Garcia-Molina, Krishnaram Kenthapadi, Rajeev Motwani 0001, Utkarsh Srivastava, Dilys Thomas, Ying Xu 0002
CIDR2
2005 LSH forest: self-tuning indexes for similarity search
abstract
We consider the problem of indexing high-dimensional data for answering (approximate) similarity-search queries. Similarity indexes prove to be important in a wide variety of settings: Web search engines desire fast, parallel, main-memory-based indexes for similarity search on text data; database systems desire disk-based similarity indexes for high-dimensional data, including text and images; peer-to-peer systems desire distributed similarity indexes with low communication cost. We propose an indexing scheme called LSH Forest which is applicable in all the above contexts. Our index uses the well-known technique of locality-sensitive hashing (LSH), but improves upon previous designs by (a) eliminating the different data-dependent parameters for which LSH must be constantly hand-tuned, and (b) improving on LSH's performance guarantees for skewed data distributions while retaining the same storage and query overhead. We show how to construct this index in main memory, on disk, in parallel systems, and in peer-to-peer systems. We evaluate the design with experiments on multiple text corpora and demonstrate both the self-tuning nature and the superior performance of LSH Forest.
Mayank Bawa, Tyson Condie, Prasanna Ganesan
WWW1
2005 Authenticity and availability in PIPE networks
Brian F. Cooper, Mayank Bawa, Neil Daswani, Sergio Marti, Hector Garcia-Molina
Future Gener. Comput. Syst.2
2004 The Price of Validity in Dynamic Networks
abstract
Massive-scale self-administered networks like Peer-to-Peer and Sensor Networks have data distributed across thousands of participant hosts. These networks are highly dynamic with short-lived hosts being the norm rather than an exception. In recent years, researchers have investigated best-effort algorithms to efficiently process aggregate queries (e.g., sum, count, average, minimum and maximum) [6, 13, 21, 34, 35, 37] on these networks. Unfortunately, query semantics for best-effort algorithms are ill-defined, making it hard to reason about guarantees associated with the result returned. In this paper, we specify a correctness condition, single-site validity, with respect to which the above algorithms are best-effort. We present a class of algorithms that guarantee validity in dynamic networks. Experiments on real-life and synthetic network topologies validate performance of our algorithms, revealing the hitherto unknown price of validity.
Mayank Bawa, Aristides Gionis, Hector Garcia-Molina, Rajeev Motwani 0001
SIGMOD Conference1
2004 Vision Paper: Enabling Privacy for the Paranoids
Gagan Aggarwal, Mayank Bawa, Prasanna Ganesan, Hector Garcia-Molina, Krishnaram Kenthapadi, Nina Mishra, Rajeev Motwani 0001, Utkarsh Srivastava, Dilys Thomas, Jennifer Widom, Ying Xu 0002
VLDB2
2004 Online Balancing of Range-Partitioned Data with Applications to Peer-to-Peer Systems
Prasanna Ganesan, Mayank Bawa, Hector Garcia-Molina
VLDB2
2003 SETS: search enhanced by topic segmentation
abstract
We present SETS, an architecture for efficient search in peer-to-peer networks, building upon ideas drawn from machine learning and social network theory. The key idea is to arrange participating sites in a topic-segmented overlay topology in which most connections are short-distance, connecting pairs of sites with similar content. Topically focused sets of sites are then joined together into a single network by long-distance links. Queries are matched and routed to only the topically closest regions. We discuss a variety of design issues and tradeoffs that an implementor of SETS would face. We show that SETS is efficient in network traffic and query processing load.
Mayank Bawa, Gurmeet Singh Manku, Prabhakar Raghavan
SIGIR1
2003 Privacy-Preserving Indexing of Documents on the Network
Mayank Bawa, Roberto J. Bayardo, Rakesh Agrawal 0001
VLDB1
2003 Make it fresh, make it quick: searching a network of personal webservers
abstract
Personal webservers have proven to be a popular means of sharing files and peer collaboration. Unfortunately, the transient availability and rapidly evolving content on such hosts render centralized, crawl-based search indices stale and incomplete. To address this problem, we propose YouSearch, a distributed search application for personal webservers operating within a shared context (e.g., a corporate intranet). With YouSearch, search results are always fast, fresh and complete -- properties we show arise from an architecture that exploits both the extensive distributed resources available at the peer webservers in addition to a centralized repository of summarized network state. YouSearch extends the concept of a shared context within web communities by enabling peers to aggregate into groups and users to search over specific groups. In this paper, we describe the challenges, design, implementation and experiences with a successful intranet deployment of YouSearch.
Mayank Bawa, Roberto J. Bayardo, Sridhar Rajagopalan, Eugene J. Shekita
WWW1
2001 Minimizing View Sets without Losing Query-Answering Power
Chen Li 0001, Mayank Bawa, Jeffrey D. Ullman
ICDT2
2000 Turbo-charging Vertical Mining of Large Databases
abstract
In a vertical representation of a market-basket database, each item is associated with a column of values representing the transactions in which it is present. The association-rule mining algorithms that have been recently proposed for this representation show performance improvements over their classical horizontal counterparts, but are either efficient only for certain database sizes, or assume particular characteristics of the database contents, or are applicable only to specific kinds of database schemas. We present here a new vertical mining algorithm called VIPER, which is general-purpose, making no special requirements of the underlying database. VIPER stores data in compressed bit-vectors called “snakes” and integrates a number of novel optimizations for efficient snake generation, intersection, counting and storage. We analyze the performance of VIPER for a range of synthetic database workloads. Our experimental results indicate significant performance gains, especially for large databases, over previously proposed vertical and horizontal mining algorithms. In fact, there are even workload regions where VIPER outperforms an optimal, but practically infeasible, horizontal mining algorithm.
Pradeep Shenoy, Jayant R. Haritsa, S. Sudarshan 0001, Gaurav Bhalotia, Mayank Bawa, Devavrat Shah
SIGMOD Conference5