Bhuvan Bamba

dblp:26/6942 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 13 · 4 first-authorArtificial intelligence and machine learning · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 first-authorSystems, architecture and hardware · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorComputer networks · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Spatial and temporal data management · 36% Information retrieval · 28% Query processing and optimization · 16%
Network and information security
2 papers
Privacy and data protection · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 62% Storage systems · 19% Distributed systems · 19%

Topics — the 20 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › data reduction › data summarization
histogram construction
0.212013
Statistics Collection in Oracle Spatial and Graph: Fast Histogram Construction for Complex Geometry Objects · Proc. VLDB Endow. 2013
Query processing and optimization
selectivity estimation
0.212013
Statistics Collection in Oracle Spatial and Graph: Fast Histogram Construction for Complex Geometry Objects · Proc. VLDB Endow. 2013
Spatial and temporal data management
spatial databases
0.212013
Statistics Collection in Oracle Spatial and Graph: Fast Histogram Construction for Complex Geometry Objects · Proc. VLDB Endow. 2013
Information retrieval
ranking
0.122007
DSphere: A Source-Centric Approach to Crawling, Indexing and Searching the World Wide Web · ICDE 2007
Information Retrieval and Knowledge Discovery Utilizing a BioMedical Patent Semantic Web · IEEE Trans. Knowl. Data Eng. 2005
Spatial and temporal data management
location-based services
0.112010
RoadTrack: Scaling Location Updates for Mobile Clients on Road Networks with Query Awareness · Proc. VLDB Endow. 2010
Spatial and temporal data management › moving object databases
location update
0.112010
RoadTrack: Scaling Location Updates for Mobile Clients on Road Networks with Query Awareness · Proc. VLDB Endow. 2010
Privacy and data protection › location privacy
proximity privacy
0.112009
A General Proximity Privacy Principle · ICDE 2009
Privacy and data protection › anonymization
location anonymity
0.112008
Supporting anonymous location queries in mobile environments with privacygrid · WWW 2008
Privacy and data protection › location privacy
location cloaking
0.112008
Supporting anonymous location queries in mobile environments with privacygrid · WWW 2008
Privacy and data protection
location privacy
0.112008
Supporting anonymous location queries in mobile environments with privacygrid · WWW 2008
Information retrieval › web search
link analysis
0.112007
DSphere: A Source-Centric Approach to Crawling, Indexing and Searching the World Wide Web · ICDE 2007
Information retrieval
web search
0.112007
DSphere: A Source-Centric Approach to Crawling, Indexing and Searching the World Wide Web · ICDE 2007
Cloud and datacenter computing › resource management
datacenter resource management
0.112007
Integrated resource allocation in heterogeneous SAN data centers · PODC 2007
Data integration and cleaning › semantic integration
ontology-based data integration
0.112005
Information Retrieval and Knowledge Discovery Utilizing a BioMedical Patent Semantic Web · IEEE Trans. Knowl. Data Eng. 2005
Information retrieval › document retrieval › domain-specific retrieval › legal information retrieval
patent retrieval
0.112005
Information Retrieval and Knowledge Discovery Utilizing a BioMedical Patent Semantic Web · IEEE Trans. Knowl. Data Eng. 2005
Knowledge graphs
semantic web
0.112005
Information Retrieval and Knowledge Discovery Utilizing a BioMedical Patent Semantic Web · IEEE Trans. Knowl. Data Eng. 2005
Query processing and optimization
query optimization
0.012013
Statistics Collection in Oracle Spatial and Graph: Fast Histogram Construction for Complex Geometry Objects · Proc. VLDB Endow. 2013
Distributed systems › distributed system architecture
decentralized architecture
0.012007
DSphere: A Source-Centric Approach to Crawling, Indexing and Searching the World Wide Web · ICDE 2007
Storage systems › networked storage
storage area network
0.012007
Integrated resource allocation in heterogeneous SAN data centers · PODC 2007
Bioinformatics and computational biology
biomedical text mining
0.012004
BioPatentMiner: An Information Retrieval System for BioMedical Patents · VLDB 2004

Methods — techniques the papers use, named apart from their topics

r-tree index · 0.2partitioning-based algorithms · 0.2top-down cloaking · 0.2temporal cloaking · 0.2bottom-up cloaking · 0.2source-based link analysis · 0.1(ε,δ)k-dissimilarity · 0.1resource allocation · 0.1semantic association discovery · 0.1keyword search · 0.1RDF triple queries · 0.1
YearPublicationVenuePosition
2018 PrivacyZone: A Novel Approach to Protecting Location Privacy of Mobile Users
abstract
While location-based services and applications are increasing in popularity, there are growing concerns over users' location privacy. Although there exist general purpose mobile permission systems and cloaking techniques, these techniques suffer from several problems when applied to continuous location and GPS access, as they are often rigid, coarse-grained, not sufficiently personalizable, and unaware of road network semantics. This paper proposes PrivacyZone, a novel system for constructing personalized fine-grained privacy quarantine regions and protecting users' privacy within these regions. PrivacyZone allows users to seamlessly enter their privacy specifications under spatial, temporal, and semantic customization. Novel challenges arise from having to enforce privacy zones for large volume and variety of users with frequent location updates. We show that naive privacy zone processing techniques are inefficient and cause excessive energy consumption. We therefore develop advanced processing techniques based on the concept of safe hibernation. We empirically evaluate our techniques to demonstrate their trade-offs with respect to hibernation time, computation effort, and network bandwidth usage. Our results show that PrivacyZone is efficient, scalable, and flexible, while preserving users' location privacy.
Emre Yigitoglu, Mehmet Emre Gursoy, Ling Liu 0001, Margaret L. Loper, Bhuvan Bamba, Kisung Lee
IEEE BigData5
2015 Road Network-Aware Anonymization in Mobile Systems with Reciprocity Support
abstract
Most existing solutions on location anonymization fail to address the issue of location privacy protection for mobile users traveling on road networks. In this paper, we present a road network-aware privacy model to handle the location privacy problem in mobile systems. We present third party supported anonymization which can be plugged into existing mobile systems without major modifications. As opposed to current road network specific solutions, our system has the following salient features. First, we argue that the graph density of a cloaked subgraph is an important location cloaking quality measure for road network-based location privacy protection. High graph density leads to high utility of the cloaked location with respect to spatial resolution. Second, we devise a suite of graph-based cloaking algorithms which guarantee reciprocity - an important location cloaking property that many existing approaches fail to support - under an enhanced privacy model. Further, we use controlled randomization in the cloaking process to provide higher privacy strength while maintaining the utility of the cloaked location. Last but not the least, our experimental evaluation shows the efficiency and robustness of the approach in terms of privacy, utility and performance of the model.
Bhuvan Bamba, Ling Liu 0001, Emre Yigitoglu
ICCCN1
2014 Distance queries for complex spatial objects in oracle spatial
abstract
With the proliferation of global positioning systems (GPS) enabled devices, a growing number of database systems are capable of storing and querying different spatial objects including points, polylines and polygons. In this paper, we present our experience with supporting one important class of spatial queries in these database systems: distance queries. For example, a traveler may want to find hotels within 500 meters of a nearby beach. In addition, this paper presents new techniques implemented in Oracle Spatial for some distance-related problems, such as the maximum distance between complex spatial objects, and the diameter, the convex hull and the minimum bounding circle of complex spatial objects. We conduct our experiments by utilizing real-world data sets and demonstrate that these distance and distance-related queries can be significantly improved.
Siva Ravada, Richard Anderson 0003, Bhuvan Bamba
SIGSPATIAL/GIS4
2013 Supporting topological relationship queries for complex line and collection geometries in oracle spatial
abstract
Transportation networks including roads and railways, and linear hydrography features like streams and canals are traditionally represented as complex lines in geographic information systems (GIS) and spatial database systems. In addition, as the Global Positioning System (GPS) becomes increasingly ubiquitous, GIS and spatial database systems are also encountering increasing use of trajectories of moving objects, which can also be represented as complex lines with each vertex not only containing location information, but also associated with some additional measures such as time. In this paper, we present our experience with supporting topological relationship queries for these complex lines. Furthermore, this paper presents our experience with supporting topological relationship queries for complex geometry collections, such as a composite hydrography feature, which can be comprised of complex lines (for narrow portions of rivers) and complex polygons (for wide portions of rivers, and lakes). We conduct our experiments by utilizing real-world data sets and demonstrate that topological relationship query performance for both complex lines and complex collections can be significantly improved.
Siva Ravada, Richard Anderson 0003, Bhuvan Bamba
SIGSPATIAL/GIS4
2013 SLIM: A Scalable Location-Sensitive Information Monitoring Service
abstract
Location-sensitive information monitoring services are a centerpiece of the technology for disseminating content-rich information from massive data streams to mobile users. The key challenges for such monitoring services are characterized by the combination of spatial and non-spatial attributes being monitored and the wide spectrum of update rates. A typical example of such services is "alert me when the gas price at a gas station within 5 miles of my current location drops to 4 per gallon". Such a service needs to monitor the gas price changes in conjunction with the highly dynamic nature of location information. Scalability of such location sensitive and content rich information monitoring services in the presence of different update rates and monitoring thresholds poses a big technical challenge. In this paper, we present SLIM, a scalable location sensitive information monitoring service framework with two unique features. First, we make intelligent use of the correlation between spatial and non-spatial attributes involved in the information monitoring service requests to devise a highly scalable distributed spatial trigger evaluation engine. Second, we introduce single and multi-dimensional safe value containment techniques to efficiently perform selective distributed processing of spatial triggers to reduce the amount of unnecessary trigger evaluations. Through extensive experiments, we show that SLIM offers high scalability for location-sensitive, content-rich information monitoring services in terms of the number of information sources being monitored, number of users and monitoring requests.
Bhuvan Bamba, Kun-Lung Wu, Bugra Gedik, Ling Liu 0001
ICWS1
2013 Statistics Collection in Oracle Spatial and Graph: Fast Histogram Construction for Complex Geometry Objects
abstract
Oracle Spatial and Graph is a geographic information system (GIS) which provides users the ability to store spatial data alongside conventional data in Oracle. As a result of the coexistence of spatial and other data, we observe a trend towards users performing increasingly complex queries which involve spatial as well as non-spatial predicates. Accurate selectivity values, especially for queries with multiple predicates requiring joins among numerous tables, are essential for the database optimizer to determine a good execution plan. For queries involving spatial predicates, this requires that reasonably accurate statistics collection has been performed on the spatial data. For extensible data cartridges such as Oracle Spatial and Graph, the optimizer expects to receive accurate predicate selectivity and cost values from functions implemented within the data cartridge. Although statistics collection for spatial data has been researched in academia for a few years; to the best of our knowledge, this is the first work to present spatial statistics collection implementation details for a commercial GIS database. In this paper, we describe our experiences with implementation of statistics collection methods for complex geometry objects within Oracle Spatial and Graph. Firstly, we exemplify issues with previous partitioning-based algorithms in presence of complex geometry objects and suggest enhancements which resolve the issues. Secondly, we propose a main memory implementation which not only speeds up the disk-based partitioning algorithms but also utilizes existing R-tree indexes to provide surprisingly accurate selectivity estimates. Last but not the least, we provide extensive experimental results and an example study which displays the efficacy of our approach on Oracle query performance.
Bhuvan Bamba, Siva Ravada, Richard Anderson 0003
Proc. VLDB Endow.1
2012 Topological relationship query processing for complex regions in Oracle Spatial
abstract
Although geographic information systems (GIS) and spatial database communities have extensively studied topological relationships for more than two decades, there is little literature describing how to efficiently implement them in GIS and spatial database systems. This is rather surprising considering that topological relationship queries are supported in many GIS and spatial database systems including IBM Informix Spatial and Geodetic DataBlades, ESRI SDE, Microsoft SQL server 2008, Oracle Spatial and PostGIS. In order to bridge this gap, we report our experience with implementing several optimization techniques in Oracle Spatial to speed up topological relationship query processing for query windows represented by complex regions (such as polygons or multi-polygons). Our experiments, utilizing real-world data sets, demonstrate that topological relationship query performance can be significantly improved using the proposed techniques.
Siva Ravada, Richard Anderson 0003, Bhuvan Bamba
SIGSPATIAL/GIS4
2010 RoadTrack: Scaling Location Updates for Mobile Clients on Road Networks with Query Awareness
abstract
Mobile commerce and location based services (LBS) are some of the fastest growing IT industries in the last five years. Location update of mobile clients is a fundamental capability in mobile commerce and all types of LBS. Higher update frequency leads to higher accuracy, but incurs unacceptably high cost of location management at the location servers. We propose RoadTrack -- a road-network based, query-aware location update framework with two unique features. First, we introduce the concept of precincts to control the granularity of location update resolution for mobile clients that are not of interest to any active location query services. Second, we define query encounter points for mobile objects that are targets of active location query services, and utilize these encounter points to define the adequate location update schedule for each mobile. The RoadTrack framework offers three unique advantages. First, encounter points as a fundamental query awareness mechanism enable us to control and differentiate location update strategies for mobile clients in the vicinity of active location queries, while meeting the needs of location query evaluation. Second, we employ system-defined precincts to manage the desired spatial resolution of location updates for different mobile clients and to control the scope of query awareness to be capitalized by a location update strategy. Third, our road-network based check-free interval optimization further enhances the effectiveness of the Road-Track query-aware location update scheduling algorithm. This optimization provides significant cost reduction for location update management at both mobile clients and location servers. We evaluate the RoadTrack location update approach using a real world road-network based mobility simulator. Our experimental results demonstrate that the RoadTrack query aware location update approach outperforms existing representative location update strategies in terms of both client energy efficiency and server processing load.
Péter Pesti, Ling Liu 0001, Bhuvan Bamba, Arun Iyengar, Matt Weber
Proc. VLDB Endow.3
2009 Distributed Processing of Spatial Alarms: A Safe Region-Based Approach
abstract
Spatial alarms are considered as one of the basic capabilities in future mobile computing systems for enabling personalization of location-based services. In this paper, we propose a distributed architecture and a suite of safe region techniques for scalable processing of spatial alarms. We show that safe region-based processing enables resource optimal distribution of partial alarm processing tasks from the server to the mobile clients. We propose three different safe region computation algorithms to explore the impact of size and shape of the safe region on network bandwidth, server load and client energy consumption. Concretely, we show that the maximum weighted perimeter rectangular safe region approach outperforms previous techniques in terms of performance and accuracy. We further explore finer granularity safe regions by introducing grid-based and pyramid-based representation of rectilinear polygonal shapes using bitmap encoding. Our experimental evaluation shows that the distributed safe region-based architecture outperforms the two most popular server-centric approaches, periodic and safe period-based, for spatial alarm processing.
Bhuvan Bamba, Ling Liu 0001, Arun Iyengar, Philip S. Yu
ICDCS1
2009 A General Proximity Privacy Principle
abstract
This work presents a systematic study of the problem of protecting general proximity privacy, with findings applicable to most existing data models. Our contributions are multi-folded: we highlighted and formulated proximity privacy breaches in a data-model-neutral manner; we proposed a new privacy principle (epsiv,delta)k-dissimilarity, with theoretically guaranteed protection against linking attacks in terms of both exact and proximate QI-SA associations; we provided a theoretical analysis regarding the satisfiability of (epsiv,delta)k-dissimilarity, and pointed to promising solutions to fulfilling this principle.
Ting Wang 0006, Shicong Meng, Bhuvan Bamba, Ling Liu 0001, Calton Pu
ICDE3
2009 Scalable and Reliable Location Services through Decentralized Replication
abstract
One of the critical challenges for service oriented computing systems is the capability to guarantee scalable and reliable service provision. This paper presents Reliable GeoGrid, a decentralized service computing architecture based on geographical location aware overlay network for supporting reliable and scalable mobile information delivery services. The reliable GeoGrid approach offers two distinct features. First, we develop a distributed replication scheme, aiming at providing scalable and reliable processing of location service requests in decentralized pervasive computing environments. Our replica management operates on a network of heterogeneous nodes and utilizes a shortcut-based optimization to increase the resilience of the system against node failures and network failures. Second, we devise a dynamic load balancing technique that exploits the service processing capabilities of replicas to scale the system in anticipation of unexpected workload changes and node failures by taking into account of node heterogeneity, network proximity, and changing workload at each node. Our experimental evaluation shows that the reliable GeoGrid architecture is highly scalable under changing service workloads with moving hotspots and highly reliable in the presence of both individual node failures and massive node failures.
Gong Zhang 0008, Ling Liu 0001, Sangeetha Seshadri, Bhuvan Bamba, Yuehua Wang
ICWS4
2009 Coupled placement in modern data centers
abstract
We introduce the coupled placement problem for modern data centers spanning placement of application computation and data among available server and storage resources. While the two have traditionally been addressed independently in data centers, two modern trends make it beneficial to consider them together in a coupled manner: (a) rise in virtualization technologies, which enable applications packaged as VMs to be run on any server in the data center with spare compute resources, and (b) rise in multi-purpose hardware devices in the data center which provide compute resources of varying capabilities at different proximities from the storage nodes. We present a novel framework called CPA for addressing such coupled placement of application data and computation in modern data centers. Based on two well-studied problems - Stable Marriage and Knapsacks - the CPA framework is simple, fast, versatile and automatically enables high throughput applications to be placed on nearby server and storage node pairs. While a theoretical proof of CPA's worst-case approximation guarantee remains an open question, we use extensive experimental analysis to evaluate CPA on large synthetic data centers comparing it to Linear Programming based methods and other traditional methods. Experiments show that CPA is consistently and surprisingly within 0 to 4% of the Linear Programming based optimal values for various data center topologies and workload patterns. At the same time it is one to two orders of magnitude faster than the LP based methods and is able to scale to much larger problem sizes. The fast running time of CPA makes it highly suitable for large data center environments where hundreds to thousands of server and storage nodes are common. LP based approaches are prohibitively slow in such environments. CPA is also suitable for fast interactive analysis during consolidation of such environments from physical to virtual resources.
Madhukar R. Korupolu, Aameek Singh, Bhuvan Bamba
IPDPS3
2008 Scalable Processing of Spatial Alarms
Bhuvan Bamba, Ling Liu 0001, Philip S. Yu, Gong Zhang 0008, Myungcheol Doo
HiPC1
2008 Supporting anonymous location queries in mobile environments with privacygrid
abstract
This paper presents PrivacyGrid - a framework for supporting anonymous location-based queries in mobile information delivery systems. The PrivacyGrid framework offers three unique capabilities. First, it provides a location privacy protection preference profile model, called location P3P, which allows mobile users to explicitly define their preferred location privacy requirements in terms of both location hiding measures (e.g., location k-anonymity and location l-diversity) and location service quality measures (e.g., maximum spatial resolution and maximum temporal resolution). Second, it provides fast and effective location cloaking algorithms for location k-anonymity and location l-diversity in a mobile environment. We develop dynamic bottom-up and top-down grid cloaking algorithms with the goal of achieving high anonymization success rate and efficiency in terms of both time complexity and maintenance cost. A hybrid approach that carefully combines the strengths of both bottom-up and top-down cloaking approaches to further reduce the average anonymization time is also developed. Last but not the least, PrivacyGrid incorporates temporal cloaking into the location cloaking process to further increase the success rate of location anonymization. We also discuss PrivacyGrid mechanisms for supporting anonymous location queries. Experimental evaluation shows that the PrivacyGrid approach can provide close to optimal location k-anonymity as defined by per user location P3P without introducing significant performance penalties.
Bhuvan Bamba, Ling Liu 0001, Péter Pesti, Ting Wang 0006
WWW1
2007 DSphere: A Source-Centric Approach to Crawling, Indexing and Searching the World Wide Web
abstract
We describe DSphere - a decentralized system for crawling, indexing, searching and ranking of documents in the World Wide Web. Unlike most of the existing search technologies that depend heavily on a page-centric view of the Web, we advocate a source-centric view of the Web and propose a decentralized architecture for crawling, indexing and searching the Web in a distributed source-specific fashion. A fully decentralized crawler is developed to crawl the World Wide Web where each peer is assigned the responsibility of crawling a specific set of documents referred to as a source collection. Link analysis techniques are used for ranking documents. Traditional link analysis techniques suffer from problems like slow refresh rate and vulnerabilities to Web Spam. We propose a source-based link analysis approach, which computes fast and accurate ranking scores for all crawled documents.
Bhuvan Bamba, Ling Liu 0001, James Caverlee, Vaibhav Padliya, Mudhakar Srivatsa, Tushar Bansal, Mahesh Palekar, Joseph Patrao, Suiyang Li, Aameek Singh
ICDE1
2007 Integrated resource allocation in heterogeneous SAN data centers
abstract
Modern data centers are complex distributed environments with application workloads requiring multiple resources like processing (CPU), storage and network. Allocation of these resources to workloads needs to be handled in an integrated manner to adequately capture the relationships between different resource nodes like connectivity between an application server and storage controller in the storage area network (SAN). As data centers grow over time, heterogeneous resources coexist at the same time and this heterogeneity adds further complexity to manual resource allocation.
Aameek Singh, Madhukar R. Korupolu, Bhuvan Bamba
PODC3
2005 OSQR: overlapping clustering of query results
abstract
No abstract available.
Bhuvan Bamba, Prasan Roy, Mukesh K. Mohania
CIKM1
2005 Towards automatic association of relevant unstructured content with structured query results
abstract
Faced with growing knowledge management needs, enterprises are increasingly realizing the importance of seamlessly integrating critical business information distributed across both structured and unstructured data sources. In existing information integration solutions, the application needs to formulate the SQL logic to retrieve the needed structured data on one hand, and identify a set of keywords to retrieve the related unstructured data on the other. This paper proposes a novel approach wherein the application specifies its information needs using only a SQL query on the structured data, and this query is automatically ``translated'' into a set of keywords that can be used to retrieve relevant unstructured data. We describe the techniques used for obtaining these keywords from (i) the query result, and (ii) additional related information in the underlying database. We further show that these techniques achieve high accuracy with very reasonable overheads.
Prasan Roy, Mukesh K. Mohania, Bhuvan Bamba, Shree Raman
CIKM3
2005 Information Retrieval and Knowledge Discovery Utilizing a BioMedical Patent Semantic Web
abstract
Before undertaking new biomedical research, identifying concepts that have already been patented is essential. A traditional keyword-based search on patent databases may not be sufficient to retrieve all the relevant information, especially for the biomedical domain. This paper presents BioPatentMiner, a system that facilitates information retrieval and knowledge discovery from biomedical patents. The system first identifies biological terms and relations from the patents and then integrates the information from the patents with knowledge from biomedical ontologies to create a semantic Web. Besides keyword search and queries linking the properties specified by one or more RDF triples, the system can discover semantic associations between the Web resources. The system also determines the importance of the resources to rank the results of a search and prevent information overload while determining the semantic associations.
Sougata Mukherjea, Bhuvan Bamba, Pankaj Kankar
IEEE Trans. Knowl. Data Eng.2
2004 BioPatentMiner: An Information Retrieval System for BioMedical Patents
Sougata Mukherjea, Bhuvan Bamba
VLDB2