Amit Manjhi

dblp:69/2448 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
0since 2021 · last 2009
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 5 first-authorComputer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Query processing and optimization · 77% Data stream processing · 23%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
Cloud and datacenter computing · 58% Distributed systems · 42%
Network and information security
2 papers
Privacy and data protection · 53% Cryptographic primitives and cryptanalysis · 36% Web and mobile security · 11%
Computer networks
2 papers
Content delivery and video streaming · 50% Internet of things and sensor networks · 43% Vehicular, aerial and satellite networks · 8%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
query result caching
0.222008
Scalable query result caching for web applications · Proc. VLDB Endow. 2008
Invalidation Clues for Database Scalability Services · ICDE 2007
Query processing and optimization › query rewriting
query transformation
0.112009
Holistic Query Transformations for Dynamic Web Applications · ICDE 2009
Compilers and program optimization
program transformation
0.112009
Holistic Query Transformations for Dynamic Web Applications · ICDE 2009
Distributed systems › consistency models
cache consistency
0.112008
Scalable query result caching for web applications · Proc. VLDB Endow. 2008
Privacy and data protection
privacy-preserving data management
0.112007
Invalidation Clues for Database Scalability Services · ICDE 2007
Cryptographic primitives and cryptanalysis
encryption
0.112006
Simultaneous scalability and security for data-intensive web applications · SIGMOD Conference 2006
Query processing and optimization › approximate query processing
approximate aggregation
0.112005
Tributaries and Deltas: Efficient and Robust Aggregation in Sensor Network Streams · SIGMOD Conference 2005
Query processing and optimization
approximate query processing
0.112005
Finding (Recently) Frequent Items in Distributed Data Streams · ICDE 2005
Data stream processing
distributed data streams
0.112005
Finding (Recently) Frequent Items in Distributed Data Streams · ICDE 2005
Data stream processing
frequency estimation
0.112005
Finding (Recently) Frequent Items in Distributed Data Streams · ICDE 2005
Internet of things and sensor networks › wireless sensor network
in-network aggregation
0.112005
Tributaries and Deltas: Efficient and Robust Aggregation in Sensor Network Streams · SIGMOD Conference 2005
Content delivery and video streaming › caching
web caching
0.012003
Improving Web Performance in Broadcast-Unicast Networks · INFOCOM 2003
Content delivery and video streaming
web performance
0.012003
Improving Web Performance in Broadcast-Unicast Networks · INFOCOM 2003
Privacy and data protection
data confidentiality
0.012006
Simultaneous scalability and security for data-intensive web applications · SIGMOD Conference 2006
Web and mobile security
web application security
0.012006
Simultaneous scalability and security for data-intensive web applications · SIGMOD Conference 2006
Internet of things and sensor networks › iot networks › iot communication
sensor network communication
0.012005
Tributaries and Deltas: Efficient and Robust Aggregation in Sensor Network Streams · SIGMOD Conference 2005
Distributed systems
communication optimization
0.012005
Finding (Recently) Frequent Items in Distributed Data Streams · ICDE 2005
Vehicular, aerial and satellite networks
satellite networks
0.012003
Improving Web Performance in Broadcast-Unicast Networks · INFOCOM 2003

Methods — techniques the papers use, named apart from their topics

source-to-source compilation · 0.3latency hiding · 0.3invalidation clues · 0.2encryption · 0.2publish-subscribe · 0.2proxy-based cooperative caching · 0.2simulation · 0.1static analysis of database segments · 0.1precision gradient · 0.1multi-path aggregation · 0.1hierarchical aggregation · 0.1tree-based aggregation · 0.1heuristic scheduler · 0.0approximation algorithm · 0.0
YearPublicationVenuePosition
2009 Holistic Query Transformations for Dynamic Web Applications
abstract
A promising approach to scaling Web applications is to distribute the server infrastructure on which they run. This approach, unfortunately, can introduce latency between the application and database servers, which in turn increases the network latency of Web interactions for the clients (end users). In this paper we introduce the concept of source-to-source holistic transformations - transformations that seek to optimize both the application code and the database requests made by it, to reduce client latency. As examples of our concept, we propose and evaluate two source-to-source holistic transformations that focus on hiding the latencies of database queries. We argue that opportunities for applying these transformations will continue to exist in Web applications. We then present algorithms for automating these transformations in a source-to-source compiler. Finally, we evaluate the effect of these two transformations on three realistic Web benchmark applications, both in the traditional centralized setting and a distributed setting.
Amit Manjhi, Charles Garrod, Bruce M. Maggs, Todd C. Mowry, Anthony Tomasic
ICDE1
2008 Scalable query result caching for web applications
abstract
The backend database system is often the performance bottleneck when running web applications. A common approach to scale the database component is query result caching, but it faces the challenge of maintaining a high cache hit rate while efficiently ensuring cache consistency as the database is updated. In this paper we introduce Ferdinand, the first proxy-based cooperative query result cache with fully distributed consistency management. To maintain a high cache hit rate, Ferdinand uses both a local query result cache on each proxy server and a distributed cache. Consistency management is implemented with a highly scalable publish/subscribe system. We implement a fully functioning Ferdinand prototype and evaluate its performance compared to several alternative query-caching approaches, showing that our high cache hit rate and consistency management are both critical for Ferdinand's performance gains over existing systems.
Charles Garrod, Amit Manjhi, Anastasia Ailamaki, Bruce M. Maggs, Todd C. Mowry, Christopher Olston, Anthony Tomasic
Proc. VLDB Endow.2
2007 Invalidation Clues for Database Scalability Services
abstract
For their scalability needs, data-intensive Web applications can use a database scalability service (DBSS), which caches applications' query results and answers queries on their behalf. One way for applications to address their security/privacy concerns when using a DBSS is to encrypt all data that passes through the DBSS. Doing so, however, causes the DBSS to invalidate large regions of its cache when data updates occur. To invalidate more precisely, the DBSS needs help in order to know which results to invalidate; such help inevitably reveals some properties about the data. In this paper, we present invalidation clues, a general technique that enables applications to reveal little data to the DBSS, yet limit the number of unnecessary invalidations. Compared with previous approaches, invalidation clues provide applications significantly improved tradeoffs between security/privacy and scalability. Our experiments using three Web application benchmarks, on a prototype DBSS we have built, confirm that invalidation clues are indeed a low-overhead, effective, and general technique for applications to balance their privacy and scalability needs.
Amit Manjhi, Phillip B. Gibbons, Anastasia Ailamaki, Charles Garrod, Bruce M. Maggs, Todd C. Mowry, Christopher Olston, Anthony Tomasic
ICDE1
2006 Simultaneous scalability and security for data-intensive web applications
abstract
For Web applications in which the database component is the bottleneck, scalability can be provided by a third-party Database Scalability Service Provider (DSSP) that caches application data and supplies query answers on behalf of the application. Cost-effective DSSPs will need to cache data from many applications, inevitably raising concerns about security. However, if all data passing through a DSSP is encrypted to enhance security, then data updates trigger invalidation of large regions of cache. Consequently, achieving good scalability becomes virtually impossible. There is a tradeoff between security and scalability, which requires careful consideration.In this paper we study the security-scalability tradeoff, both formally and empirically. We begin by providing a method for statically identifying segments of the database that can be encrypted without impacting scalability. Experiments over a prototype DSSP system show the effectiveness of our static analysis method--for all three realistic bench-mark applications that we study, our method enables a significant fraction of the database to be encrypted without impacting scalability. Moreover, most of the data that can be encrypted without impacting scalability is of the type that application designers will want to encrypt, all other things being equal. Based on our static analysis method, we propose a new scalability-conscious security design methodology that features: (a) compulsory encryption of highly sensitive data like credit card information, and (b) encryption of data for which encryption does not impair scalability. As a result, the security-scalability tradeoff needs to be considered only over data for which encryption impacts scalability, thus greatly simplifying the task of managing the tradeoff.
Amit Manjhi, Anastasia Ailamaki, Bruce M. Maggs, Todd C. Mowry, Christopher Olston, Anthony Tomasic
SIGMOD Conference1
2005 A Scalability Service for Dynamic Web Applications
Christopher Olston, Amit Manjhi, Charles Garrod, Anastasia Ailamaki, Bruce M. Maggs, Todd C. Mowry
CIDR2
2005 Finding (Recently) Frequent Items in Distributed Data Streams
abstract
We consider the problem of maintaining frequency counts for items occurring frequently in the union of multiple distributed data streams. Naive methods of combining approximate frequency counts from multiple nodes tend to result in excessively large data structures that are costly to transfer among nodes. To minimize communication requirements, the degree of precision maintained by each node while counting item frequencies must be managed carefully. We introduce the concept of a precision gradient for managing precision when nodes are arranged in a hierarchical communication structure. We then study the optimization problem of how to set the precision gradient so as to minimize communication, and provide optimal solutions that minimize worst-case communication load over all possible inputs. We then introduce a variant designed to perform well in practice, with input data that does not conform to worst-case characteristics. We verify the effectiveness of our approach empirically using real-world data, and show that our methods incur substantially less communication than naive approaches while providing the same error guarantees on answers.
Amit Manjhi, Vladislav Shkapenyuk, Kedar Dhamdhere, Christopher Olston
ICDE1
2005 Tributaries and Deltas: Efficient and Robust Aggregation in Sensor Network Streams
abstract
Existing energy-efficient approaches to in-network aggregation in sensor networks can be classified into two categories, tree-based and multi-path-based, with each having unique strengths and weaknesses. In this paper, we introduce Tributary-Delta, a novel approach that combines the advantages of the tree and multi-path approaches by running them simultaneously in different regions of the network. We present schemes for adjusting the regions in response to changes in network conditions, and show how many useful aggregates can be readily computed within this new framework. We then show how a difficult aggregate for this context---finding frequent items---can be efficiently computed within the framework. To this end, we devise the first algorithm for frequent items (and for quantiles) that provably minimizes the worst case total communication for non-regular trees. In addition, we give a multi-path algorithm for frequent items that is considerably more accurate than previous approaches. These algorithms form the basis for our efficient Tributary-Delta frequent items algorithm. Through extensive simulation with real-world and synthetic data, we show the significant advantages of our techniques. For example, in computing Count under realistic loss rates, our techniques reduce answer error by up to a factor of 3 compared to any previous technique.
Amit Manjhi, Suman Nath, Phillip B. Gibbons
SIGMOD Conference1
2003 Improving Web Performance in Broadcast-Unicast Networks
abstract
Satellite operators have recently begun offering Internet access over their networks. Typically, users connect to the network using a modem for uplink, and a satellite dish for downlink. We investigate how the performance of these networks might be improved by two simple techniques: caching and use of the return path on the modem link. We examine the problem from a theoretical perspective and via simulation. We show that the general problem is NP-hard, as are several special cases, and we give approximation algorithms for them. We then use insights from these cases to design practical heuristic schedulers which leverage caching and the modem downlinks. Via simulation, we show that caching alone can simultaneously reduce bandwidth requirements by 33% and improve response times by 62%. We further show that the proposed schedulers, combined with caching, yield a system that performs far better under high loads than existing systems.
Mukesh Agrawal 0002, Amit Manjhi, Nikhil Bansal 0001, Srinivasan Seshan
INFOCOM2