Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

William Acosta

dblp:62/4256 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
0since 2021 · last 2012
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 first-authorComputer networks · 2 · 1 first-authorSecurity and privacy · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 77% Database system architecture and tuning · 23%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › distributed information retrieval
peer-to-peer search
0.112008
Exploiting the Properties of Query Workload and File Name Distributions to Improve P2P Synopsis-Based Searches · INFOCOM 2008
Distributed systems › peer-to-peer systems › peer-to-peer search
hybrid search
0.112008
Exploiting the Properties of Query Workload and File Name Distributions to Improve P2P Synopsis-Based Searches · INFOCOM 2008
Distributed systems
peer-to-peer systems
0.112008
Exploiting the Properties of Query Workload and File Name Distributions to Improve P2P Synopsis-Based Searches · INFOCOM 2008
Database system architecture and tuning › workload management
query workload analysis
0.012008
Exploiting the Properties of Query Workload and File Name Distributions to Improve P2P Synopsis-Based Searches · INFOCOM 2008

Methods — techniques the papers use, named apart from their topics

distribution analysis · 0.2adaptive algorithm design · 0.2
YearPublicationVenuePosition
2012 Implications of the file names and user requested queries on Gnutella performance
Surendar Chandra, William Acosta
Peer-to-Peer Netw. Appl.2
2008 Exploiting the Properties of Query Workload and File Name Distributions to Improve P2P Synopsis-Based Searches
abstract
Modern P2P systems use hybrid searches to improve search efficiency. They use a synopsis of neighborhood content to determine whether to use a structured or unstructured overlay to satisfy a particular query. Because of their size restrictions, a synopsis cannot hold all the terms from every file in the neighborhood. The challenge is to choose the terms that should be represented in the synopsis. In this work, we investigated the distribution of query terms and file terms in Gnutella networks. We observed that there was a mismatch between terms that were popular among file names and the terms that were popular among the queries generated by the user. Because the query behavior changed with time, a synopsis based on only static set of popular file terms was ill-suited to support efficient searches. We used these observations to design a synopsis creation algorithm that dynamically adapted to the query workload and selected terms for the synopsis to reflect popular terms in both the query workload and file distribution. Our preliminary experimental analysis showed that our Query-Adaptive synopsis improved the search performance over the traditional file-based synopsis model.
William Acosta, Surendar Chandra
INFOCOM1
2008 On the need for query-centric unstructured peer-to-peer overlays
abstract
Hybrid P2P systems rely on the assumption that sufficient objects exist nearby in order to make the unstructured search component efficient. This availability depends on the object annotations as well as on the terms in the queries. Earlier work assumed that the object annotations and query terms follow Zipf-like long-tail distribution. We show that the queries in real systems exhibit more complex temporal behavior. To support our position, first we analyzed the names and annotations of objects that were stored in two popular P2P sharing systems; Gnutella and Apple iTunes. We showed that the names and annotations exhibited a Zipf like long tail distribution. The long tail meant that over 98% of the objects were insufficiently replicated (less than 0.1% of the peers). We also analyzed a query trace of the Gnutella network and identified the popularity distribution of the terms used in the queries. We showed that the set of popular query terms remained stable over time and exhibited a similarity of over 90%. We also showed that despite the Zipf popularity distributions of both query terms and file annotation terms, there was little similarity over time (<20%) between popular file annotation terms and popular file terms. Prior P2P search performance analysis did not take this mismatch between the query terms and object annotations into account and thus overestimated the system performance. There is a need to develop unstructured P2P systems that are aware of the temporal mismatch of the object and query popularity distributions.
William Acosta, Surendar Chandra
IPDPS1
2007 Improving Search Using a Fault-Tolerant Overlay in Unstructured P2P Systems
abstract
Gnutella overlays have evolved to use a two-tier topology. However, we observed that the new topology had only achieved modest improvements in search success rates. Also, the new two-tier topology had not reduced the message routing overhead and bandwidth consumption. In this work, we used local information at each node to construct an overlay, Makalu, that improved search performance and reduced bandwidth consumption. The overlay maximized the expansion from each node's neighborhood while minimizing the latency to its neighbors. We show that for a 100,000 node system, wild card searches using flooding successfully resolved most queries within four hops for object replications ratios as lows as 0.05% (50 randomly distributed copies) with less than 3% duplicate messages. Using attenuated bloom filters to route messages for exact identifier searches, we show that Makalu resolved most queries with less than ten messages for networks as large as 100,000 nodes. The performance of this search is comparable to that of structured P2P systems. Finally, using data from traffic traces of Gnutella in 2003 and 2006, we demonstrated search success rates that were up to five times more successful and required 75% less bandwidth on a Makalu overlay than on a modern Gnutella overlay.
William Acosta, Surendar Chandra
ICPP1
2007 Trace Driven Analysis of the Long Term Evolution of Gnutella Peer-to-Peer Traffic
William Acosta, Surendar Chandra
PAM1