Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sriharsha Gangam

dblp:57/9801 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 2 · 2 first-authorSystems, architecture and hardware · 1Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
2 papers
Network measurement and analytics · 100%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 8 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
materialized view
0.212016
Kodiak: Leveraging Materialized Views For Very Low-Latency Analytics Over High-Dimensional Web-Scale Data · Proc. VLDB Endow. 2016
Network measurement and analytics
anomaly detection
0.212013
Pegasus: Precision hunting for icebergs and anomalies in network flows · INFOCOM 2013
Network measurement and analytics
flow monitoring
0.212013
Pegasus: Precision hunting for icebergs and anomalies in network flows · INFOCOM 2013
Network measurement and analytics
heavy hitter detection
0.212013
Pegasus: Precision hunting for icebergs and anomalies in network flows · INFOCOM 2013
Network measurement and analytics
network inference
0.112011
On the Cost of Network Inference Mechanisms · IEEE Trans. Parallel Distributed Syst. 2011
Network measurement and analytics › network inference
path inference
0.112011
On the Cost of Network Inference Mechanisms · IEEE Trans. Parallel Distributed Syst. 2011
Network measurement and analytics › measurement infrastructure
distributed monitoring
0.012013
Pegasus: Precision hunting for icebergs and anomalies in network flows · INFOCOM 2013
Network measurement and analytics
end-to-end measurement
0.012011
On the Cost of Network Inference Mechanisms · IEEE Trans. Parallel Distributed Syst. 2011

Methods — techniques the papers use, named apart from their topics

view materialization · 0.5query auto-selection · 0.5sketching · 0.2adaptive data transfer · 0.2synthetic data evaluation · 0.1algorithmic complexity analysis · 0.1
YearPublicationVenuePosition
2016 Kodiak: Leveraging Materialized Views For Very Low-Latency Analytics Over High-Dimensional Web-Scale Data
abstract
Turn's online advertising campaigns produce petabytes of data. This data is composed of trillions of events, e.g. impressions, clicks, etc., spanning multiple years. In addition to a timestamp, each event includes hundreds of fields describing the user's attributes, campaign's attributes, attributes of where the ad was served, etc. Advertisers need advanced analytics to monitor their running campaigns' performance, as well as to optimize future campaigns. This involves slicing and dicing the data over tens of dimensions over arbitrary time ranges. Many of these queries need to power the web portal to provide reports and dashboards. For an interactive response time, they have to have tens of milliseconds latency. At Turn's scale of operations, no existing system was able to deliver this performance in a cost effective manner. Kodiak, a distributed analytical data platform for web-scale high-dimensional data, was built to serve this need. It relies on pre-computations to materialize thousands of views to serve these advanced queries. These views are partitioned and replicated across Kodiak's storage nodes for scalability and reliability. They are system maintained as new events arrive. At query time, the system auto-selects the most suitable view to serve each query. Kodiak has been used in production for over a year. It hosts 2490 views for over three petabytes of raw data serving over 200K queries daily. It has median and 99% query latencies of 8 ms and 252 ms respectively. Our experiments show that its query latency is 3 orders of magnitude faster than leading big data platforms on head-to-head comparisons using Turn's query workload. Moreover, Kodiak uses 4 orders of magnitude less resources to run the same workload.
Shaosu Liu, Sriharsha Gangam, Lawrence Lo, Khaled Elmeleegy
Proc. VLDB Endow.3
2013 Pegasus: Precision hunting for icebergs and anomalies in network flows
abstract
Accurate online network monitoring is crucial for detecting attacks, faults, and anomalies, and determining traffic properties across the network. With high bandwidth links and consequently increasing traffic volumes, it is difficult to collect and analyze detailed flow records in an online manner. Traditional solutions that decouple data collection from analysis resort to sampling and sketching to handle large monitoring traffic volumes. We propose a new system, Pegasus, to leverage commercially available co-located compute and storage devices near routers and switches. Pegasus adaptively manages data transfers between monitors and aggregators based on traffic patterns and user queries. We use Pegasus to detect global icebergs or global heavy-hitters. Icebergs are flows with a common property that contribute a significant fraction of network traffic. For example, DDoS attack detection is an iceberg detection problem with a common destination IP. Other applications include identification of “top talkers,” top destinations, and detection of worms and port scans. Experiments with Abilene traces, sFlow traces from an enterprise network, and deployment of Pegasus as a live monitoring service on PlanetLab show that our system is accurate and scales well with increasing traffic and number of monitors.
Sriharsha Gangam, Puneet Sharma 0001, Sonia Fahmy
INFOCOM1
2013 Estimating TCP Latency Approximately with Passive Measurements
Sriharsha Gangam, Jaideep Chandrashekar, Ítalo S. Cunha, James F. Kurose
PAM1
2011 Mitigating interference in a network measurement service
abstract
Shared measurement services offer key advantages over conventional ad-hoc techniques for network monitoring. A measurement service may receive measurement requests concurrently from different applications and network administrators. These measurement requests are often served by injecting active network measurement traffic between two hosts. Two active measurements are said to interfere when the probe packets of one measurement tool are viewed as network traffic by the other. This may lead to faulty measurement readings. In this paper, we model the measurement interference problem, and show how to schedule measurement tasks to reduce interference and hence increase measurement accuracy. We propose twelve computationally tractable algorithms that decrease the total completion time (makespan) of measurement tasks, while avoiding interference. Our evaluation shows that the algorithm we refer to as Largest Area First, Busiest Node First - Earliest Interval Schedule (LAFBNF-EIS) has a mean makespan of about 5% more than the theoretical lower bound over our set of measurement workloads.
Sriharsha Gangam, Sonia Fahmy
IWQoS1
2011 On the Cost of Network Inference Mechanisms
abstract
A number of network path delay, loss, or bandwidth inference mechanisms have been proposed over the past decade. Concurrently, several network measurement services have been deployed over the Internet and intranets. We consider inference mechanisms that use O(n) end-to-end measurements to predict the O(n2) end-to-end pairwise measurements among n nodes, and investigate when it is beneficial to use them in measurement services. In particular, we address the following questions : 1) For which measurement request patterns would using an inference mechanism be advantageous? 2) How does a measurement service determine the set of hosts that should utilize inference mechanisms, as opposed to those that are better served using direct end-to-end measurements? We explore three solutions that identify groups of hosts which are likely to benefit from inference. We compare these solutions in terms of effectiveness and algorithmic complexity. Results with synthetic data sets and data sets from a popular peer-to-peer system demonstrate that our techniques accurately identify host subsets that benefit from inference, in significantly less time than an algorithm that identifies optimal subsets. The measurement savings are large when measurement request patterns exhibit small-world characteristics, which is often the case. (Part of this work (focusing on one of three solutions presented in this paper) appeared in).
Ethan Blanton, Sonia Fahmy, Greg N. Frederickson, Sriharsha Gangam
IEEE Trans. Parallel Distributed Syst.4