Arpit Gupta

dblp:122/9141 · DBLP profile ↗
← Back
42ranked-venue papers
9as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 25 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 9 · 8 since 2021Security and privacy · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TurboTest: Learning When Less is Enough through Early Termination of Internet Speed Tests
Haarika Manda, Manshi Sagar, Yogesh, Kartikay Singh, Cindy Zhao, Tarun Mangla, Phillipa Gill, Elizabeth M. Belding, Arpit Gupta
NSDI9
2026 SPLIDT: Partitioned Decision Trees for Scalable Stateful Inference at Line Rate
Murayyiam Parvez, Annus Zulfiqar, Roman Beltiukov, Shir Landau Feibish, Walter Willinger, Arpit Gupta, Muhammad Shahbaz 0001
NSDI6
2025 Poster: Uncovering Lesser-Studied RTC Applications using RTP
abstract
Video conferencing and, more generally, real-time communications (RTC) applications have become indispensable across domains such as remote work or telemedicine, with new platforms competing for performance and market share. Much of the research to date has focused on understanding how these systems operate and perform under varying network conditions. However, the scale of the ecosystem for RTC applications beyond a handful of key players remains largely understudied and unquantified. Frameworks such as WebRTC have democratized the ability to easily extend an application to support RTC capabilities, resulting in a diverse ecosystem beyond the most commonly studied applications.
Sanjay Chandrasekaran, Jaber Daneshamooz, Oliver Michel, Arpit Gupta
IMC4
2025 Poster: Enabling Data-Driven Policymaking using the Broadband-Plan Querying Tool (BQT+)
abstract
Poor broadband access undermines civic and economic life, a challenge exacerbated by the fact that millions of Americans still lack access to reliable high-speed broadband connectivity. Federal broadband funding initiatives such as the Connect America Fund (CAF), the Rural Digital Opportunity Fund (RDOF), and the Broadband Equity, Access, and Deployment (BEAD) seek to address these gaps, but their success relies on the accuracy of broadband availability and affordability data. This data is often based on self-reported ISP information that overstates coverage and speeds. Such inaccuracies risk misallocating funds, leaving unserved and underserved communities without high-speed Internet. In this work, we present BQT+, an ML-based data collection platform that queries ISP web interfaces by inputting residential street addresses and extracting the returned data on service availability, quality, and pricing. BQT+ has been applied in policy evaluation studies, including an independent assessment of the state of broadband availability, quality (available speed tiers), and affordability in areas expected to benefit from the $42.45 billion BEAD program.
Laasya Koduru, Tejas N. Narechania, Elizabeth M. Belding, Arpit Gupta
IMC4
2025 Demystifying Network Foundation Models
abstract
This work presents a systematic investigation into the latent knowledge encoded within Network Foundation Models (NFMs). Different from existing efforts, we focus on hidden representations analysis rather than pure downstream task performance and analyze NFMs through a three-part evaluation: Embedding Geometry Analysis to assess representation space utilization, Metric Alignment Assessment to measure correspondence with domain-expert features, and Causal Sensitivity Testing to evaluate robustness to protocol perturbations. Using five diverse network datasets spanning controlled and real-world environments, we evaluate four state-of-the-art NFMs, revealing that they all exhibit significant anisotropy, inconsistent feature sensitivity patterns, an inability to separate the high-level context, payload dependency, and other properties. Our work identifies numerous limitations across all models and demonstrates that addressing them can significantly improve model performance (up to 0.35 increase in $F_1$ scores without architectural changes).
Roman Beltiukov, Satyandra Guthula, Wenbo Guo 0002, Walter Willinger, Arpit Gupta
NeurIPS5
2025 PreTE: Traffic Engineering with Predictive Failures
abstract
Fiber links in wide-area networks (WANs) are exposed to complicated environments and hence are vulnerable to failures like fiber cuts. The conventional approach of using static probabilistic failures falls short in fiber-cut scenarios because these fiber cuts are rare but disruptive, making it difficult for network operators to balance network utilization and availability in WAN traffic engineering. Our large-scale measurements of per-second optical-layer data reveal that the fiber's failure probability increases by several orders of magnitude when experiencing a rare and ephemeral degradation state. Therefore, we present a novel traffic engineering (TE) system called PreTE to factor in the dynamic fiber cut probabilities directly into TE systems. At the core of the PreTE system, fiber degradation facilitates failure predictions and traffic tunnels to be proactively updated, followed by traffic allocation optimizations among updated tunnels. We evaluate PreTE using a production-level WAN testbed and large-scale simulations. The testbed evaluation quantifies PreTE's runtime to demonstrate the feasibility to implement in large-scale WANs. Our large-scale simulation results show that PreTE can support up to 2× more demand at the same level of availability as compared to existing TE schemes.
Congcong Miao, Zhizhen Zhong, Arpit Gupta, Ying Zhang 0022, Zekun He, Xianneng Zou, Jilong Wang 0001
SIGCOMM4
2025 SpliDT: Partitioned Decision Trees for Scalable Stateful Inference at Line Rate
abstract
Machine learning is increasingly used in programmable data planes, such as switches [4, 12, 13] and smartNICs [1, 16], to enable real-time traffic analysis and security monitoring at line rate. Decision trees (DTs) are particularly well-suited for these tasks due to their interpretability and compatibility with the Reconfigurable Match-Action Table (RMT) architecture. However, current DT implementations require collecting all features upfront, which limits scalability and accuracy due to constrained data plane resources.
Murayyiam Parvez, Annus Zulfiqar, Roman Beltiukov, Shir Landau Feibish, Walter Willinger, Arpit Gupta, Muhammad Shahbaz 0001
SIGCOMM6
2024 DiNADO: Norm-Disentangled Neurally-Decomposed Oracles for Controlling Language Models
abstract
NeurAlly-Decomposed Oracle (NADO) is a powerful approach for controllable generation with large language models. It is designed to avoid catastrophic forgetting while achieving guaranteed convergence to an entropy-maximized closed-form optimal solution with reasonable modeling capacity. Despite the success, several challenges arise when apply NADO to a wide range of scenarios. Vanilla NADO suffers from gradient vanishing for low-probability control signals and is highly reliant on a regularization to satisfy the stochastic version of Bellman equation. In addition, the vanilla implementation of NADO introduces a few additional transformer layers, suffering from a limited capacity especially compared to other finetune-based model adaptation methods like LoRA. In this paper, we propose a improved version of the NADO algorithm, namely DiNADO (norm-Disentangled NeurAlly-Decomposed Oracles), which improves the performance of the NADO algorithm through disentangling the step-wise global norm over the approximated oracle $R$-value for all potential next-tokens, allowing DiNADO to be combined with finetuning methods like LoRA. We discuss in depth how DiNADO achieves better capacity, stability and flexibility with both empirical and theoretical results. Experiments on formality control in machine translation and the lexically constrained generation task CommonGen demonstrates the significance of the improvements.
Sidi Lu, Wenbo Zhao 0006, Chenyang Tao, Arpit Gupta, Shanchan Wu, Tagyoung Chung, Nanyun Peng 0001
ICML4
2024 Mitigating Bias for Question Answering Models by Tracking Bias Influence
abstract
Mingyu Ma, Jiun-Yu Kao, Arpit Gupta, Yu-Hsiang Lin, Wenbo Zhao, Tagyoung Chung, Wei Wang, Kai-Wei Chang, Nanyun Peng. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Mingyu Derek Ma, Jiun-Yu Kao, Arpit Gupta, Yu-Hsiang Lin, Wenbo Zhao 0006, Tagyoung Chung, Wei Wang 0010, Kai-Wei Chang 0001, Nanyun Peng 0001
NAACL-HLT3
2024 NetworkGym: Reinforcement Learning Environments for Multi-Access Traffic Management in Network Simulation
abstract
Mobile devices such as smartphones, laptops, and tablets can often connect to multiple access networks (e.g., Wi-Fi, LTE, and 5G) simultaneously.Recent advancements facilitate seamless integration of these connections below the transport layer, enhancing the experience for apps that lack inherent multi-path support.This optimization hinges on dynamically determining the traffic distribution across networks for each device, a process referred to as multi-access traffic splitting.This paper introduces NetworkGym, a high-fidelity network environment simulator that facilitates generating multiple network traffic flows and multi-access traffic splitting.This simulator facilitates training and evaluating different RL-based solutions for the multi-access traffic splitting problem.Our initial explorations demonstrate that the majority of existing state-of-the-art offline RL algorithms (e.g. CQL) fail to outperform certain hand-crafted heuristic policies on average.This illustrates the urgent need to evaluate offline RL algorithms against a broader range of benchmarks, rather than relying solely on popular ones such as D4RL.We also propose an extension to the TD3+BC algorithm, named Pessimistic TD3 (PTD3), and demonstrate that it outperforms many state-of-the-art offline RL algorithms.PTD3's behavioral constraint mechanism, which relies on value-function pessimism, is theoretically motivated and relatively simple to implement.We open source our code and offline datasets at github.com/hmomin/networkgym.
Momin Haider, Ming Yin 0003, Menglei Zhang, Arpit Gupta, Yu-Xiang Wang 0003
NeurIPS4
2024 Watching Stars in Pixels: The Interplay Of Traffic Shaping and YouTube Streaming QoE over GEO Satellite Networks
Jiamo Liu, David Lerner, Jae Chung, Udita Paul, Arpit Gupta, Elizabeth M. Belding
PAM (2)5
2024 The Efficacy of the Connect America Fund in Addressing US Internet Access Inequities
abstract
Residential fixed broadband internet access in the US remains inequitable, despite significant taxpayer investment. This paper evaluates the efficacy of the Connect America Fund (CAF), which subsidizes new broadband monopolies in underserved areas to provide internet access comparable to that in urban regions. CAF's oversight relies heavily on self-reported data from internet service providers (ISPs). Unfortunately, the reliability of this self-reported data has always been open to question. We use the broadband-plan querying tool (BQT) to create a novel dataset that complements ISP-reported information with ISP-advertised broadband plan details from publicly accessible websites for 537k residential addresses across 15 states. Our analysis reveals significant discrepancies, with a serviceability rate of only 55.45%, indicating that a significant fraction of addresses certified as served are still unserved. Furthermore, we observe a compliance rate of only 33.03%, indicating that a significant fraction of served addresses receive download speeds that are non-compliant with the FCC's 10 Mbps threshold for CAF-served addresses. Although we observe that CAF-served addresses occasionally receive higher download speeds than their monopoly-served neighbors, overall, the CAF program has largely failed to achieve its intended goal, leaving many targeted rural communities with inadequate or no broadband connectivity.
Haarika Manda, Varshika Srinivasavaradhan, Laasya Koduru, Xuanhe Zhou, Udita Paul, Elizabeth M. Belding, Arpit Gupta, Tejas N. Narechania
SIGCOMM8
2024 Leveraging Prefix Structure to Detect Volumetric DDoS Attack Signatures with Programmable Switches
abstract
As increasingly complex and dynamic volumetric DDoS attacks continue to wreak havoc on edge networks, two recent developments promise to bolster DDoS defense at the edge. First, programmable switches have emerged as promising means for achieving scalable and cost-effective attack signature detection. However, their practical application in edge networks remains a challenging open problem. Second, machine learning (ML)-based solutions have demonstrated potential in accurately detecting attack signatures based on per-flow traffic features. Yet, their inability to effectively scale to the traffic volumes and number of flows in actual production edge networks has largely excluded them from practical considerations.In this paper, we introduce ZAPDOS, a novel approach to accurately, quickly, and scalably detect volumetric DDoS attack signatures at the source prefix level. ZAPDOS is the first to utilize a key characteristic of the observed structure of measured attack and benign source prefixes (i.e., a pronounced cluster-within-cluster property) and effectively apply it in practice against modern attacks. ZAPDOS operates by monitoring aggregate prefix-level features in switch hardware, employing a learning model to identify prefixes suspected of containing attack sources, and using several innovative algorithmic methods to pinpoint attack sources efficiently. We have built a hardware prototype of ZAPDOS and a packet-level software simulator which achieve comparable accuracy results. Since existing datasets are inadequate for training and evaluating prefix-level models, we have developed a new data-fusion methodology for training and evaluating ZAPDOS. We use our prototype and simulator to show that ZAPDOS can detect volumetric DDoS attack signatures with orders of magnitude lower error rates than state-of-the-art under comparable monitoring resource budgets and for a range of different attack scenarios.
Chris Misa, Ramakrishnan Durairajan, Arpit Gupta, Reza Rejaie, Walter Willinger
SP3
2023 In Search of netUnicorn: A Data-Collection Platform to Develop Generalizable ML Models for Network Security Problems
abstract
The remarkable success of the use of machine learning-based solutions for network security problems has been impeded by the developed ML models' inability to maintain efficacy when used in different network environments exhibiting different network behaviors. This issue is commonly referred to as the generalizability problem of ML models. The community has recognized the critical role that training datasets play in this context and has developed various techniques to improve dataset curation to overcome this problem. Unfortunately, these methods are generally ill-suited or even counterproductive in the network security domain, where they often result in unrealistic or poor-quality datasets.
Roman Beltiukov, Wenbo Guo 0002, Arpit Gupta, Walter Willinger
CCS3
2023 Estimating WebRTC Video QoE Metrics Without Using Application Headers
abstract
The increased use of video conferencing applications (VCAs) has made it critical to understand and support end-user quality of experience (QoE) by all stakeholders in the VCA ecosystem, especially network operators, who typically do not have direct access to client software. Existing VCA QoE estimation methods use passive measurements of application-level Real-time Transport Protocol (RTP) headers. However, a network operator does not always have access to RTP headers in all cases, particularly when VCAs use custom RTP protocols (e.g., Zoom) or due to system constraints (e.g., legacy measurement systems). Given this challenge, this paper considers the use of more standard features in the network traffic, namely, IP and UDP headers, to provide per-second estimates of key VCA QoE metrics such as frames rate and video resolution. We develop a method that uses machine learning with a combination of flow statistics (e.g., throughput) and features derived based on the mechanisms used by the VCAs to fragment video frames into packets. We evaluate our method for three prevalent VCAs running over WebRTC: Google Meet, Microsoft Teams, and Cisco Webex. Our evaluation consists of 54,696 seconds of VCA data collected from both (1), controlled in-lab network conditions, and (2) real-world networks from 15 households. We show that the ML-based approach yields similar accuracy compared to the RTP-based methods, despite using only IP/UDP data. For instance, we can estimate FPS within 2 FPS for up to 83.05% of one-second intervals in the real-world data, which is only 1.76% lower than using the application-level RTP headers.
Taveesh Sharma, Tarun Mangla, Arpit Gupta, Junchen Jiang, Nick Feamster
IMC3
2023 Parameter-Efficient Low-Resource Dialogue State Tracking by Prompt Tuning
Mingyu Derek Ma, Jiun-Yu Kao, Shuyang Gao, Arpit Gupta, Di Jin 0005, Tagyoung Chung, Nanyun Peng 0001
INTERSPEECH4
2023 Poster: Traffic Shaping and YouTube Performance Interaction in GEO Satellite Networks
abstract
Geosynchronous satellite (GEO) networks are a crucial option for users beyond terrestrial connectivity. However, unlike terrestrial networks, GEO networks exhibit high latency and deploy TCP proxies and traffic shapers. The deployment of proxies mitigates the impact of high network latency, while traffic shapers help realize customer-controlled data-saver options that optimize data usage. It is unclear how the interplay between GEO networks' high latency, TCP proxies, and traffic-shaping policies affects the quality of experience (QoE) for commonly used video applications. In our study, we examine this relationship through a series of video streaming experiments at a shaped rate of 900kbps. Our preliminary analysis reveals that 28% of TCP sessions (with TCP proxies) and 18% of gQUIC sessions (without TCP proxies) experience rebuffering events, while the median average resolution is only 380p for TCP and 299p for gQUIC. Additionally, we identify two key factors contributing to sub-optimal performance: (i) unlike TCP, gQUIC only utilizes 63% of network capacity; and (ii) YouTube's chunk request pipelining is imperfect. To avoid potential degradation in video quality, the satellite provider subsequently discontinued providing data saver options that shape video traffic to US residential customers.
Jiamo Liu, David Lerner, Jae Chung, Udita Paul, Arpit Gupta, Elizabeth M. Belding
SIGCOMM5
2023 Decoding the Divide: Analyzing Disparities in Broadband Plans Offered by Major US ISPs
abstract
Digital equity in Internet access is often measured along three axes: availability, affordability, and adoption. Most prior work focuses on availability; the other two aspects have received less attention. In this paper, we study broadband affordability in the US by focusing on the nature of broadband plans offered by major ISPs. To this end, we develop a broadband plan querying tool (BQT) that obtains broadband plans (upload/download speed and price) offered by seven major wireline US ISPs for any street address in the US. We then use this tool to curate a dataset, querying broadband plans for over 837 k street addresses in thirty cities for these ISPs. We use a plan's carriage value, defined as the Mbps of a user's traffic that an ISP carries for one dollar, to compare plans. Our analysis provides us with the following new insights: (1) ISP plans vary inter-city. Specifically, up to 60% of the census block groups in a city can receive low carriage value plans from an ISP; (2) ISP plans intra-city are spatially clustered, and the carriage value can vary as much as 600% within a city; (3) Cable-based ISPs offer up to 30% higher carriage value to users when they are competing with fiber-based ISPs in a block group compared to when they are operating alone or in conjunction with a DSL-based ISP; and (4) Fiber deployments, which have better carriage values, are associated with higher average income block groups. While we hope our tool, dataset, and analysis in their current form are helpful for policymakers at different levels (city, county, state), they are only a small step toward quantifying digital inequity. We conclude with recommendations to further advance our understanding of broadband affordability.
Udita Paul, Vinothini Gunasekaran, Jiamo Liu, Tejas N. Narechania, Arpit Gupta, Elizabeth M. Belding
SIGCOMM5
2023 Panakos: Chasing the Tails for Multidimensional Data Streams
abstract
System operators are often interested in extracting different feature streams from multi-dimensional data streams; and reporting their distributions at regular intervals, including the heavy hitters that contribute to the tail portion of the feature distribution. Satisfying these requirements to increase data rates with limited resources is challenging. This paper presents the design and implementation of Panakos that makes the best use of available resources to report a given feature's distribution accurately, its tail contributors, and other stream statistics (e.g., cardinality, entropy, etc.). Our key idea is to leverage the skewness inherent to most feature streams in the real world. We leverage this skewness by disentangling the feature stream into hot, warm, and cold items based on their feature values. We then use different data structures for tracking objects in each category. Panakos provides solid theoretical guarantees and achieves high performance for various tasks. We have implemented Panakos on both software and hardware and compared Panakos to other state-of-the-art sketches using synthetic and real-world datasets. The experimental results demonstrate that Panakos often achieves one order of magnitude better accuracy than the state-of-the-art solutions for a given memory budget.
Fuheng Zhao, Punnal Ismail Khan, Divyakant Agrawal, Amr El Abbadi, Arpit Gupta, Zaoxing Liu
Proc. VLDB Endow.5
2022 AI/ML for Network Security: The Emperor has no Clothes
abstract
Several recent research efforts have proposed Machine Learning (ML)-based solutions that can detect complex patterns in network traffic for a wide range of network security problems. However, without understanding how these black-box models are making their decisions, network operators are reluctant to trust and deploy them in their production settings. One key reason for this reluctance is that these models are prone to the problem of underspecification, defined here as the failure to specify a model in adequate detail. Not unique to the network security domain, this problem manifests itself in ML models that exhibit unexpectedly poor behavior when deployed in real-world settings and has prompted growing interest in developing interpretable ML solutions (e.g., decision trees) for "explaining'' to humans how a given black-box model makes its decisions. However, synthesizing such explainable models that capture a given black-box model's decisions with high fidelity while also being practical (i.e., small enough in size for humans to comprehend) is challenging.
Arthur Selle Jacobs, Roman Beltiukov, Walter Willinger, Ronaldo A. Ferreira, Arpit Gupta, Lisandro Z. Granville
CCS5
2022 GRAVL-BERT: Graphical Visual-Linguistic Representations for Multimodal Coreference Resolution
abstract
Learning from multimodal data has become a popular research topic in recent years. Multimodal coreference resolution (MCR) is an important task in this area. MCR involves resolving the references across different modalities, e.g., text and images, which is a crucial capability for building next-generation conversational agents. MCR is challenging as it requires encoding information from different modalities and modeling associations between them. Although significant progress has been made for visual-linguistic tasks such as visual grounding, most of the current works involve single turn utterances and focus on simple coreference resolutions. In this work, we propose an MCR model that resolves coreferences made in multi-turn dialogues with scene images. We present GRAVL-BERT, a unified MCR framework which combines visual relationships between objects, background scenes, dialogue, and metadata by integrating Graph Neural Networks with VL-BERT. We present results on the SIMMC 2.0 multimodal conversational dataset, achieving the rank-1 on the DSTC-10 SIMMC 2.0 MCR challenge with F1 score 0.783. Our code is available at https://github.com/alexa/gravl-bert.
Danfeng Guo, Arpit Gupta, Sanchit Agarwal, Jiun-Yu Kao, Shuyang Gao, Arijit Biswas, Chien-Wei Lin, Tagyoung Chung, Mohit Bansal
COLING2
2022 Characterizing Internet Access and Quality Inequities in California M-Lab Measurements
abstract
It is well documented that, in the United States (U.S.), the availability of Internet access is related to several demographic attributes. Data collected through end user network diagnostic tools, such as the one provided by the Measurement Lab (M-Lab) Speed Test, allows the extension of prior work by exploring the relationship between the quality, as opposed to only the availability, of Internet access and demographic attributes of users of the platform. In this study, we use network measurements collected from the users of Speed Test by M-Lab and demographic data to characterize the relationship between the quality-of-service (QoS) metric download speed, and various critical demographic attributes, such as income, education level, and poverty. For brevity, we limit our focus to the state of California. For users of the M-Lab Speed Test, our study has the following key takeaways: (1) geographic type (urban/rural) and income level in an area have the most significant relationship to download speed; (2) average download speed in rural areas is 2.5 times lower than urban areas; (3) the COVID-19 pandemic had a varied impact on download speeds for different demographic attributes; and (4) the U.S. Federal Communication Commission’s (FCC’s) broadband speed data significantly over-represents the download speed for rural and low-income communities compared to what is recorded through Speed Test.
Udita Paul, Jiamo Liu, David Farias-llerenas, Vivek Adarsh, Arpit Gupta, Elizabeth M. Belding
COMPASS5
2022 The importance of contextualization of crowdsourced active speed test measurements
abstract
Crowdsourced speed test measurements, such as those by Ookla® and Measurement Lab (M-Lab), offer a critical view of network access and performance from the user's perspective. However, we argue that taking these measurements at surface value is problematic. It is essential to contextualize these measurements to understand better what the attained upload and download speeds truly measure. To this end, we develop a novel Broadband Subscription Tier (BST) methodology that associates a speed test data point with a residential broadband subscription plan. Our evaluation of this methodology with the FCC's MBA dataset shows over 96% accuracy. We augment approximately 1.5M Ookla and M-Lab speed test measurements from four major U.S. cities with the BST methodology. We show that many low-speed data points are attributable to lower-tier subscriptions and not necessarily poor access. Then, for a subset of the measurement sample (80k data points), we quantify the impact of access link type (WiFi or wired), WiFi spectrum band and RSSI (if applicable), and device memory on speed test performance. Interestingly, we observe that measurement time of day only marginally affects the reported speeds. Finally, we show that the median throughput reported by Ookla speed tests can be up to two times greater than M-Lab measurements for the same subscription tier, city, and ISP due to M-Lab's employment of different measurement methodologies. Based on our results, we put forward a set of recommendations for both speed test vendors and the FCC to con-textualize speed test data points and correctly interpret measured performance.
Udita Paul, Jiamo Liu, Mengyang Gu, Arpit Gupta, Elizabeth M. Belding
IMC4
2022 Detecting Ephemeral Optical Events with OpTel
Congcong Miao, Minggang Chen, Arpit Gupta, Zili Meng, Lianjin Ye, Jingyu Xiao, Zekun He, Xulong Luo, Jilong Wang 0001, Heng Yu 0005
NSDI3
2021 AutIS: Artificial Intelligent Based Automated Interviewing System
Rupesh Kumar Dewang, Arpit Gupta, Anisha Kumari, Ritik Raj, Raj Nath Shah, Tanmay Jaiswal, Arvind Mewada
HIS2
2021 Coverage is Not Binary: Quantifying Mobile Broadband Quality in Urban, Rural, and Tribal Contexts
abstract
Cellular network performance does not cleanly generalize. A variety of factors, such as location, terrain, signal quality and network load, affect the performance of services delivered over LTE networks. As a result, the presence of LTE coverage does not always equate to usable service; coverage can be of poor quality, or it can be congested and difficult to access. Given that reliance on LTE networks for Internet connectivity has exploded, it is critical to understand the quality of experience for applications delivered over these networks in a variety of scenarios. To this end, we develop a robust measurement suite that we use to conduct a unique measurement campaign in tribal, rural, congested urban and uncongested urban regions, representing a variety of under-provisioned, congested, and well-provisioned operational LTE networks run by four major providers. Our analysis confirms that the performance of LTE networks in tribal and rural areas is typically worse than even heavily congested urban networks. More specifically, in the regions that we study, LTE networks in under-provisioned (tribal/rural) areas have $ 9\times$ poorer video streaming quality, $ 10\times$ higher video start-up delay, undergo more than $ 10\times$ the number of resolution switches, and lead to more than $ 2\times$ slower Web browsing experience as compared to urban deployments. We show that throughput and latency are $ 11\times$ and $ 3\times$ worse in tribal and rural locations, despite identical LTE carrier subscription plans.
Vivek Adarsh, Michael Nekrasov, Udita Paul, Tarun Mangla, Arpit Gupta, Morgan Vigil-Hayes, Ellen Zegura, Elizabeth M. Belding
ICCCN5
2021 Continuous Flow Measurement with SuperFlow
abstract
Flow-based network measurement enables operators to perform a wide range of network management tasks in a scalable manner. Recently, various algorithms have been proposed for flow record collection at very high speed. However, they all focus on processing traffic in a short time window, but overlook the fact that flow measurements are typically needed continuously for unlimited time. To this end, we propose a new algorithm named SuperFlow to support continuous and accurate flow record collection at very high speed by monitoring the flow activeness and exporting the inactive records from the data plane automatically. Our data structures and the corresponding algorithms are carefully designed and analyzed, so the above goal is achieved with limited memory and bandwidth consumption. We implement SuperFlow on both x86 CPU and state-of-the-art PISA target. Comprehensive experiments show that SuperFlow consistently outperforms its competitors significantly. Especially, compared with the best competitor, it records around 136.7% more flows, reduces the error in flow size estimation by 51.5%, and reduces the memory or bandwidth consumption by up to 71.0%, while bringing only negligible throughput degradation.
Zongyi Zhao, Xingang Shi, Arpit Gupta, Qing Li 0006, Bin Xiong, Xia Yin 0001
IWQoS3
2021 Too Late for Playback: Estimation of Video Stream Quality in Rural and Urban Contexts
Vivek Adarsh, Michael Nekrasov, Udita Paul, Alexander Ermakov, Arpit Gupta, Morgan Vigil-Hayes, Ellen Zegura, Elizabeth M. Belding
PAM5
2020 (How Much) Does a Private WAN Improve Cloud Performance?
abstract
The construction of private WANs by cloud providers enables them to extend their networks to more locations and establish direct connectivity with end user ISPs. Tenants of the cloud providers benefit from this proximity to users, which is supposed to provide improved performance by bypassing the public Internet. However, the performance impact of cloud providers' private WANs is not widely understood.To isolate the impact of a private WAN, we measure from globally distributed vantage points to two large cloud providers, comparing performance when using their worldwide WAN and when instead using the public Internet. The benefits are not universal. While 48% of our vantage points saw improved performance when using the WAN, 43% had statistically indistinguishable median performance, and 9% had better performance over the public Internet. We find that the benefits of the private WAN tend to improve with client-to-server distance, but the benefits (or drawbacks) for a particular vantage point depend on specifics of its geographic and network connectivity.
Todd Arnold, Ege Gürmeriçliler, Georgia Essig, Arpit Gupta, Matt Calder, Vasileios Giotsas, Ethan Katz-Bassett
INFOCOM4
2019 Beating BGP is Harder than we Thought
abstract
Online services all seek to provide their customers with the best Quality of Experience (QoE) possible. Milliseconds of delay can cause users to abandon a cat video or move onto a different shopping site, which translates into lost revenue. Thus, minimizing latency between users and content is crucial. To reduce latency, content and cloud providers have built massive, global networks. However, their networks must interact with customer ISPs via BGP, which has no concept of performance.
Todd Arnold, Matt Calder, Ítalo S. Cunha, Arpit Gupta, Harsha V. Madhyastha, Michael Schapira, Ethan Katz-Bassett
HotNets4
2019 An Effort to Democratize Networking Research in the Era of AI/ML
abstract
A growing concern within today's networking community is that with the proliferation of Artificial Intelligence/Machine Learning (AI/ML) techniques, a lack of access to real-world production networks is putting academic researchers at a significant disadvantage. Indeed, compared to a select few research groups in industry that can leverage access to their global-scale production networks in their data-driven efforts to develop and evaluate learning models, academic researchers not only struggle to get their hands on real-world data sets but find it almost impossible to adequately train and assess their learning models under realistic conditions.
Arpit Gupta, Chris Mac-Stoker, Walter Willinger
HotNets1
2018 Preserving Privacy at IXPs
abstract
Autonomous systems (ASes) on the Internet increasingly rely on Internet Exchange Points (IXPs) for peering. A single IXP may interconnect several 100s or 1000s of participants (ASes) all of which might peer with each other through BGP sessions. IXPs have addressed this scaling challenge through the use of route servers. However, route servers require participants to trust the IXP and reveal their policies, a drastic change from the accepted norm where all policies are kept private. In this paper we look at techniques to build route servers which provide the same functionality as existing route servers without requiring participants to reveal their policies thus preserving the status quo and enabling wider adoption of IXPs. Prior work has looked at secure multiparty computation (SMPC) as a means of implementing such route servers however this affects performance and reduces policy flexibility. In this paper we take a different tack and build on trusted execution environments (TEEs) such as Intel SGX to keep policies private and flexible. We present results from an initial route server implementation that runs under Intel SGX and show that our approach has 20x better performance than SMPC based approaches. Furthermore, we demonstrate that the additional privacy provided by our approach comes at minimal cost and our implementation is at worse 2.1x slower than a current route server implementation (and in some situations up to 2x faster).
Xiaohe Hu, Arpit Gupta, Nick Feamster, Aurojit Panda, Scott Shenker
APNet2
2018 Contextual Slot Carryover for Disparate Schemas
abstract
In the slot-filling paradigm, where a user can refer back to slots in the context during a conversation, the goal of the contextual understanding system is to resolve the referring expressions to the appropriate slots in the context. In large-scale multi-domain systems, this presents two challenges - scaling to a very large and potentially unbounded set of slot values, and dealing with diverse schemas. We present a neural network architecture that addresses the slot value scalability challenge by reformulating the contextual interpretation as a decision to carryover a slot from a set of possible candidates. To deal with heterogenous schemas, we introduce a simple data-driven method for trans- forming the candidate slots. Our experiments show that our approach can scale to multiple domains and provides competitive results over a strong baseline.
Chetan Naik, Arpit Gupta, Hancheng Ge, Lambert Mathias, Ruhi Sarikaya
INTERSPEECH2
2018 Sonata: query-driven streaming network telemetry
abstract
Managing and securing networks requires collecting and analyzing network traffic data in real time. Existing telemetry systems do not allow operators to express the range of queries needed to perform management or scale to large traffic volumes and rates. We present Sonata, an expressive and scalable telemetry system that coordinates joint collection and analysis of network traffic. Sonata provides a declarative interface to express queries for a wide range of common telemetry tasks; to enable real-time execution, Sonata partitions each query across the stream processor and the data plane, running as much of the query as it can on the network switch, at line rate. To optimize the use of limited switch memory, Sonata dynamically refines each query to ensure that available resources focus only on traffic that satisfies the query. Our evaluation shows that Sonata can support a wide range of telemetry tasks while reducing the workload for the stream processor by as much as seven orders of magnitude compared to existing telemetry systems.
Arpit Gupta, Rob Harrison, Marco Canini, Nick Feamster, Jennifer Rexford, Walter Willinger
SIGCOMM1
2016 Network Monitoring as a Streaming Analytics Problem
abstract
Programmable switches potentially make it easier to perform flexible network monitoring queries at line rate, and scalable stream processors make it possible to fuse data streams to answer more sophisticated queries about the network in real-time. However, processing such network monitoring queries at high traffic rates requires both the switches and the stream processors to filter the traffic iteratively and adaptively so as to extract only that traffic that is of interest to the query at hand. While the realization that network monitoring is a streaming analytics problem has been made earlier, our main contribution in this paper is the design and implementation of Sonata, a closed-loop system that enables network operators to perform streaming analytics for network monitoring applications at scale. To achieve this objective, Sonata allows operators to express a network monitoring query by considering each packet as a tuple. More importantly, Sonata allows them to partition the query across both the switches and the stream processor, and through iterative refinement, Sonata's runtime attempts to extract only the traffic that pertains to the query, thus ensuring that the stream processor can scale to satisfy a large number of queries for traffic at very high rates. We show with a simple example query involving DNS reflection attacks and traffic traces from one of the world's largest IXPs that Sonata can capture 95% of all traffic pertaining to the query, while reducing the overall data rate by a factor of about 400 and the number of required counters by four orders of magnitude.
Arpit Gupta, Rüdiger Birkner, Marco Canini, Nick Feamster, Chris Mac-Stoker, Walter Willinger
HotNets1
2016 An Industrial-Scale Software Defined Internet Exchange Point
Arpit Gupta, Robert MacDavid, Rüdiger Birkner, Marco Canini, Nick Feamster, Jennifer Rexford, Laurent Vanbever
NSDI1
2016 An Industrial-Scale Software Defined Internet Exchange Point
Arpit Gupta, Robert MacDavid, Rüdiger Birkner, Marco Canini, Nick Feamster, Jennifer Rexford, Laurent Vanbever
USENIX ATC1
2015 Kinetic: Verifiable Dynamic Network Control
Hyojoon Kim, Joshua Reich, Arpit Gupta, Muhammad Shahbaz 0001, Nick Feamster, Russell J. Clark 0001
NSDI3
2014 Peering at the Internet's Frontier: A First Look at ISP Interconnectivity in Africa
Arpit Gupta, Matt Calder, Nick Feamster, Marshini Chetty, Enrico Calandro, Ethan Katz-Bassett
PAM1
2014 SDX: a software defined internet exchange
abstract
BGP severely constrains how networks can deliver traffic over the Internet. Today's networks can only forward traffic based on the destination IP prefix, by selecting among routes offered by their immediate neighbors. We believe Software Defined Networking (SDN) could revolutionize wide-area traffic delivery, by offering direct control over packet-processing rules that match on multiple header fields and perform a variety of actions. Internet exchange points (IXPs) are a compelling place to start, given their central role in interconnecting many networks and their growing importance in bringing popular content closer to end users.
Arpit Gupta, Laurent Vanbever, Muhammad Shahbaz 0001, Sean Patrick Donovan, Brandon Schlinker, Nick Feamster, Jennifer Rexford, Scott Shenker, Russell J. Clark 0001, Ethan Katz-Bassett
SIGCOMM1
2014 SDX: a software defined internet exchange
abstract
BGP severely constrains how networks can deliver traffic over the Internet. Today's networks can only forward traffic based on the destination IP prefix, by selecting among routes offered by their immediate neighbors. We believe Software Defined Networking (SDN) could revolutionize wide-area traffic delivery, by offering direct control over packet-processing rules that match on multiple header fields and perform a variety of actions. Internet exchange points (IXPs) are a compelling place to start, given their central role in interconnecting many networks and their growing importance in bringing popular content closer to end users. To realize a Software Defined IXP (an "SDX"), we need new programming abstractions that allow participating networks to create and run these applications and a runtime that both behaves correctly when interacting with BGP and ensures that applications do not interfere with each other. We must also ensure that the system scales, both in rule-table size and computational overhead. In this demo, we show how we tackle these challenges demonstrating the flexibility and scalability of our SDX platform. The paper also appears in the main program.
Arpit Gupta, Laurent Vanbever, Muhammad Shahbaz 0001, Sean Patrick Donovan, Brandon Schlinker, Nick Feamster, Jennifer Rexford, Scott Shenker, Russell J. Clark 0001, Ethan Katz-Bassett
SIGCOMM1
2012 WiFox: scaling WiFi performance for large audience environments
abstract
WiFi-based wireless LANs (WLANs) are widely used for Internet access. They were designed such that an Access Points (AP) serves few associated clients with symmetric uplink/downlink traffic patterns. Usage of WiFi hotspots in locations such as airports and large conventions frequently experience poor performance in terms of downlink goodput and responsiveness. We study the various factors responsible for this performance degradation. We analyse and emulate a large conference network environment on our testbed with 45 nodes. We find that presence of asymmetry between the uplink/downlink traffic results in backlogged packets at WiFi Access Point's (AP's) transmission queue and subsequent packet losses. This traffic asymmetry results in maximum performance loss for such an environment along with degradation due to rate diversity, fairness and TCP behaviour. We propose our solution WiFox, which (1) adaptively prioritizes AP's channel access over competing STAs avoiding traffic asymmetry (2) provides a fairness framework alleviating the problem of performance loss due to rate-diversity/fairness and (3) avoids degradation due to TCP behaviour. We demonstrate that WiFox not only improves downlink goodput by 400-700 % but also reduces request's average response time by 30-40 %.
Arpit Gupta, Jeongki Min, Injong Rhee
CoNEXT1