VLDB 2026 Research / reviewers in the wild / expert
Anand Padmanabha Iyer
dblp:72/309 · also Anand P. Iyer
· DBLP profile ↗
21ranked-venue papers
9as first author
8since 2021 · last 2025
0000-0002-5952-3346ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 5 first-author · 3 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Flex: Fast, Accurate DNN Inference on Low-Cost Edges Using Heterogeneous Accelerator ExecutionabstractSignificant b reakthroughs in machine learning (ML) and the advantages of on-device processing have led to edge devices increasingly incorporating accelerators like GPUs, NPUs, and DSPs. However, these accelerators consume energy, prompting users to limit their floating-point precision. Many edge device users are in regions where including high-fidelity accelerators is too costly, leading to low-cost devices with low precision, sacrificing accuracy. Previous work predetermined layer assignments between the CPU and accelerator offline for high accuracy and low latency without considering the input, but we observe that input affects optimal layer assignment. To address this, we present Flex, a system for Fast, Accurate DNN Inference on Low-Cost Edges using Heterogeneous Accelerator eXecution. Leveraging common observations from models on various edge devices, Flex uses a lightweight heuristic and reinforcement learning (RL) to dynamically assign layers across the CPU and accelerator. Experiments show Flex improves average inference time by up to 39%, accuracy by up to 22%, and energy consumption by up to 61% compared to state-of-the-art methods, and is only 4.2% less optimal than the best achievable results. Tanmoy Sen, Haiying Shen, Anand Padmanabha Iyer |
EuroSys | 3 |
| 2024 | Vulcan: Automatic Query Planning for Live ML Analytics
Yiwen Zhang 0008, Xumiao Zhang, Ganesh Ananthanarayanan, Anand Padmanabha Iyer, Yuanchao Shu, Paramvir Bahl, Z. Morley Mao, Mosharaf Chowdhury |
NSDI | 4 |
| 2024 | USHER: Holistic Interference Avoidance for Resource Optimized ML Inference
Sudipta Saha Shubha, Haiying Shen, Anand Padmanabha Iyer |
OSDI | 3 |
| 2024 | Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML ServingabstractMachine learning (ML) inference platforms are tasked with balancing two competing goals: ensuring high throughput given many requests, and delivering low-latency responses to support interactive applications. Unfortunately, existing platform knobs (e.g., batch sizes) fail to ease this fundamental tension, and instead only enable users to harshly trade off one property for the other. This paper explores an alternate strategy to taming throughput-latency tradeoffs by changing the granularity at which inference is performed. We present Apparate, a system that automatically applies and manages early exits (EEs) in ML models, whereby certain inputs can exit with results at intermediate layers. To cope with the time-varying overhead and accuracy challenges that EEs bring, Apparate repurposes exits to provide continual feedback that powers several novel runtime monitoring and adaptation strategies. Apparate lowers median response latencies by 40.5--91.5% and 10.0--24.2% for diverse CV and NLP classification workloads, and median time-per-token latencies by 22.6--77.9% for generative scenarios, without affecting throughputs or violating tight accuracy constraints. Yinwei Dai, Rui Pan 0003, Anand Padmanabha Iyer, Kai Li 0001, Ravi Netravali |
SOSP | 3 |
| 2024 | Improving DNN Inference Throughput Using Practical, Per-Input Compute AdaptationabstractMachine learning inference platforms continue to face high request rates and strict latency constraints. Existing solutions largely focus on compressing models to substantially lower compute costs (and time) with mild accuracy degradations. This paper explores an alternate (but complementary) technique that trades off accuracy and resource costs on a perinput granularity: early exit models, which selectively allow certain inputs to exit a model from an intermediate layer. Though intuitive, early exits face fundamental deployment challenges, largely owing to the effects that exiting inputs have on batch size (and resource utilization) throughout model execution. We present E3, the first system that makes early exit models practical for realistic inference deployments. Our key insight is to split and replicate blocks of layers in models in a manner that maintains a constant batch size throughout execution, all the while accounting for resource requirements and communication overheads. Evaluations with NLP and vision models show that E3 can deliver up to 1.74× improvement in goodput (for a fixed cost) or 1.78× reduction in cost (for a fixed goodput). Additionally, E3's goodput wins generalize to autoregressive LLMs (2.8--3.8×) and compressed models (1.67×). Anand Padmanabha Iyer, Mingyu Guan, Yinwei Dai, Rui Pan 0003, Swapnil Gandhi, Ravi Netravali |
SOSP | 1 |
| 2023 | Gemel: Model Merging for Memory-Efficient, Real-Time Video Analytics at the Edge
Arthi Padmanabhan, Neil Agarwal, Anand Padmanabha Iyer, Ganesh Ananthanarayanan, Yuanchao Shu, Nikolaos Karianakis, Guoqing Harry Xu, Ravi Netravali |
NSDI | 3 |
| 2021 | TEGRA: Efficient Ad-Hoc Analytics on Evolving Graphs
Anand Padmanabha Iyer, Qifan Pu, Kishan Patel, Joseph Gonzalez 0001, Ion Stoica |
NSDI | 1 |
| 2021 | P3: Distributed Deep Graph Learning at Scale
Swapnil Gandhi, Anand Padmanabha Iyer |
OSDI | 2 |
| 2018 | Mitigating the Latency-Accuracy Trade-off in Mobile Data Analytics SystemsabstractAn increasing amount of mobile analytics is performed on data that is procured in a real-time fashion to make real-time decisions. Such tasks include simple reporting on streams to sophisticated model building. However, the practicality of these analyses are impeded in several domains because they are faced with a fundamental trade-off between data collection latency and analysis accuracy. In this paper, we first study this trade-off in the context of a specific domain, Cellular Radio Access Networks (RAN). We find that the trade-off can be resolved using two broad, general techniques: intelligent data grouping and task formulations that leverage domain characteristics. Based on this, we present CellScope, a system that applies a domain specific formulation and application of Multi-task Learning (MTL) to RAN performance analysis. It uses three techniques: feature engineering to transform raw data into effective features, a PCA inspired similarity metric to group data from geographically nearby base stations sharing performance commonalities, and a hybrid online-offline model for efficient model updates. Our evaluation shows that CellScope's accuracy improvements over direct application of ML range from 2.5× to 4.4× while reducing the model update overhead by up to 4.8×. We have also used CellScope to analyze an LTE network of over 2 million subscribers, where it reduced troubleshooting efforts by several magnitudes. We then apply the underlying techniques in CellScope to another domain specific problem, mobile phone energy bug diagnosis, and show that the techniques are general. Anand Padmanabha Iyer, Li Erran Li, Mosharaf Chowdhury, Ion Stoica |
MobiCom | 1 |
| 2018 | ASAP: Fast, Approximate Graph Pattern Mining at Scale
Anand Padmanabha Iyer, Zaoxing Liu, Xin Jin 0008, Shivaram Venkataraman, Vladimir Braverman, Ion Stoica |
OSDI | 1 |
| 2017 | A scalable distributed spatial index for the internet-of-thingsabstractThe increasing interest in the Internet-of-Things (IoT) suggests that a new source of big data is imminent---the machines and sensors in the IoT ecosystem. The fundamental characteristic of the data produced by these sources is that they are inherently geospatial in nature. In addition, they exhibit unprecedented and unpredictable skews. Thus, big data systems designed for IoT applications must be able to efficiently ingest, index and query spatial data having heavy and unpredictable skews. Spatial indexing is well explored area of research in literature, but little attention has been given to the topic of efficient distributed spatial indexing. Anand Padmanabha Iyer, Ion Stoica |
SoCC | 1 |
| 2017 | Automating Diagnosis of Cellular Radio Access Network ProblemsabstractIn an increasingly mobile connected world, our user experience of mobile applications more and more depends on the performance of cellular radio access networks (RAN). To achieve high quality of experience for the user, it is imperative that operators identify and diagnose performance problems quickly. In this paper, we describe our experience in understanding the challenges in automating the diagnosis of RAN performance problems. Working with a major cellular network operator on a part of their RAN that services more than 2 million users, we demonstrate that fine-grained modeling and analysis could be the key towards this goal. We describe our methodology in analyzing RAN problems, and highlight a few of our findings, some previously unknown. We also discuss lessons from our attempt at building automated diagnosis solutions. Anand Padmanabha Iyer, Li Erran Li, Ion Stoica |
MobiCom | 1 |
| 2015 | FastLane: making short flows shorter with agile drop notificationabstractThe drive towards richer and more interactive web content places increasingly stringent requirements on datacenter network performance. Applications running atop these networks typically partition an incoming query into multiple subqueries, and generate the final result by aggregating the responses for these subqueries. As a result, a large fraction --- as high as 80% --- of the network flows in such workloads are short and latency-sensitive. The speed with which existing networks respond to packet drops limits their ability to meet high-percentile flow completion time SLOs. Indirect notifications indicating packet drops (e.g., duplicates in an end-to-end acknowledgement sequence) are an important limitation to the agility of response to packet drops. David Zats, Anand Padmanabha Iyer, Ganesh Ananthanarayanan, Rachit Agarwal 0001, Randy H. Katz, Ion Stoica, Amin Vahdat |
SoCC | 2 |
| 2015 | CellIQ : Real-Time Cellular Network Analytics at Scale
Anand Padmanabha Iyer, Li Erran Li, Ion Stoica |
NSDI | 1 |
| 2013 | Carat: collaborative energy diagnosis for mobile devicesabstractWe aim to detect and diagnose energy anomalies, abnormally heavy battery use. This paper describes a collaborative black-box method, and an implementation called Carat, for diagnosing anomalies on mobile devices. A client app sends intermittent, coarse-grained measurements to a server, which correlates higher expected energy use with client properties like the running apps, device model, and operating system. The analysis quantifies the error and confidence associated with a diagnosis, suggests actions the user could take to improve battery life, and projects the amount of improvement. During a deployment to a community of more than 500,000 devices, Carat diagnosed thousands of energy anomalies in the wild. Carat detected all synthetically injected anomalies, produced no known instances of false positives, projected the battery impact of anomalies with 95% accuracy, and, on average, increased a user's battery life by 11% after 10 days (compared with 1.9% for the control group). Adam J. Oliner, Anand Padmanabha Iyer, Ion Stoica, Eemil Lagerspetz, Sasu Tarkoma |
SenSys | 2 |
| 2012 | Blink and It's Done: Interactive Queries on Very Large DataabstractIn this demonstration, we present BlinkDB, a massively parallel, sampling-based approximate query processing framework for running interactive queries on large volumes of data. The key observation in BlinkDB is that one can make reasonable decisions in the absence of perfect answers. BlinkDB extends the Hive/HDFS stack and can handle the same set of SPJA (selection, projection, join and aggregate) queries as supported by these systems. BlinkDB provides real-time answers along with statistical error guarantees, and can scale to petabytes of data and thousands of machines in a fault-tolerant manner. Our experiments using the TPC-H benchmark and on an anonymized real-world video content distribution workload from Conviva Inc. show that BlinkDB can execute a wide range of queries up to 150x faster than Hive on MapReduce and 10--150x faster than Shark (Hive on Spark) over tens of terabytes of data stored across 100 machines, all with an error of 2--10%. Sameer Agarwal 0002, Aurojit Panda, Barzan Mozafari, Anand Padmanabha Iyer, Samuel Madden 0001, Ion Stoica |
Proc. VLDB Endow. | 4 |
| 2010 | Cyclostationary-Based Architectures for Spectrum Sensing in IEEE 802.22 WRANabstractThe well known noise rejection property of the cyclostationary spectrum makes it an ideal candidate for spectrum sensing in low SNR environments such as the IEEE 802.22 WRAN, which stipulates detection of primary signals at -20.8dB. In this paper, we propose two novel detector architectures that exploit cyclostationary properties: the Spectral Correlation Density (SCD), and the Magnitude Squared Coherence (MSC). Through extensive simulations, both on generated data and real world ATSC capture data, we show that our detector achieves an improvement of 2.5dB compared to existing proposals. Additionally, we compare our proposal against two popular choices for spectrum sensing in cognitive radio - the matched filter detection and energy detection, and show the superiority of cyclostationary spectrum sensing in such low SNR environments. Deepa Bhargavi, Anand Padmanabha Iyer, Chandra R. Murthy |
GLOBECOM | 2 |
| 2010 | Indoor localization without the painabstractWhile WiFi-based indoor localization is attractive, the need for a significant degree of pre-deployment effort is a key challenge. In this paper, we ask the question: can we perform indoor localization with no pre-deployment effort? Our setting is an indoor space, such as an office building or a mall, with WiFi coverage but where we do not assume knowledge of the physical layout, including the placement of the APs. Users carrying WiFi-enabled devices such as smartphones traverse this space in normal course. The mobile devices record Received Signal Strength (RSS) measurements corresponding to APs in their view at various (unknown) locations and report these to a localization server. Occasionally, a mobile device will also obtain and report a location fix, say by obtaining a GPS lock at the entrance or near a window. The centerpiece of our work is the EZ Localization algorithm, which runs on the localization server. The key intuition is that all of the observations reported to the server, even the many from unknown locations, are constrained by the physics of wireless propagation. EZ models these constraints and then uses a genetic algorithm to solve them. The results from our deployment in two different buildings are promising. Despite the absence of any explicit pre-deployment calibration, EZ yields a median localization error of 2m and 7m, respectively, in a small building and a large building, which is only somewhat worse than the 0.7m and 4m yielded by the best-performing but calibrationintensive Horus scheme [29] from prior work. Krishna Chintalapudi, Anand Padmanabha Iyer, Venkat N. Padmanabhan |
MobiCom | 2 |
| 2009 | Handling mobility across WiFi and WiMAXabstractPerformance of wireless data networks can be improved by integrating heterogeneous networks. Hence, emerging wireless Internet networks consist of heterogeneous wireless networks working in synergy. WiFi and WiMAX are particularly interesting in their ability towards mobile data oriented networking, and a scheme that enables mobility across these two would provide several advantages to end-users, wireless operators as well as Wireless Internet Service Providers (WISPs). In this work, we propose a novel, cost-effective and end-user friendly mobility scheme. Our approach does not require additional client software to handle WiFi-WiMAX mobility, or hardware changes in any of the network entities involved. We demonstrate the feasibility of our solution by developing an actual prototype. Anand Padmanabha Iyer, Jayaraman Iyer |
IWCMC | 1 |
| 2009 | Fast Resilient Jumbo frames in wireless LANsabstractWith the phenomenal growth of wireless networks and applications, it is increasingly important to deliver content efficiently and reliably over wireless links. However, wireless performance is still far from satisfactory due to limited wireless spectrum, inherent lossy wireless medium, and imperfect packet scheduling. While significant research has been done to improve wireless performance, much of the existing work focuses on individual design space. We take a holistic approach to optimizing wireless performance and resilience. We propose Fast Resilient Jumbo frames (FRJ), which exploit the synergy between three important design spaces: (i) frame size selection, (ii) partial packet recovery, and (iii) rate adaptation. While these design spaces are seemingly unrelated, we show that there are strong interactions between them and effectively leveraging these techniques can provide increased robustness and performance benefits in wireless LANs. FRJ uses jumbo frames to boost network throughput under good channel conditions and uses partial packet recovery to efficiently recover packet losses under bad channel conditions. FRJ also utilizes partial recovery aware rate adaptation to maximize throughput under partial recovery. Using real implementation and testbed experiments, we show that FRJ out-performs existing approaches in a wide range of scenarios. Anand Padmanabha Iyer, Gaurav Deshpande, Eric Rozner, Apurv Bhartia, Lili Qiu |
IWQoS | 1 |
| 2007 | ER: efficient retransmission scheme for wireless LANsabstractWireless LANs (WLANs) have been deployed at a remarkable rate at university campuses, office buildings, airports, hotels, and malls. Providing efficient and reliable wireless communications is challenging due to inherent lossy wireless medium and imperfect packet scheduling that results in packet collisions. In this paper, we develop an efficient retransmission scheme (ER) for wirless LANs. Instead of retransmitting the lost packets in their original forms, ER codes packets lost at different destinations and uses a single retransmission to potentially recover multiple packet losses. We develop a simple and practical protocol to realize the idea and implement it in both simulation and testbed, and our results demonstrate the effectiveness of this approach. Eric Rozner, Anand Padmanabha Iyer, Yogita Mehta, Lili Qiu, Mansoor Jafry |
CoNEXT | 2 |