VLDB 2026 Research / reviewers in the wild / expert
Arvind Krishnamurthy
dblp:k/AKrishnamurthy
· DBLP profile ↗
173ranked-venue papers
3as first author
34since 2021 · last 2026
0000-0002-9505-9528ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 90 · 19 since 2021Systems, architecture and hardware · 41 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 33 · 2 first-author · 8 since 2021Security and privacy · 7Artificial intelligence and machine learning · 5 · 1 since 2021Databases, data management, data science and information retrieval · 5Theory of computation · 4Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SG-IOV: Socket-Granular I/O Virtualization for SmartNIC-Based Container NetworksabstractI/O Virtualization (IOV) is a cornerstone of cloud computing, with container networking as a critical form of IOV in modern cloud paradigms. While container networks serve as feature-rich infrastructure, they incur a high CPU tax yet leave room for efficiency improvement. A natural idea is to offload container networks onto hardware such as SmartNICs via IOV interfaces. However, existing IOV mechanisms, such as SR-IOV, are misaligned with container requirements: limited device scalability versus high container density, packet-layer abstraction versus application-layer processing demands, and coarse-grained virtualization versus fine-grained container workloads. Chenxingyu Zhao, Jaehong Min, Shengkai Lin, Wei Zhang 0052, Kaiyuan Zhang 0001, Ming Liu 0027, Arvind Krishnamurthy |
ASPLOS (2) | 8 |
| 2026 | Reducing the GPU Memory Bottleneck with Lossless Compression for MLabstractMachine learning (ML) training and inference often process data sets far exceeding GPU memory capacity, forcing them to rely on PCIe for on-demand tensor transfers, causing critical transfer bottlenecks. Lossy compression has been proposed to relieve bottlenecks but introduces workload-dependent accuracy loss, making it complex or even prohibitive to use in existing ML deployments. Aditya K. Kamath, Arvind Krishnamurthy, Marco Canini, Simon Peter 0001 |
EuroSys | 2 |
| 2026 | FAST: An Efficient Scheduler for All-to-All GPU Communication
Yiran Lei, Dongjoo Lee 0001, Liangyu Zhao, Daniar Kurniawan, Chanmyeong Kim, Heetaek Jeong, Changsu Kim 0004, Hyeonseong Choi, Liangcheng Yu, Arvind Krishnamurthy, Justine Sherry, Eriko Nurvitadhi |
NSDI | 10 |
| 2026 | RLBoost: Harvesting Preemptible Cloud Resources for Cost-Efficient Reinforcement Learning on LLMs
Xueshen Liu, Haizhong Zheng, Juncheng Gu, Beidi Chen, Z. Morley Mao, Arvind Krishnamurthy, Ion Stoica |
NSDI | 7 |
| 2026 | ForestColl: Throughput-Optimal Collective Communications on Heterogeneous Network Fabrics
Liangyu Zhao, Saeed Maleki, Yuanhong Wang, Zezhou Wang, Hossein Pourreza, Arvind Krishnamurthy |
NSDI | 7 |
| 2026 | Efficient and Flexible Datapaths for Fine-Grained Rack-Scale Interconnects with Elastic QPabstractRack-scale interconnects serve as critical datapaths for emerging communication-intensive systems to scale up. Innovative solutions for this datapath are rising at a rapid pace, especially those based on Ethernet. However, existing hardware-based solutions, such as RDMA, face performance issues, particularly for small-message memory access, and suffer from the inflexibility of hardware-fixed processing. The community is actively pursuing efficient, flexible, and cost-effective rack-scale datapaths. Chenxingyu Zhao, Jaehong Min, Ming Liu 0027, Arvind Krishnamurthy |
SIGCOMM | 6 |
| 2025 | Rethinking RPC Communication for Microservices-based ApplicationsabstractFast and efficient RPCs are key to the performance of applications based on microservices. But RPC communication suffers from significant overhead today because it relies on the standard, layered protocol stack and loose coupling between the end host and in-network proxies that process RPCs. We propose delayering the RPC communication stack and tightly coupling the end host and in-network processing using high-level abstractions. This approach leads to more efficient and performant RPC communication because it eliminates many sources of overhead. Xiangfeng Zhu, Arvind Krishnamurthy, Sam Kumar, Ratul Mahajan, Danyang Zhuo |
HotOS | 5 |
| 2025 | White-Boxing RDMA with Packet-Granular Software Control
Chenxingyu Zhao, Jaehong Min, Ming Liu 0027, Arvind Krishnamurthy |
NSDI | 4 |
| 2025 | Efficient Direct-Connect Topologies for Collective Communications
Liangyu Zhao, Siddharth Pal, Tapan Chugh, Weiyang Wang, Jason Fantl, Prithwish Basu, Joud Khoury, Arvind Krishnamurthy |
NSDI | 8 |
| 2025 | High-level Programming for Application Networks
Xiangfeng Zhu, Banruo Liu, Yongtong Wu, Nikola Bojanic, Jingrong Chen 0002, Gilbert Louis Bernstein, Arvind Krishnamurthy, Sam Kumar, Ratul Mahajan, Danyang Zhuo |
NSDI | 8 |
| 2025 | NanoFlow: Towards Optimal Large Language Model Serving Throughput
Kan Zhu, Yilong Zhao 0002, Liangyu Zhao, Gefei Zuo, Yile Gu, Dedong Xie, Zihao Ye 0001, Keisuke Kamahori, Chien-Yu Lin, Ziren Wang, Stephanie Wang, Arvind Krishnamurthy, Baris Kasikci |
OSDI | 13 |
| 2025 | IC-Cache: Efficient Large Language Model Serving via In-context CachingabstractLarge language models (LLMs) have excelled in various applications, yet serving them at scale is challenging due to their substantial resource demands and high latency. Our real-world studies reveal that over 70% of user requests to LLMs have semantically similar counterparts, suggesting the potential for knowledge transfer among requests. However, naively caching and reusing past responses leads to a big quality drop. Yu Gan 0002, Nikhil Sarda, Lillian Tsai, Yanqi Zhou, Arvind Krishnamurthy, Fan Lai 0001, Henry M. Levy, David E. Culler |
SOSP | 7 |
| 2024 | CC-NIC: a Cache-Coherent Interface to the NICabstractEmerging interconnects make peripherals, such as the network interface controller (NIC), accessible through the processor's cache hierarchy, allowing these devices to participate in the CPU cache coherence protocol. This is a fundamental change from the separate I/O data paths and read-write transaction primitives of today's PCIe NICs. Our experiments show that the I/O data path characteristics cause NICs to prioritize CPU efficiency at the expense of inflated latency, an issue that can be mitigated by the emerging low-latency coherent interconnects. But, the coherence abstraction is not suited to current host-NIC access patterns. Applying existing signaling mechanisms and data structure layouts in a cache-coherent setting results in extraneous communication and cache retention, limiting performance. Redesigning the interface is necessary to minimize overheads and benefit from the new interactions coherence enables. This work contributes CC-NIC, a host-NIC interface design for coherent interconnects. We model CC-NIC using Intel's Ice Lake and Sapphire Rapids UPI interconnects, demonstrating the potential of optimizing for coherence. Our results show a maximum packet rate of 1.5Gpps and 980Gbps packet throughput. CC-NIC has 77% lower minimum latency, and 88% lower at 80% load, than today's PCIe NICs. We also demonstrate application-level core savings. Finally, we show that CC-NIC's benefits hold across a range of interconnect performance characteristics. Henry Schuh, Arvind Krishnamurthy, David E. Culler, Henry M. Levy, Luigi Rizzo, Samira Manabi Khan, Brent E. Stephens |
ASPLOS (1) | 2 |
| 2024 | SuperNIC: An FPGA-Based, Cloud-Oriented SmartNICabstractWith CPU scaling slowing down in today's data centers, more functionalities are being offloaded from the CPU to auxiliary devices. One such device is the SmartNIC, which is being increasingly adopted in data centers. In today's cloud environment, VMs on the same server can each have their own network computation (or network tasks) or workflows of network tasks to offload to a SmartNIC. These network tasks can be dynamically added/removed as VMs come and go and can be shared across VMs. Such dynamism demands that a SmartNIC not only schedules and processes packets but also manages and executes offloaded network tasks for different users. Although software solutions like an OS exist for managing software-based network tasks, such software-based SmartNICs cannot keep up with the quickly increasing data-center network speed. This paper proposes a new SmartNIC platform called SuperNIC that allows multiple tenants to efficiently and safely offload FPGA-based network computation DAGs. For efficiency and scalability, our core idea is to group network tasks into virtual chains that are dynamically mapped to different forms of physical chains depending on load and FPGA space availability. We further propose techniques to automatically scale network task chains with different types of parallelism. Moreover, we propose a fair sharing mechanism that considers both fair space sharing and fair time sharing of different types of hardware resources. Our FPGA prototype of SuperNIC achieves high bandwidth and low latency performance whilst efficiently utilizing and fairly sharing resources. Will Lin, Yizhou Shan, Ryan Kosta, Arvind Krishnamurthy, Yiying Zhang 0005 |
FPGA | 4 |
| 2024 | Efficient all-to-all Collective Communication Schedules for Direct-connect TopologiesabstractThe all-to-all collective communications primitive is widely used in machine learning (ML) and high performance computing (HPC) workloads, and optimizing its performance is of interest to both ML and HPC communities. All-to-all is a particularly challenging workload that can severely strain the underlying interconnect bandwidth at scale. This paper takes a holistic approach to optimize the performance of all-to-all collective communications on supercomputer-scale direct-connect interconnects. We address several algorithmic and practical challenges in developing efficient and bandwidth-optimal all-to-all schedules for any topology and lowering the schedules to various runtimes and interconnect technologies. We also propose a novel topology that delivers near-optimal all-to-all performance. Prithwish Basu, Liangyu Zhao, Jason Fantl, Siddharth Pal, Arvind Krishnamurthy, Joud Khoury |
HPDC | 5 |
| 2024 | Principles for Internet Congestion ManagementabstractGiven the technical flaws with---and the increasing non-observance of---the TCP-friendliness paradigm, we must rethink how the Internet should manage bandwidth allocation. We explore this question from first principles, but remain within the constraints of the Internet's current architecture and commercial arrangements. We propose a new framework, Recursive Congestion Shares (RCS), that provides bandwidth allocations independent of which congestion control algorithms flows use but consistent with the Internet's economics. We show that RCS achieves this goal using game-theoretic calculations and simulations as well as network emulation. Lloyd Brown, Albert Gran Alcoz, Frank Cangialosi, Akshay Narayan 0001, Mohammad Alizadeh, Hari Balakrishnan, Eric J. Friedman, Ethan Katz-Bassett, Arvind Krishnamurthy, Michael Schapira, Scott Shenker |
SIGCOMM | 9 |
| 2024 | An Architecture For Edge Networking ServicesabstractThe layered Internet architecture, while far from perfect, has provided a global and neutral platform for the development of a wide range of applications. However, this core architecture has been increasingly augmented with additional in-network functionality that improves the performance, security, and privacy of these applications. These additional in-network functions, which are typically implemented at the network edge, are consistent with the layering of the Internet architecture but deviate from two of the core tenets of the Internet: interconnection and end-to-end simplicity. In this paper, we propose an architecture for these edge networking services called the InterEdge that applies these two Internet tenets in a manner appropriate to edge services while not requiring changes to the underlying Internet architecture or infrastructure. Lloyd Brown, Emily Marx, Dev Bali, Emmanuel Amaro, Debnil Sur, Ezra Kissel, Inder Monga, Ethan Katz-Bassett, Arvind Krishnamurthy, James Murphy McCauley, Tejas Narechania, Aurojit Panda, Scott Shenker |
SIGCOMM | 9 |
| 2024 | Understanding the Host NetworkabstractThe host network integrates processor, memory, and peripheral interconnects to enable data transfer within the host. Several recent studies from production datacenters show that contention within the host network can have significant impact on end-to-end application performance. The goal of this paper is to build an in-depth understanding of such contention within the host network. Midhul Vuppalapati, Saksham Agarwal, Henry Schuh, Baris Kasikci, Arvind Krishnamurthy, Rachit Agarwal 0001 |
SIGCOMM | 5 |
| 2024 | Relational Network VerificationabstractRelational network verification is a new approach for validating network changes. In contrast to traditional network verification, which analyzes specifications for a single network snapshot, it analyzes specifications that capture similarities and differences between two network snapshots (e.g., pre- and post-change snapshots). Relational specifications are compact and precise because they focus on the flows and paths that change between snapshots and then simply mandate that all other network behaviors "stay the same", without enumerating them. To achieve similar guarantees, single-snapshot specifications would need to enumerate all flow and path behaviors that are not expected to change in order to enable checking that nothing has accidentally changed. Such specifications are proportional to network size, which makes them impractical to generate for many real-world networks. Xieyang Xu, Yifei Yuan 0001, Zachary Kincaid, Arvind Krishnamurthy, Ratul Mahajan, David Walker 0001, Ennan Zhai |
SIGCOMM | 4 |
| 2024 | eZNS: Elastic Zoned Namespace for Enhanced Performance Isolation and Device UtilizationabstractEmerging Zoned Namespace (ZNS) SSDs, providing the coarse-grained zone abstraction, hold the potential to significantly enhance the cost efficiency of future storage infrastructure and mitigate performance unpredictability. However, existing ZNS SSDs have a static zoned interface, making them in-adaptable to workload runtime behavior, unscalable to underlying hardware capabilities, and interfering with co-located zones. Applications either under-provision the zone resources yielding unsatisfied throughput, create over-provisioned zones and incur costs, or experience unexpected I/O latencies. We propose eZNS, an elastic-ZNS interface that exposes an adaptive zone with predictable characteristics. eZNS comprises two major components: a zone arbiter that manages zone allocation and active resources on the control plane, and a hierarchical I/O scheduler with read congestion control and write admission control on the data plane. Together, eZNS enables the transparent use of a ZNS SSD and closes the gap between application requirements and zone interface properties. Our evaluations over RocksDB demonstrate that eZNS outperforms a static zoned interface by 17.7% and 80.3% in throughput and tail latency, respectively, at most. Jaehong Min, Chenxingyu Zhao, Ming Liu 0027, Arvind Krishnamurthy |
ACM Trans. Storage | 4 |
| 2023 | Anticipatory Resource Allocation for ML TrainingabstractOur analysis of a large public cloud ML training service shows that resources remain unused likely because users statically (over-)allocate resources for their jobs given a desire for predictable performance, and state-of-the-art schedulers do not exploit idle resources lest they slow down some jobs excessively. We consider if an anticipatory scheduler, which schedules based on predictions of future job arrivals and durations, can improve over the state-of-the-art. We find that realizing gains from anticipation requires dealing effectively with prediction errors, and even the best predictors have errors that do not conform to simple models (such as bounded or i.i.d. error). We devise a novel anticipatory scheduler called SIA that is robust to such errors. On real workloads, SIA reduces job latency by an average of 2.83× over the current production scheduler, while reducing the likelihood of job slowdowns by orders of magnitude relative to schedulers that naïvely share resources. Tapan Chugh, Srikanth Kandula, Arvind Krishnamurthy, Ratul Mahajan, Ishai Menache |
SoCC | 3 |
| 2023 | Dissecting Overheads of Service Mesh SidecarsabstractService meshes play a central role in the modern application ecosystem by providing an easy and flexible way to connect microservices of a distributed application. However, because of how they interpose on application traffic, they can substantially increase application latency and its resource consumption. We develop a tool called MeshInsight to help developers quantify the overhead of service meshes in deployment scenarios of interest and make informed trade-offs about their functionality vs. overhead. Using MeshInsight, we confirm that service meshes can have high overhead---up to 269% higher latency and up to 163% more virtual CPU cores for our benchmark applications---but the severity is intimately tied to how they are configured and the application workload. IPC (inter-process communication) and socket writes dominate when the service mesh operates as a TCP proxy, but protocol parsing dominates when it operates as an HTTP proxy. MeshInsight also enables us to study the end-to-end impact of optimizations to service meshes. We show that not all seemingly-promising optimizations lead to a notable overhead reduction in realistic settings. Xiangfeng Zhu, Guozhen She, Yu Zhang 0209, Yongsu Zhang, Xuan Kelvin Zou, Xiongchun Duan, Peng He 0003, Arvind Krishnamurthy, Matthew Lentz, Danyang Zhuo, Ratul Mahajan |
SoCC | 9 |
| 2023 | How I Learned to Stop Worrying About CCA ContentionabstractThis paper asks whether inter-flow contention between congestion control algorithms (CCAs) is a dominant factor in determining a flow's bandwidth allocation in today's Internet. We hypothesize that CCA contention typically does not determine a flow's bandwidth allocation, present an initial analysis in support of this hypothesis, propose a measurement technique and study to settle this question, and discuss the implications should the hypothesis prove true. Lloyd Brown, Yash Kothari, Akshay Narayan 0001, Arvind Krishnamurthy, Aurojit Panda, Justine Sherry, Scott Shenker |
HotNets | 4 |
| 2023 | Application Defined NetworksabstractWith the rise of microservices, the execution environment of many cloud applications has become a set of virtual machines or containers connected by a flexible and feature-rich virtual network. We argue that the implementation of such virtual networks should be completely application-specific and not layered on top of general-purpose network abstractions from the Internet age. Such layering tends to more than double the latency and CPU usage of applications. We propose application-defined networks in which developers specify network functionality in a high-level language and a controller generates a custom distributed implementation that runs across available hardware and software resources. Experiments with a preliminary prototype suggest that, compared to the state of the art, ADN reduces latency by up to 20x and increases the throughput by up to 6x. Xiangfeng Zhu, Weixin Deng, Banruo Liu, Jingrong Chen 0002, Thomas E. Anderson, Arvind Krishnamurthy, Ratul Mahajan, Danyang Zhuo |
HotNets | 7 |
| 2023 | eZNS: An Elastic Zoned Namespace for Commodity ZNS SSDs
Jaehong Min, Chenxingyu Zhao, Ming Liu 0027, Arvind Krishnamurthy |
OSDI | 4 |
| 2023 | Host Congestion ControlabstractThe conventional wisdom in systems and networking communities is that congestion happens primarily within the network fabric. However, adoption of high-bandwidth access links and relatively stagnant technology trends for resources within hosts have led to emergence of host congestion---that is, congestion within the host network that enables data exchange between NIC and CPU/memory. Such host congestion alters the many assumptions entrenched within decades of research and practice of congestion control. Saksham Agarwal, Arvind Krishnamurthy, Rachit Agarwal 0001 |
SIGCOMM | 2 |
| 2023 | Unleashing SmartNIC Packet Processing Performance in P4abstractSmartNICs are on the rise as a packet processing platform, with the trend towards a uniform P4 programming model. However, unleashing SmartNIC packet processing performance in P4 is a formidable task. Traditional SmartNIC optimizations rely on low-level program tuning, but P4 abstractions operate at one level above. At the same time, today's P4 optimizations primarily focus on resource packing rather than performance tuning. We develop Pipeleon, an automated performance optimization framework for P4 programmable SmartNICs. We introduce techniques that are tailored to the performance characteristics of SmartNICs, and further leverage dynamic workload patterns for profile-guided optimization. Pipeleon pinpoints program hotspots at the P4 level and computes runtime optimization plans to specialize the program layout based on the latest profile. We have prototyped Pipeleon and applied it to optimize two popular P4 SmartNICs---Nvidia BlueField2 and Netronome Agilio CX---as well as a software SmartNIC emulator extended based on BMv2. Our results show that Pipeleon significantly improves SmartNIC packet processing performance in realistic scenarios. Jiarong Xing, Yiming Qiu 0001, Kuo-Feng Hsu, Songyuan Sui, Khalid Manaa, Omer Shabtai, Yonatan Piasetzky, Matty Kadosh, Arvind Krishnamurthy, T. S. Eugene Ng, Ang Chen 0001 |
SIGCOMM | 9 |
| 2023 | A Cloud-Scale Characterization of Remote Procedure CallsabstractThe global scale and challenging requirements of modern cloud applications have led to the development of complex, widely distributed, service-oriented applications. One enabler of such applications is the remote procedure call (RPC), which provides location-independent communication and hides the myriad of cloud communication complexities and requirements within the RPC stack. Understanding RPCs is thus one key to understanding the behavior of cloud applications. While there have been numerous studies of RPCs in distributed systems, as well as attempts to optimize RPC overheads with both software and hardware, there is still a lack of knowledge about the characteristics of RPCs "in the wild" in the modern cloud environment. Korakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu 0001, Hassan M. G. Wassel, Soheil Hassas Yeganeh, Alex C. Snoeren, Arvind Krishnamurthy, David E. Culler, Henry M. Levy |
SOSP | 8 |
| 2022 | Runtime Programmable Switches
Jiarong Xing, Kuo-Feng Hsu, Matty Kadosh, Alan Lo, Yonatan Piasetzky, Arvind Krishnamurthy, Ang Chen 0001 |
NSDI | 6 |
| 2021 | A Vision for Runtime Programmable NetworksabstractOur community has made significant progress in developing programmable network infrastructure, starting from the control plane and expanding to the data plane. As a latest trend, network devices are becoming runtime programmable while serving live traffic. This allows for reprogramming of individual device programs at fine-grained timescales to add or remove network functions. Many applications and services, however, need control over a combination of devices, including end host stacks, NICs, and switches, to accomplish their goals. We lay out our vision for runtime programmable networks, building upon device-level features to provide live, network-wide, runtime reprogramming. A whole-stack approach is needed with new programming models, compiler support, and network management abstractions. We outline a research agenda as a call to arms to the community. Jiarong Xing, Yiming Qiu 0001, Kuo-Feng Hsu, Matty Kadosh, Alan Lo, Aditya Akella, Thomas E. Anderson, Arvind Krishnamurthy, T. S. Eugene Ng, Ang Chen 0001 |
HotNets | 9 |
| 2021 | AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the Fly
Tianyi Zhou 0001, Liangyu Zhao, Yibo Zhu 0001, Chuanxiong Guo, Marco Canini, Arvind Krishnamurthy |
ICLR | 7 |
| 2021 | Scaling Distributed Machine Learning with In-Network Aggregation
Amedeo Sapio, Marco Canini, Chen-Yu Ho 0001, Jacob Nelson 0001, Panos Kalnis, Changhoon Kim, Arvind Krishnamurthy, Masoud Moshref, Dan R. K. Ports, Peter Richtárik |
NSDI | 7 |
| 2021 | Gimbal: enabling multi-tenant storage disaggregation on SmartNIC JBOFsabstractEmerging SmartNIC-based disaggregated NVMe storage has become a promising storage infrastructure due to its competitive IO performance and low cost. These SmartNIC JBOFs are shared among multiple co-resident applications, and there is a need for the platform to ensure fairness, QoS, and high utilization. Unfortunately, given the limited computing capability of the SmartNICs and the non-deterministic nature of NVMe drives, it is challenging to provide such support on today's SmartNIC JBOFs. Jaehong Min, Ming Liu 0027, Tapan Chugh, Chenxingyu Zhao, Andrew Wei, In Hwan Doh, Arvind Krishnamurthy |
SIGCOMM | 7 |
| 2021 | Xenic: SmartNIC-Accelerated Distributed TransactionsabstractHigh-performance distributed transactions require efficient remote operations on database memory and protocol metadata. The high communication cost of this workload calls for hardware acceleration. Recent research has applied RDMA to this end, leveraging the network controller to manipulate host memory without consuming CPU cycles on the target server. However, the basic read/write RDMA primitives demand trade-offs in data structure and protocol design, limiting their benefits. SmartNICs are a flexible alternative for fast distributed transactions, adding programmable compute cores and on-board memory to the network interface. Applying measured performance characteristics, we design Xenic, a SmartNIC-optimized transaction processing system. Xenic applies an asynchronous, aggregated execution model to maximize network and core efficiency. Xenic's co-designed data store achieves low-overhead remote object accesses. Additionally, Xenic uses flexible, point-to-point communication patterns between SmartNICs to minimize transaction commit latency. We compare Xenic against prior RDMA- and RPC-based transaction systems with the TPC-C, Retwis, and Smallbank benchmarks. Our results for the three benchmarks show 2.42x, 2.07x, and 2.21x throughput improvement, 59%, 42%, and 22% latency reduction, while saving 2.3, 8.1, and 10.1 threads per server. Henry Schuh, Weihao Liang, Ming Liu 0027, Jacob Nelson 0001, Arvind Krishnamurthy |
SOSP | 5 |
| 2020 | Talek: Private Group Messaging with Hidden Access PatternsabstractTalek is a private group messaging system that sends messages through potentially untrustworthy servers, while hiding both data content and the communication patterns among its users. Talek explores a new point in the design space of private messaging; it guarantees access sequence indistinguishability, which is among the strongest guarantees in the space, while assuming an anytrust threat model, which is only slightly weaker than the strongest threat model currently found in related work. Our results suggest that this is a pragmatic point in the design space, since it supports strong privacy and good performance: we demonstrate a 3-server Talek cluster that achieves throughput of 9,433 messages/second for 32,000 active users with 1.7-second end-to-end latency. To achieve its security goals without coordination between clients, Talek relies on information-theoretic private information retrieval. To achieve good performance and minimize server-side storage, Talek introduces new techniques and optimizations that may be of independent interest, e.g., a novel use of blocked cuckoo hashing and support for private notifications. The latter provide a private, efficient mechanism for users to learn, without polling, which logs have new messages. Raymond Cheng 0001, William Scott 0002, Elisaweta Masserova, Irene Zhang, Vipul Goyal, Thomas E. Anderson, Arvind Krishnamurthy, Bryan Parno |
ACSAC | 7 |
| 2020 | Meerkat: multicore-scalable replicated transactions following the zero-coordination principleabstractTraditionally, the high cost of network communication between servers has hidden the impact of cross-core coordination in replicated systems. However, new technologies, like kernel-bypass networking and faster network links, have exposed hidden bottlenecks in distributed systems. Adriana Szekeres, Michael J. Whittaker, Jialin Li 0001, Naveen Kr. Sharma, Arvind Krishnamurthy, Dan R. K. Ports, Irene Zhang |
EuroSys | 5 |
| 2020 | Remote Memory CallsabstractIn this paper we propose an extension to RDMA, called Remote Memory Calls (RMCs), that allows applications to install a customized set of 1-sided RDMA operations. We then explain how RMCs can be implemented on the forthcoming generation of SmartNICs and discuss the resulting tradeoffs between RMCs, 1-sided and 2-sided RDMA operations. Emmanuel Amaro, Zhihong Luo, Amy Ousterhout, Arvind Krishnamurthy, Aurojit Panda, Sylvia Ratnasamy, Scott Shenker |
HotNets | 4 |
| 2020 | On the Future of Congestion Control for the Public InternetabstractThe conventional wisdom requires that all congestion control algorithms deployed on the public Internet be TCP-friendly. If universally obeyed, this requirement would greatly constrain the future of such congestion control algorithms. If partially ignored, as is increasingly likely, then there could be significant inequities in the bandwidth received by different flows. To avoid this dilemma, we propose an alternative to the TCP-friendly paradigm that can accommodate innovation, is consistent with the Internet's current economic model, and is feasible to deploy given current usage trends. Lloyd Brown, Ganesh Ananthanarayanan, Ethan Katz-Bassett, Arvind Krishnamurthy, Sylvia Ratnasamy, Michael Schapira, Scott Shenker |
HotNets | 4 |
| 2020 | Bertha: Tunneling through the Network APIabstractNetwork APIs such as UNIX sockets, DPDK, Netmap, etc. assume that networks provide only end-to-end connectivity. However, networks increasingly include smart NICs and programmable switches that can implement both network and application functions. Several recent works have shown the benefit of offloading application functionality to the network, but using these approaches requires changing not just the applications, but also network and system configuration. In this paper we propose Bertha, a network API that provides a uniform abstraction for offloads, aiming to simplify their use. Akshay Narayan 0001, Aurojit Panda, Mohammad Alizadeh, Hari Balakrishnan, Arvind Krishnamurthy, Scott Shenker |
HotNets | 5 |
| 2020 | Fine-Grained Replicated State Machines for a Cluster Storage System
Ming Liu 0027, Arvind Krishnamurthy, Harsha V. Madhyastha, Rishi Bhardwaj, Chinmay Kamat, Huapeng Yuan, Aditya Jaltade, Roger Liao, Pavan Konka, Anoop Jawahar |
NSDI | 2 |
| 2020 | Programmable Calendar Queues for High-speed Packet Scheduling
Naveen Kr. Sharma, Chenxingyu Zhao, Ming Liu 0027, Pravein G. Kannan, Changhoon Kim, Arvind Krishnamurthy, Anirudh Sivaraman |
NSDI | 6 |
| 2020 | Automated Verification of Customizable Middlebox Properties with Gravel
Kaiyuan Zhang 0001, Danyang Zhuo, Aditya Akella, Arvind Krishnamurthy, Xi Wang 0005 |
NSDI | 4 |
| 2020 | A Public Option for the CoreabstractThis paper is focused not on the Internet architecture - as defined by layering, the narrow waist of IP, and other core design principles - but on the Internet infrastructure, as embodied in the technologies and organizations that provide Internet service. In this paper we discuss both the challenges and the opportunities that make this an auspicious time to revisit how we might best structure the Internet's infrastructure. Currently, the tasks of transit-between-domains and last-mile-delivery are jointly handled by a set of ISPs who interconnect through BGP. In this paper we propose cleanly separating these two tasks. For transit, we propose the creation of a "public option" for the Internet's core backbone. This public option core, which complements rather than replaces the backbones used by large-scale ISPs, would (i) run an open market for backbone bandwidth so it could leverage links offered by third-parties, and (ii) structure its terms-of-service to enforce network neutrality so as to encourage competition and reduce the advantage of large incumbents. Yotam Harchol, Dirk Bergemann, Nick Feamster, Eric J. Friedman, Arvind Krishnamurthy, Aurojit Panda, Sylvia Ratnasamy, Michael Schapira, Scott Shenker |
SIGCOMM | 5 |
| 2020 | Gallium: Automated Software Middlebox Offloading to Programmable SwitchesabstractResearchers have shown that offloading software middleboxes (e.g., NAT, firewall, load balancer) to programmable switches can yield orders-of-magnitude performance gains. However, it requires manually selecting the middle-box components to offload and rewriting the offloaded code in P4, a domain-specific language for programmable switches. We design and implement Gallium, a compiler that transforms an input software middlebox into two parts---a P4 program that runs on a programmable switch and an x86 non-offloaded program that runs on a regular middlebox server. Gallium ensures that (1) the combined effect of the P4 program and the non-offloaded program is functionally equivalent to the input middlebox program, (2) the P4 program respects the resource constraints in the programmable switch, and (3) the run-to-completion semantics are met under concurrent execution. Our evaluations show that Gallium saves 21-79% of processing cycles and reduces latency by about 31% across various software middleboxes. Kaiyuan Zhang 0001, Danyang Zhuo, Arvind Krishnamurthy |
SIGCOMM | 3 |
| 2020 | End the Senseless Killing: Improving Memory Management for Mobile Operating Systems
Niel Lebeck, Arvind Krishnamurthy, Henry M. Levy, Irene Zhang |
USENIX ATC | 2 |
| 2019 | TAS: TCP Acceleration as an OS ServiceabstractAs datacenter network speeds rise, an increasing fraction of server CPU cycles is consumed by TCP packet processing, in particular for remote procedure calls (RPCs). To free server CPUs from this burden, various existing approaches have attempted to mitigate these overheads, by bypassing the OS kernel, customizing the TCP stack for an application, or by offloading packet processing to dedicated hardware. In doing so, these approaches trade security, agility, or generality for efficiency. Neither trade-off is fully desirable in the fast-evolving commodity cloud. Antoine Kaufmann, Tim Stamler, Simon Peter 0001, Naveen Kr. Sharma, Arvind Krishnamurthy, Thomas E. Anderson |
EuroSys | 5 |
| 2019 | Practical Safe Linux Kernel ExtensibilityabstractThe ability to extend kernel functionality safely has long been a design goal for operating systems. Modern operating systems, such as Linux, are structured for extensibility to enable sharing a single code base among many environments. Unfortunately, safety has lagged behind, and bugs in kernel extensions continue to cause problems. We study three recent kernel extensions critical to Docker containers (Overlay File System, Open vSwitch Datapath, and AppArmor) to guide further research in extension safety. We find that all the studied kernel extensions suffer from the same set of low-level memory, concurrency, and type errors. Though safe kernel extensibility is a well-studied area, existing solutions are heavyweight, requiring extensive changes to the kernel and/or expensive runtime checks. We then explore the feasibility of writing kernel extensions in a high-level, type safe language (i.e., Rust) while preserving compatibility with Linux and find this to be an appealing approach. We show that there are key challenges to implementing this approach and propose potential solutions. Samantha Miller, Kaiyuan Zhang 0001, Danyang Zhuo, Shibin Xu, Arvind Krishnamurthy, Thomas E. Anderson |
HotOS | 5 |
| 2019 | Stable and Practical AS Relationship Inference with ProbLink
Colin Scott, Amogh Dhamdhere, Vasileios Giotsas, Arvind Krishnamurthy, Scott Shenker |
NSDI | 5 |
| 2019 | Slim: OS Kernel Support for a Low-Overhead Container Overlay Network
Danyang Zhuo, Kaiyuan Zhang 0001, Yibo Zhu 0001, Hongqiang Harry Liu, Matthew Rockett, Arvind Krishnamurthy, Thomas E. Anderson |
NSDI | 6 |
| 2019 | Zooming in on wide-area latencies to a global cloud providerabstractThe network communications between the cloud and the client have become the weak link for global cloud services that aim to provide low latency services to their clients. In this paper, we first characterize WAN latency from the viewpoint of a large cloud provider Azure, whose network edges serve hundreds of billions of TCP connections a day across hundreds of locations worldwide. In particular, we focus on instances of latency degradation and design a tool, BlameIt, that enables cloud operators to localize the cause (i.e., faulty AS) of such degradation. BlameIt uses passive diagnosis, using measurements of existing connections between clients and the cloud locations, to localize the cause to one of cloud, middle, or client segments. Then it invokes selective active probing (within a probing budget) to localize the cause more precisely. We validate BlameIt by comparing its automatic fault localization results with that arrived at by network engineers manually, and observe that BlameIt correctly localized the problem in all the 88 incidents. Further, BlameIt issues 72X fewer active probes than a solution relying on active probing alone, and is deployed in production at Azure. Sundararajan Renganathan, Ganesh Ananthanarayanan, Junchen Jiang, Venkat N. Padmanabhan, Manuel Schröder, Matt Calder, Arvind Krishnamurthy |
SIGCOMM | 8 |
| 2019 | Offloading distributed applications onto smartNICs using iPipeabstractEmerging Multicore SoC SmartNICs, enclosing rich computing resources (e.g., a multicore processor, onboard DRAM, accelerators, programmable DMA engines), hold the potential to offload generic datacenter server tasks. However, it is unclear how to use a SmartNIC efficiently and maximize the offloading benefits, especially for distributed applications. Towards this end, we characterize four commodity SmartNICs and summarize the offloading performance implications from four perspectives: traffic control, computing capability, onboard memory, and host communication. Ming Liu 0027, Tianyi Cui, Henry Schuh, Arvind Krishnamurthy, Simon Peter 0001 |
SIGCOMM | 4 |
| 2019 | Nexus: a GPU cluster engine for accelerating DNN-based video analysisabstractWe address the problem of serving Deep Neural Networks (DNNs) efficiently from a cluster of GPUs. In order to realize the promise of very low-cost processing made by accelerators such as GPUs, it is essential to run them at sustained high utilization. Doing so requires cluster-scale resource management that performs detailed scheduling of GPUs, reasoning about groups of DNN invocations that need to be co-scheduled, and moving from the conventional whole-DNN execution model to executing fragments of DNNs. Nexus is a fully implemented system that includes these innovations. In large-scale case studies on 16 GPUs, when required to stay within latency constraints at least 99% of the time, Nexus can process requests at rates 1.8-12.7X higher than state of the art systems can. A long-running multi-application deployment stays within 84% of optimal utilization and, on a 100-GPU cluster, violates latency SLOs on 0.27% of requests. Haichen Shen, Lequn Chen 0001, Liangyu Zhao, Bingyu Kong, Matthai Philipose, Arvind Krishnamurthy, Ravi Sundaram |
SOSP | 7 |
| 2019 | E3: Energy-Efficient Microservices on SmartNIC-Accelerated Servers
Ming Liu 0027, Simon Peter 0001, Arvind Krishnamurthy, Phitchaya Mangpo Phothilimthana |
USENIX ATC | 3 |
| 2018 | Parameter Hub: a Rack-Scale Parameter Server for Distributed Deep Neural Network TrainingabstractDistributed deep neural network (DDNN) training constitutes an increasingly important workload that frequently runs in the cloud. Larger DNN models and faster compute engines are shifting DDNN training bottlenecks from computation to communication. This paper characterizes DDNN training to precisely pinpoint these bottlenecks. We found that timely training requires high performance parameter servers (PSs) with optimized network stacks and gradient processing pipelines, as well as server and network hardware with balanced computation and communication resources. We therefore propose PHub, a high performance multi-tenant, rack-scale PS design. PHub co-designs the PS software and hardware to accelerate rack-level and hierarchical cross-rack parameter exchange, with an API compatible with many DDNN training frameworks. PHub provides a performance improvement of up to 2.7x compared to state-of-the-art cloud-based distributed training techniques for image classification workloads, with 25% better throughput per dollar. Jacob Nelson 0001, Luis Ceze, Amar Phanishayee, Arvind Krishnamurthy |
SoCC | 5 |
| 2018 | MultiNyx: a multi-level abstraction framework for systematic analysis of hypervisorsabstractMultiNyx is a new framework designed to systematically analyze modern virtual machine monitors (VMMs), which rely on complex processor extensions to enhance their efficiency. To achieve better scalability, MultiNyx introduces selective, multi-level symbolic execution: it analyzes most instructions at a high semantic level, and leverages an executable specification (e.g., the Bochs CPU emulator) to analyze complex instructions at a low semantic level. MultiNyx seamlessly transitions between these different semantic levels of analysis by converting their state. Pedro Fonseca 0001, Xi Wang 0005, Arvind Krishnamurthy |
EuroSys | 3 |
| 2018 | Learning to Optimize Tensor ProgramsabstractWe introduce a learning-based framework to optimize tensor programs for deep learning workloads. Efficient implementations of tensor operators, such as matrix multiplication and high dimensional convolution are key enablers of effective deep learning systems. However, existing systems rely on manually optimized libraries such as cuDNN where only a narrow range of server class GPUs are well-supported. The reliance on hardware specific operator libraries limits the applicability of high-level graph optimizations and incurs significant engineering costs when deploying to new hardware targets. We use learning to remove this engineering burden. We learn domain specific statistical cost models to guide the search of tensor operator implementations over billions of possible program variants. We further accelerate the search by effective model transfer across workloads. Experimental results show that our framework delivers performance competitive with state-of-the-art hand-tuned libraries for low-power CPU, mobile GPU, and server-class GPU. Tianqi Chen 0001, Lianmin Zheng, Eddie Q. Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, Arvind Krishnamurthy |
NeurIPS | 8 |
| 2018 | Approximating Fair Queueing on Reconfigurable Switches
Naveen Kr. Sharma, Ming Liu 0027, Kishore Atreya, Arvind Krishnamurthy |
NSDI | 4 |
| 2018 | Deepview: Virtual Disk Failure Diagnosis and Pattern Detection for Azure
Qiao Zhang 0001, Chuanxiong Guo, Yingnong Dang, Nick Swanson, Xinsheng Yang, Randolph Yao, Murali Chintalapati, Arvind Krishnamurthy, Thomas E. Anderson |
NSDI | 9 |
| 2018 | TVM: An Automated End-to-End Optimizing Compiler for Deep Learning
Tianqi Chen 0001, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Q. Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Luis Ceze, Carlos Guestrin, Arvind Krishnamurthy |
OSDI | 12 |
| 2018 | Revisiting network support for RDMAabstractThe advent of RoCE (RDMA over Converged Ethernet) has led to a significant increase in the use of RDMA in datacenter networks. To achieve good performance, RoCE requires a lossless network which is in turn achieved by enabling Priority Flow Control (PFC) within the network. However, PFC brings with it a host of problems such as head-of-the-line blocking, congestion spreading, and occasional deadlocks. Rather than seek to fix these issues, we instead ask: is PFC fundamentally required to support RDMA over Ethernet? Radhika Mittal, Alexander Shpiner, Aurojit Panda, Eitan Zahavi, Arvind Krishnamurthy, Sylvia Ratnasamy, Scott Shenker |
SIGCOMM | 5 |
| 2017 | IncBricks: Toward In-Network Computation with an In-Network CacheabstractThe emergence of programmable network devices and the increasing data traffic of datacenters motivate the idea of in-network computation. By offloading compute operations onto intermediate networking devices (e.g., switches, network accelerators, middleboxes), one can (1) serve network requests on the fly with low latency; (2) reduce datacenter traffic and mitigate network congestion; and (3) save energy by running servers in a low-power mode. However, since (1) existing switch technology doesn't provide general computing capabilities, and (2) commodity datacenter networks are complex (e.g., hierarchical fat-tree topologies, multipath communication), enabling in-network computation inside a datacenter is challenging. Ming Liu 0027, Jacob Nelson 0001, Luis Ceze, Arvind Krishnamurthy, Kishore Atreya |
ASPLOS | 5 |
| 2017 | Fast Video Classification via Adaptive Cascading of Deep ModelsabstractRecent advances have enabled oracle classifiers that can classify across many classes and input distributions with high accuracy without retraining. However, these classifiers are relatively heavyweight, so that applying them to classify video is costly. We show that day-to-day video exhibits highly skewed class distributions over the short term, and that these distributions can be classified by much simpler models. We formulate the problem of detecting the short-term skews online and exploiting models based on it as a new sequential decision making problem dubbed the Online Bandit Problem, and present a new algorithm to solve it. When applied to recognizing faces in TV shows and movies, we realize end-to-end classification speedups of 2.4-7.8x/2.6-11.2x (on GPU/CPU) relative to a state-of-the-art convolutional neural network, at competitive accuracy. Haichen Shen, Seungyeop Han, Matthai Philipose, Arvind Krishnamurthy |
CVPR | 4 |
| 2017 | An Empirical Study on the Correctness of Formally Verified Distributed SystemsabstractRecent advances in formal verification techniques enabled the implementation of distributed systems with machine-checked proofs. While results are encouraging, the importance of distributed systems warrants a large scale evaluation of the results and verification practices. Pedro Fonseca 0001, Kaiyuan Zhang 0001, Xi Wang 0005, Arvind Krishnamurthy |
EuroSys | 4 |
| 2017 | High-resolution measurement of data center microburstsabstractData centers house some of the largest, fastest networks in the world. In contrast to and as a result of their speed, these networks operate on very small timescales---a 100 Gbps port processes a single packet in at most 500 ns with end-to-end network latencies of under a millisecond. In this study, we explore the fine-grained behaviors of a large production data center using extremely high-resolution measurements (10s to 100s of microsecond) of rack-level traffic. Our results show that characterizing network events like congestion and synchronized behavior in data centers does indeed require the use of such measurements. In fact, we observe that more than 70% of bursts on the racks we measured are sustained for at most tens of microseconds: a range that is orders of magnitude higher-resolution than most deployed measurement frameworks. Congestion events observed by less granular measurements are likely collections of smaller μbursts. Thus, we find that traffic at the edge is significantly less balanced than other metrics might suggest. Beyond the implications for measurement granularity, we hope these results will inform future data center load balancing and congestion control protocols. Qiao Zhang 0001, Vincent Liu 0001, Hongyi Zeng, Arvind Krishnamurthy |
Internet Measurement Conference | 4 |
| 2017 | Curator: Self-Managing Storage for Enterprise Clusters
Ignacio Cano 0001, Srinivas Aiyar, Varun Arora, Manosiz Bhattacharyya, Akhilesh Chaganti, Chern Cheah, Brent N. Chun, Vinayak Khot, Arvind Krishnamurthy |
NSDI | 10 |
| 2017 | SCL: Simplifying Distributed SDN Control Planes
Aurojit Panda, Wenting Zheng, Xiaohe Hu, Arvind Krishnamurthy, Scott Shenker |
NSDI | 4 |
| 2017 | Evaluating the Power of Flexible Packet Processing for Network Resource Allocation
Naveen Kr. Sharma, Antoine Kaufmann, Thomas E. Anderson, Arvind Krishnamurthy, Jacob Nelson 0001, Simon Peter 0001 |
NSDI | 4 |
| 2017 | RAIL: A Case for Redundant Arrays of Inexpensive Links in Data Center Networks
Danyang Zhuo, Manya Ghobadi, Ratul Mahajan, Amar Phanishayee, Xuan Kelvin Zou, Hang Guan, Arvind Krishnamurthy, Thomas E. Anderson |
NSDI | 7 |
| 2017 | Understanding and Mitigating Packet Corruption in Data Center NetworksabstractWe take a comprehensive look at packet corruption in data center networks, which leads to packet losses and application performance degradation. By studying 350K links across 15 production data centers, we find that the extent of corruption losses is significant and that its characteristics differ markedly from congestion losses. Corruption impacts fewer links than congestion, but imposes a heavier loss rate; and unlike congestion, corruption rate on a link is stable over time and is not correlated with its utilization. Danyang Zhuo, Manya Ghobadi, Ratul Mahajan, Klaus-Tycho Förster, Arvind Krishnamurthy, Thomas E. Anderson |
SIGCOMM | 5 |
| 2017 | Building Consistent Transactions with Inconsistent ReplicationabstractApplication programmers increasingly prefer distributed storage systems with strong consistency and distributed transactions (e.g., Google’s Spanner) for their strong guarantees and ease of use. Unfortunately, existing transactional storage systems are expensive to use—in part, because they require costly replication protocols, like Paxos, for fault tolerance. In this article, we present a new approach that makes transactional storage systems more affordable: We eliminate consistency from the replication protocol, while still providing distributed transactions with strong consistency to applications. We present the Transactional Application Protocol for Inconsistent Replication (TAPIR), the first transaction protocol to use a novel replication protocol, called inconsistent replication , that provides fault tolerance without consistency. By enforcing strong consistency only in the transaction protocol, TAPIR can commit transactions in a single round-trip and order distributed transactions without centralized coordination. We demonstrate the use of TAPIR in a transactional key-value store, TAPIR-KV . Compared to conventional systems, TAPIR-KV provides better latency and better throughput. Irene Zhang, Naveen Kr. Sharma, Adriana Szekeres, Arvind Krishnamurthy, Dan R. K. Ports |
ACM Trans. Comput. Syst. | 4 |
| 2016 | Specifying and Checking File System Crash-Consistency ModelsabstractApplications depend on persistent storage to recover state after system crashes. But the POSIX file system interfaces do not define the possible outcomes of a crash. As a result, it is difficult for application writers to correctly understand the ordering of and dependencies between file system operations, which can lead to corrupt application state and, in the worst case, catastrophic data loss. This paper presents crash-consistency models, analogous to memory consistency models, which describe the behavior of a file system across crashes. Crash-consistency models include both litmus tests, which demonstrate allowed and forbidden behaviors, and axiomatic and operational specifications. We present a formal framework for developing crash-consistency models, and a toolkit, called Ferrite, for validating those models against real file system implementations. We develop a crash-consistency model for ext4, and use Ferrite to demonstrate unintuitive crash behaviors of the ext4 implementation. To demonstrate the utility of crash-consistency models to application writers, we use our models to prototype proof-of-concept verification and synthesis tools, as well as new library interfaces for crash-safe applications. James Bornholt, Antoine Kaufmann, Jialin Li 0001, Arvind Krishnamurthy, Emina Torlak, Xi Wang 0005 |
ASPLOS | 4 |
| 2016 | High Performance Packet Processing with FlexNICabstractThe recent surge of network I/O performance has put enormous pressure on memory and software I/O processing sub systems. We argue that the primary reason for high memory and processing overheads is the inefficient use of these resources by current commodity network interface cards (NICs). We propose FlexNIC, a flexible network DMA interface that can be used by operating systems and applications alike to reduce packet processing overheads. FlexNIC allows services to install packet processing rules into the NIC, which then executes simple operations on packets while exchanging them with host memory. Thus, our proposal moves some of the packet processing traditionally done in software to the NIC, where it can be done flexibly and at high speed. Antoine Kaufmann, Simon Peter 0001, Naveen Kr. Sharma, Thomas E. Anderson, Arvind Krishnamurthy |
ASPLOS | 5 |
| 2016 | Characterizing Private Clouds: A Large-Scale Empirical Analysis of Enterprise ClustersabstractThere is an increasing trend in the use of on-premise clusters within companies. Security, regulatory constraints, and enhanced service quality push organizations to work in these so called private cloud environments. On the other hand, the deployment of private enterprise clusters requires careful consideration of what will be necessary or may happen in the future, both in terms of compute demands and failures, as they lack the public cloud's flexibility to immediately provision new nodes in case of demand spikes or node failures. Ignacio Cano 0001, Srinivas Aiyar, Arvind Krishnamurthy |
SoCC | 3 |
| 2016 | Radiatus: a Shared-Nothing Server-Side Web ArchitectureabstractWeb applications are a frequent target of successful attacks. In most web frameworks, the damage is amplified by the fact that application code is responsible for security enforcement. In this paper, we design and evaluate Radiatus, a shared-nothing web framework where application-specific computation and storage on the server is contained within a sandbox with the privileges of the end-user. By strongly isolating users, user data and service availability can be protected from application vulnerabilities. Raymond Cheng 0001, William Scott 0002, Paul M. Ellenbogen, Jon Howell, Franziska Roesner, Arvind Krishnamurthy, Thomas E. Anderson |
SoCC | 6 |
| 2016 | Rack-level Congestion ControlabstractMany data center traffic patterns exhibit abundant concurrent connections and high churn. In the face of these characteristics, server-centric congestion control is a poor fit—each connection, no matter how small, must start from scratch when testing when and how much to send along a given path. This is despite the fact that there are a large number of flows that may have already probed the same exact path, not just at a server level, but also at a rack level. Thus, we argue for rack-level congestion control in which all connections are tunneled through rack-to-rack JumboFlows. This design allows an entire rack’s connections to cooperate with one another for better fairness and performance, particularly for short flows. In this paper, we examine situations in which JumboFlows might be useful and present a preliminary design of a system (RackCC) that implements JumboFlows. Danyang Zhuo, Qiao Zhang 0001, Vincent Liu 0001, Arvind Krishnamurthy, Thomas E. Anderson |
HotNets | 4 |
| 2016 | MCDNN: An Approximation-Based Execution Framework for Deep Stream Processing Under Resource ConstraintsabstractWe consider applying computer vision to video on cloud-backed mobile devices using Deep Neural Networks (DNNs). The computational demands of DNNs are high enough that, without careful resource management, such applications strain device battery, wireless data, and cloud cost budgets. We pose the corresponding resource management problem, which we call Approximate Model Scheduling, as one of serving a stream of heterogeneous (i.e., solving multiple classification problems) requests under resource constraints. We present the design and implementation of an optimizing compiler and runtime scheduler to address this problem. Going beyond traditional resource allocators, we allow each request to be served approximately, by systematically trading off DNN classification accuracy for resource use, and remotely, by reasoning about on-device/cloud execution trade-offs. To inform the resource allocator, we characterize how several common DNNs, when subjected to state-of-the art optimizations, trade off accuracy for resource use such as memory, computation, and energy. The heterogeneous streaming setting is a novel one for DNN execution, and we introduce two new and powerful DNN optimizations that exploit it. Using the challenging continuous mobile vision domain as a case study, we show that our techniques yield significant reductions in resource usage and perform effectively over a broad range of operating conditions. Seungyeop Han, Haichen Shen, Matthai Philipose, Sharad Agarwal, Alec Wolman, Arvind Krishnamurthy |
MobiSys | 6 |
| 2016 | Minimizing Faulty Executions of Distributed Systems
Colin Scott, Aurojit Panda, Vjekoslav Brajkovic, George C. Necula, Arvind Krishnamurthy, Scott Shenker |
NSDI | 5 |
| 2016 | Speeding up Web Page Loads with Shandian
Xiao Sophia Wang, Arvind Krishnamurthy, David Wetherall |
NSDI | 2 |
| 2016 | Scalable verification of border gateway protocol configurations with an SMT solverabstractInternet Service Providers (ISPs) use the Border Gateway Protocol (BGP) to announce and exchange routes for de- livering packets through the internet. ISPs must carefully configure their BGP routers to ensure traffic is routed reli- ably and securely. Correctly configuring BGP routers has proven challenging in practice, and misconfiguration has led to worldwide outages and traffic hijacks. This paper presents Bagpipe, a system that enables ISPs to declaratively express BGP policies and that automatically verifies that router configurations implement such policies. The novel initial network reduction soundly reduces policy verification to a search for counterexamples in a finite space. An SMT-based symbolic execution engine performs this search efficiently. Bagpipe reduces the size of its search space using predicate abstraction and parallelizes its search using symbolic variable hoisting. Bagpipe's policy specification language is expressive: we expressed policies inferred from real AS configurations, policies from the literature, and policies for 10 Juniper TechLibrary configuration scenarios. Bagpipe is efficient: we ran it on three ASes with a total of over 240,000 lines of Cisco and Juniper BGP configuration. Bagpipe is effective: it revealed 19 policy violations without issuing any false positives. Konstantin Weitz, Doug Woos, Emina Torlak, Michael D. Ernst, Arvind Krishnamurthy, Zachary Tatlock |
OOPSLA | 5 |
| 2016 | Diamond: Automating Data Management and Storage for Wide-Area, Reactive Applications
Irene Zhang, Niel Lebeck, Pedro Fonseca 0001, Brandon Holt, Raymond Cheng 0001, Ariadna Norberg, Arvind Krishnamurthy, Henry M. Levy |
OSDI | 7 |
| 2016 | Satellite: Joint Analysis of CDNs and Network-Level Interference
Will Scott, Thomas E. Anderson, Tadayoshi Kohno, Arvind Krishnamurthy |
USENIX ATC | 4 |
| 2016 | Caching Doesn't Improve Mobile Web Performance (Much)
Jamshed Vesuna, Colin Scott, Michael Buettner, Michael Piatek, Arvind Krishnamurthy, Scott Shenker |
USENIX ATC | 5 |
| 2016 | Arrakis: The Operating System Is the Control PlaneabstractRecent device hardware trends enable a new approach to the design of network server operating systems. In a traditional operating system, the kernel mediates access to device hardware by server applications to enforce process isolation as well as network and disk security. We have designed and implemented a new operating system, Arrakis, that splits the traditional role of the kernel in two. Applications have direct access to virtualized I/O devices, allowing most I/O operations to skip the kernel entirely, while the kernel is re-engineered to provide network and disk protection without kernel mediation of every operation. We describe the hardware and software changes needed to take advantage of this new abstraction, and we illustrate its power by showing improvements of 2 to 5 × in latency and 9 × throughput for a popular persistent NoSQL store relative to a well-tuned Linux implementation. Simon Peter 0001, Jialin Li 0001, Irene Zhang, Dan R. K. Ports, Doug Woos, Arvind Krishnamurthy, Thomas E. Anderson, Timothy Roscoe |
ACM Trans. Comput. Syst. | 6 |
| 2015 | Subways: a case for redundant, inexpensive data center edge linksabstractAs network demand increases, data center network operators face a number of challenges including the need to add capacity to the network. Unfortunately, network upgrades can be an expensive proposition, particularly at the edge of the network where most of the network's cost lies. Vincent Liu 0001, Danyang Zhuo, Simon Peter 0001, Arvind Krishnamurthy, Thomas E. Anderson |
CoNEXT | 4 |
| 2015 | FlexNIC: Rethinking Network DMA
Antoine Kaufmann, Simon Peter 0001, Thomas E. Anderson, Arvind Krishnamurthy |
HotOS | 4 |
| 2015 | Designing Distributed Systems Using Approximate Synchrony in Data Center Networks
Dan R. K. Ports, Jialin Li 0001, Vincent Liu 0001, Naveen Kr. Sharma, Arvind Krishnamurthy |
NSDI | 5 |
| 2015 | Rollback-Recovery for MiddleboxesabstractNetwork middleboxes must offer high availability, with automatic failover when a device fails. Achieving high availability is challenging because failover must correctly restore lost state (e.g., activity logs, port mappings) but must do so quickly (e.g., in less than typical transport timeout values to minimize disruption to applications) and with little overhead to failure-free operation (e.g., additional per-packet latencies of 10-100s of us). No existing middlebox design provides failover that is correct, fast to recover, and imposes little increased latency on failure-free operations. We present a new design for fault-tolerance in middleboxes that achieves these three goals. Our system, FTMB (for Fault-Tolerant MiddleBox), adopts the classical approach of "rollback recovery" in which a system uses information logged during normal operation to correctly reconstruct state after a failure. However, traditional rollback recovery cannot maintain high throughput given the frequent output rate of middleboxes. Hence, we design a novel solution to record middlebox state which relies on two mechanisms: (1) 'ordered logging', which provides lightweight logging of the information needed after recovery, and (2) a `parallel release' algorithm which, when coupled with ordered logging, ensures that recovery is always correct. We implement ordered logging and parallel release in Click and show that for our test applications our design adds only 30$\mu$s of latency to median per packet latencies. Our system introduces moderate throughput overheads (5-30%) and can reconstruct lost state in 40-275ms for practical systems. Justine Sherry, Peter Xiang Gao, Soumya Basu 0003, Aurojit Panda, Arvind Krishnamurthy, Christian Maciocco, Maziar Manesh, Sylvia Ratnasamy, Luigi Rizzo, Scott Shenker |
SIGCOMM | 5 |
| 2015 | Building consistent transactions with inconsistent replicationabstractApplication programmers increasingly prefer distributed storage systems with strong consistency and distributed transactions (e.g., Google's Spanner) for their strong guarantees and ease of use. Unfortunately, existing transactional storage systems are expensive to use -- in part because they require costly replication protocols, like Paxos, for fault tolerance. In this paper, we present a new approach that makes transactional storage systems more affordable: we eliminate consistency from the replication protocol while still providing distributed transactions with strong consistency to applications. Irene Zhang, Naveen Kr. Sharma, Adriana Szekeres, Arvind Krishnamurthy, Dan R. K. Ports |
SOSP | 4 |
| 2015 | MetaSync: File Synchronization Across Multiple Untrusted Storage Services
Seungyeop Han, Haichen Shen, Taesoo Kim, Arvind Krishnamurthy, Thomas E. Anderson, David Wetherall |
USENIX ATC | 4 |
| 2015 | Using Declarative Specification to Improve the Understanding, Extensibility, and Comparison of Model-Inference AlgorithmsabstractIt is a staple development practice to log system behavior. Numerous powerful model-inference algorithms have been proposed to aid developers in log analysis and system understanding. Unfortunately, existing algorithms are typically declared procedurally, making them difficult to understand, extend, and compare. This paper presents InvariMint, an approach to specify model-inference algorithms declaratively. We applied the InvariMint declarative approach to two model-inference algorithms. The evaluation results illustrate that InvariMint (1) leads to new fundamental insights and better understanding of existing algorithms, (2) simplifies creation of new algorithms, including hybrids that combine or extend existing algorithms, and (3) makes it easy to compare and contrast previously published algorithms. InvariMint's declarative approach can outperform procedural implementations. For example, on a log of 50,000 events, InvariMint's declarative implementation of the kTails algorithm completes in 12 seconds, while a procedural implementation completes in 18 minutes. We also found that InvariMint's declarative version of the Synoptic algorithm can be over 170 times faster than the procedural implementation. Ivan Beschastnikh, Yuriy Brun, Jenny Abrahamson, Michael D. Ernst, Arvind Krishnamurthy |
IEEE Trans. Software Eng. | 5 |
| 2014 | A Highly Available Software Defined FabricabstractExisting SDNs rely on a collection of intricate, mutually-dependent mechanisms to implement a logically centralized control plane. These cyclical dependencies and lack of clean separation of concerns can impact the availability of SDNs, such that a handful of link failures could render entire portions of an SDN non-functional. This paper shows why and when this could happen, and makes the case for taking a fresh look at architecting SDNs for robustness to faults from the ground up. Our approach carefully synthesizes various key distributed systems ideas -- in particular, reliable flooding, global snapshots, and replicated controllers. We argue informally that it can offer high availability in the face of a variety of network failures, but much work needs to be done to make our approach scalable and general. Thus, our paper represents a starting point for a broader discussion on approaches for building highly available SDNs. Aditya Akella, Arvind Krishnamurthy |
HotNets | 2 |
| 2014 | Towards High-Performance Application-Level Storage Management
Simon Peter 0001, Jialin Li 0001, Irene Zhang, Dan R. K. Ports, Thomas E. Anderson, Arvind Krishnamurthy, Mark Zbikowski, Doug Woos |
HotStorage | 6 |
| 2014 | Inferring models of concurrent systems from logs of their behavior with CSightabstractConcurrent systems are notoriously difficult to debug and understand. A common way of gaining insight into system behavior is to inspect execution logs and documentation. Unfortunately, manual inspection of logs is an arduous process, and documentation is often incomplete and out of sync with the implementation. Ivan Beschastnikh, Yuriy Brun, Michael D. Ernst, Arvind Krishnamurthy |
ICSE | 4 |
| 2014 | How Much Can We Micro-Cache Web Pages?abstractBrowser caches are widely used to improve the performance of Web page loads. Unfortunately, current object-based caching is too coarse-grained to minimize the costs associated with small, localized updates to a Web object. In this paper, we evaluate the benefits if caching were performed at a finer granularity and at different levels (i.e., computed layout and compiled JavaScript). By analyzing Web pages gathered over two years, we find that both layout and code are highly cacheable, suggesting that our proposal can radically reduce time to first paint. We also find that mobile pages are similar to their desktop counterparts in terms of the amount and composition of updates. Xiao Sophia Wang, Arvind Krishnamurthy, David Wetherall |
Internet Measurement Conference | 2 |
| 2014 | How Speedy is SPDY?
Xiao Sophia Wang, Aruna Balasubramanian, Arvind Krishnamurthy, David Wetherall |
NSDI | 3 |
| 2014 | Arrakis: The Operating System is the Control Plane
Simon Peter 0001, Jialin Li 0001, Irene Zhang, Dan R. K. Ports, Doug Woos, Arvind Krishnamurthy, Thomas E. Anderson, Timothy Roscoe |
OSDI | 6 |
| 2014 | Customizable and Extensible Deployment for Mobile/Cloud Applications
Irene Zhang, Adriana Szekeres, Dana Van Aken, Isaac Ackerman, Steve D. Gribble, Arvind Krishnamurthy, Henry M. Levy |
OSDI | 6 |
| 2014 | One tunnel is (often) enoughabstractA longstanding problem with the Internet is that it is vulnerable to outages, black holes, hijacking and denial of service. Although architectural solutions have been proposed to address many of these issues, they have had difficulty being adopted due to the need for widespread adoption before most users would see any benefit. This is especially relevant as the Internet is increasingly used for applications where correct and continuous operation is essential. Simon Peter 0001, Umar Javed, Qiao Zhang 0001, Doug Woos, Thomas E. Anderson, Arvind Krishnamurthy |
SIGCOMM | 6 |
| 2013 | Unifying FSM-inference algorithms through declarative specificationabstractLogging system behavior is a staple development practice. Numerous powerful model inference algorithms have been proposed to aid developers in log analysis and system understanding. Unfortunately, existing algorithms are difficult to understand, extend, and compare. This paper presents InvariMint, an approach to specify model inference algorithms declaratively. We applied InvariMint to two model inference algorithms and present evaluation results to illustrate that InvariMint (1) leads to new fundamental insights and better understanding of existing algorithms, (2) simplifies creation of new algorithms, including hybrids that extend existing algorithms, and (3) makes it easy to compare and contrast previously published algorithms. Finally, algorithms specified with InvariMint can outperform their procedural versions. Ivan Beschastnikh, Yuriy Brun, Jenny Abrahamson, Michael D. Ernst, Arvind Krishnamurthy |
ICSE | 5 |
| 2013 | F10: A Fault-Tolerant Engineered Network
Vincent Liu 0001, Daniel Halperin, Arvind Krishnamurthy, Thomas E. Anderson |
NSDI | 3 |
| 2013 | Demystifying Page Load Performance with WProf
Xiao Sophia Wang, Aruna Balasubramanian, Arvind Krishnamurthy, David Wetherall |
NSDI | 3 |
| 2013 | Expressive privacy control with pseudonymsabstractAs personal information increases in value, the incentives for remote services to collect as much of it as possible increase as well. In the current Internet, the default assumption is that all behavior can be correlated using a variety of identifying information, not the least of which is a user's IP address. Tools like Tor, Privoxy, and even NATs, are located at the opposite end of the spectrum and prevent any behavior from being linked. Instead, our goal is to provide users with more control over linkability---which activites of the user can be correlated at the remote services---not necessarily more anonymity. Seungyeop Han, Vincent Liu 0001, Qifan Pu, Simon Peter 0001, Thomas E. Anderson, Arvind Krishnamurthy, David Wetherall |
SIGCOMM | 6 |
| 2013 | PoiRoot: investigating the root cause of interdomain path changesabstractInterdomain path changes occur frequently. Because routing protocols expose insufficient information to reason about all changes, the general problem of identifying the root cause remains unsolved. In this work, we design and evaluate PoiRoot, a real-time system that allows a provider to accurately isolate the root cause (the network responsible) of path changes affecting its prefixes. First, we develop a new model describing path changes and use it to provably identify the set of all potentially responsible networks. Next, we develop a recursive algorithm that accurately isolates the root cause of any path change. We observe that the algorithm requires monitoring paths that are generally not visible using standard measurement tools. To address this limitation, we combine existing measurement tools in new ways to acquire path information required for isolating the root cause of a path change. We evaluate PoiRoot on path changes obtained through controlled Internet experiments, simulations, and "in-the-wild" measurements. We demonstrate that PoiRoot is highly accurate, works well even with partial information, and generally narrows down the root cause to a single network or two neighboring ones. On controlled experiments PoiRoot is 100% accurate, as opposed to prior work which is accurate only 61.7% of the time. Umar Javed, Ítalo S. Cunha, David R. Choffnes, Ethan Katz-Bassett, Thomas E. Anderson, Arvind Krishnamurthy |
SIGCOMM | 6 |
| 2012 | FreeDOM: a new baseline for the webabstractFree web services often face growing pains. In the current client-server access model, the cost of providing a service increases with its popularity. This leads organizations that want to provide services free-of-charge to rely to donations, advertisements, or mergers with larger companies to cope with operational costs. Raymond Cheng 0001, William Scott 0002, Arvind Krishnamurthy, Thomas E. Anderson |
HotNets | 3 |
| 2012 | LIFEGUARD: practical repair of persistent route failuresabstractThe Internet was designed to always find a route if there is a policy-compliant path. However, in many cases, connectivity is disrupted despite the existence of an underlying valid path. The research community has focused on short-term outages that occur during route convergence. There has been less progress on addressing avoidable long-lasting outages. Our measurements show that long-lasting events contribute significantly to overall unavailability. Ethan Katz-Bassett, Colin Scott, David R. Choffnes, Ítalo S. Cunha, Vytautas Valancius, Nick Feamster, Harsha V. Madhyastha, Thomas E. Anderson, Arvind Krishnamurthy |
SIGCOMM | 9 |
| 2012 | FairCloud: sharing the network in cloud computingabstractThe network, similar to CPU and memory, is a critical and shared resource in the cloud. However, unlike other resources, it is neither shared proportionally to payment, nor do cloud providers offer minimum guarantees on network bandwidth. The reason networks are more difficult to share is because the network allocation of a virtual machine (VM) X depends not only on the VMs running on the same machine with X, but also on the other VMs that X communicates with and the cross-traffic on each link used by X. In this paper, we start from the above requirements--payment proportionality and minimum guarantees--and show that the network-specific challenges lead to fundamental tradeoffs when sharing cloud networks. We then propose a set of properties to explicitly express these tradeoffs. Finally, we present three allocation policies that allow us to navigate the tradeoff space. We evaluate their characteristics through simulation and testbed experiments to show that they can provide minimum guarantees and achieve better proportionality than existing solutions. Lucian Popa 0002, Gautam Kumar 0001, Mosharaf Chowdhury, Arvind Krishnamurthy, Sylvia Ratnasamy, Ion Stoica |
SIGCOMM | 4 |
| 2012 | Making middleboxes someone else's problem: network processing as a cloud serviceabstractModern enterprises almost ubiquitously deploy middlebox processing services to improve security and performance in their networks. Despite this, we find that today's middlebox infrastructure is expensive, complex to manage, and creates new failure modes for the networks that use them. Given the promise of cloud computing to decrease costs, ease management, and provide elasticity and fault-tolerance, we argue that middlebox processing can benefit from outsourcing the cloud. Arriving at a feasible implementation, however, is challenging due to the need to achieve functional equivalence with traditional middlebox deployments without sacrificing performance or increasing network complexity. Justine Sherry, Shaddi Hasan, Colin Scott, Arvind Krishnamurthy, Sylvia Ratnasamy, Vyas Sekar |
SIGCOMM | 4 |
| 2011 | FairCloud: sharing the network in cloud computingabstractThe network is a crucial resource in cloud computing, but in contrast to other resources such as CPU or memory, the network is currently shared in a best effort manner. However, sharing the network in a datacenter is more challenging than sharing the other resources. The key difficulty is that the network allocation for a VM X depends not only on the VMs running on the same machine with X, but also on the other VMs that X communicates with, as well as on the cross-traffic on each link used by X. In this paper, we first propose a set of desirable properties for allocating the network bandwidth in a datacenter at the VM granularity, and show that there exists a fundamental tradeoff between the ability to share congested links in proportion to payment and the ability to provide minimal bandwidth guarantees to VMs. Second, we show that the existing allocation models violate one or more of these properties, and propose a mechanism that can select different points in the aforementioned tradeoff between payment proportionality and bandwidth guarantees. Lucian Popa 0002, Arvind Krishnamurthy, Sylvia Ratnasamy, Ion Stoica |
HotNets | 2 |
| 2011 | Machiavellian routing: improving internet availability with BGP poisoningabstractWe propose a new approach to mitigate disruptions of Internet connectivity. The Internet was designed to always find a route if there is a policy-compliant path; however, in many cases, connectivity is disrupted despite the existence of an underlying valid path. The research community has done considerable work on this problem, much of it focused on short-term outages that occur during route convergence. There has been less progress on addressing avoidable long-lasting outages. Our measurements show that long-lasting events contribute significantly to overall unavailability. Ethan Katz-Bassett, David R. Choffnes, Ítalo S. Cunha, Colin Scott, Thomas E. Anderson, Arvind Krishnamurthy |
HotNets | 6 |
| 2011 | Tor instead of IPabstractAs the Internet has become more popular, it has increasingly been a target and medium for monitoring, censorship, content discrimination, and denial of service. Although anonymizing overlays such as Tor [2] provide some help to end users in combating these trends, the overlays themselves have become targets in turn. In this paper, we take a fresh approach: instead of running Tor on top of IP, we propose to run Tor instead of IP. We ask: what might the Internet look like if privacy and censorship resistance had been designed in from scratch? To be practical, any proposal also needs to be robust to failures, achieve reasonable efficiency compared to today's Internet, and be consistent with ISP economic concerns. Although preliminary, we argue that our design achieves these goals. Vincent Liu 0001, Seungyeop Han, Arvind Krishnamurthy, Thomas E. Anderson |
HotNets | 3 |
| 2011 | ETTM: A Scalable Fault Tolerant Network Manager
Colin Dixon, Hardeep Uppal, Vjekoslav Brajkovic, Dane Brandon, Thomas E. Anderson, Arvind Krishnamurthy |
NSDI | 6 |
| 2011 | Scalable consistency in ScatterabstractDistributed storage systems often trade off strong semantics for improved scalability. This paper describes the design, implementation, and evaluation of Scatter, a scalable and consistent distributed key-value storage system. Scatter adopts the highly decentralized and self-organizing structure of scalable peer-to-peer systems, while preserving linearizable consistency even under adverse circumstances. Our prototype implementation demonstrates that even with very short node lifetimes, it is possible to build a scalable and consistent system with practical performance. Lisa Glendenning, Ivan Beschastnikh, Arvind Krishnamurthy, Thomas E. Anderson |
SOSP | 3 |
| 2011 | deSEO: Combating Search-Result Poisoning
John P. John, Fang Yu 0002, Yinglian Xie, Arvind Krishnamurthy, Martín Abadi |
USENIX Security Symposium | 4 |
| 2011 | Heat-seeking honeypots: design and experienceabstractMany malicious activities on the Web today make use of compromised Web servers, because these servers often have high pageranks and provide free resources. Attackers are therefore constantly searching for vulnerable servers. In this work, we aim to understand how attackers find, compromise, and misuse vulnerable servers. Specifically, we present heat-seeking honeypots that actively attract attackers, dynamically generate and deploy honeypot pages, then analyze logs to identify attack patterns. John P. John, Fang Yu 0002, Yinglian Xie, Arvind Krishnamurthy, Martín Abadi |
WWW | 4 |
| 2010 | Retaining sandbox containment despite bugs in privileged memory-safe codeabstractFlaws in the standard libraries of secure sandboxes represent a major security threat to billions of devices worldwide. The standard libraries are hard to secure because they frequently need to perform low-level operations that are forbidden in untrusted application code. Existing designs have a single, large trusted computing base that contains security checks at the boundaries between trusted and untrusted code. Unfortunately, flaws in the standard library often allow an attacker to escape the security protections of the sandbox. Justin Cappos, Armon Dadgar, Jeff Rasley, Justin Samuel, Ivan Beschastnikh, Cosmin Barsan, Arvind Krishnamurthy, Thomas E. Anderson |
CCS | 7 |
| 2010 | A cost comparison of datacenter network architecturesabstractThere is a growing body of research exploring new network architectures for the data center. These proposals all seek to improve the scalability and cost-effectiveness of current data center networks, but adopt very different approaches to doing so. For example, some proposals build networks entirely out of switches while others do so using a combination of switches and servers. How do these different network architectures compare? For that matter, by what metrics should we even begin to compare these architectures? Lucian Popa 0002, Sylvia Ratnasamy, Gianluca Iannaccone, Arvind Krishnamurthy, Ion Stoica |
CoNEXT | 4 |
| 2010 | Resolving IP aliases with prespecified timestampsabstractOperators and researchers want accurate router-level views of the Internet for purposes including troubleshooting and modeling. However, tools such as traceroute return IP addresses. Because routers may have dozens of IP addresses, or aliases, multiple measurements may return different addresses, obscuring whether they represent the same machine. While many techniques exist to address this issue by identifying some IP aliases, these techniques, even in combination, find only a subset of alias pairs. To improve this state, we design and evaluate a new alias resolution technique using the IP prespecified timestamp option. This option allows a sender to request timestamp val- ues from multiple IP addresses in the same probe. By careful arrangement of these IP addresses, we show that we can infer aliases in many cases. In this paper, we conduct a measurement study of how many routers support IP timestamps, demonstrating that enough honor the option to base our technique on it. Using our technique, and compared to the most accurate alias information available, we find that 94.7% of the aliases identified by our technique are true positives. Further, we show that our IP timestamp-based technique complements existing alias resolution techniques, providing significant gains by discovering previously unidentifiable aliases. Justine Sherry, Ethan Katz-Bassett, Mary Pimenova, Harsha V. Madhyastha, Thomas E. Anderson, Arvind Krishnamurthy |
Internet Measurement Conference | 6 |
| 2010 | Reverse traceroute
Ethan Katz-Bassett, Harsha V. Madhyastha, Vijay Kumar Adhikari, Colin Scott, Justine Sherry, Peter van Wesep, Thomas E. Anderson, Arvind Krishnamurthy |
NSDI | 8 |
| 2010 | Contracts: Practical Contribution Incentives for P2P Live Streaming
Michael Piatek, Arvind Krishnamurthy, Arun Venkataramani, Yang Richard Yang, David Zhang 0002, Alexander Jaffe |
NSDI | 2 |
| 2010 | Comet: An active distributed key-value store
Roxana Geambasu, Amit Levy 0001, Tadayoshi Kohno, Arvind Krishnamurthy, Henry M. Levy |
OSDI | 4 |
| 2010 | Privacy-preserving P2P data sharing with OneSwarmabstractPrivacy -- the protection of information from unauthorized disclosure -- is increasingly scarce on the Internet. The lack of privacy is particularly true for popular peer-to-peer data sharing applications such as BitTorrent where user behavior is easily monitored by third parties. Anonymizing overlays such as Tor and Freenet can improve user privacy, but only at a cost of substantially reduced performance. Most users are caught in the middle, unwilling to sacrifice either privacy or performance. Tomas Isdal, Michael Piatek, Arvind Krishnamurthy, Thomas E. Anderson |
SIGCOMM | 3 |
| 2010 | Searching the Searchers with SearchAudit
John P. John, Fang Yu 0002, Yinglian Xie, Martín Abadi, Arvind Krishnamurthy |
USENIX Security Symposium | 5 |
| 2009 | Pitfalls for ISP-friendly P2P design
Michael Piatek, Harsha V. Madhyastha, John P. John, Arvind Krishnamurthy, Thomas E. Anderson |
HotNets | 4 |
| 2009 | An End to the Middle
Colin Dixon, Arvind Krishnamurthy, Thomas E. Anderson |
HotOS | 2 |
| 2009 | Moving beyond end-to-end path information to optimize CDN performanceabstractReplicating content across a geographically distributed set of servers and redirecting clients to the closest server in terms of latency has emerged as a common paradigm for improving client performance. In this paper, we analyze latencies measured from servers in Google's content distribution network (CDN) to clients all across the Internet to study the effectiveness of latency-based server selection. Our main result is that redirecting every client to the server with least latency does not suffice to optimize client latencies. First, even though most clients are served by a geographically nearby CDN node, a sizeable fraction of experience latencies several tens of milliseconds higher than other in the same region. Second, we find that queueing delays often override the benefits of a client interacting with a nearby server. Rupa Krishnan, Harsha V. Madhyastha, Sridhar Srinivasan, Sushant Jain, Arvind Krishnamurthy, Thomas E. Anderson, Jie Gao 0001 |
Internet Measurement Conference | 5 |
| 2009 | Studying Spamming Botnets Using Botlab
John P. John, Alexander Moshchuk, Steve D. Gribble, Arvind Krishnamurthy |
NSDI | 4 |
| 2009 | iPlane Nano: Path Prediction for Peer-to-Peer Applications
Harsha V. Madhyastha, Ethan Katz-Bassett, Thomas E. Anderson, Arvind Krishnamurthy, Arun Venkataramani |
NSDI | 4 |
| 2009 | Seattle: a platform for educational cloud computingabstractCloud computing is rapidly increasing in popularity. Companies such as RedHat, Microsoft, Amazon, Google, and IBM are increasingly funding cloud computing infrastructure and research, making it important for students to gain the necessary skills to work with cloud-based resources. This paper presents a free, educational research platform called Seattle that is community-driven, a common denominator for diverse platform types, and is broadly deployed. Justin Cappos, Ivan Beschastnikh, Arvind Krishnamurthy, Thomas E. Anderson |
SIGCSE | 3 |
| 2008 | Phalanx: Withstanding Multimillion-Node Botnets
Colin Dixon, Thomas E. Anderson, Arvind Krishnamurthy |
NSDI | 3 |
| 2008 | Consensus Routing: The Internet as a Distributed System. (Best Paper)
John P. John, Ethan Katz-Bassett, Arvind Krishnamurthy, Thomas E. Anderson, Arun Venkataramani |
NSDI | 3 |
| 2008 | Studying Black Holes in the Internet with Hubble
Ethan Katz-Bassett, Harsha V. Madhyastha, John P. John, Arvind Krishnamurthy, David Wetherall, Thomas E. Anderson |
NSDI | 4 |
| 2008 | One Hop Reputations for Peer to Peer File Sharing Workloads
Michael Piatek, Tomas Isdal, Arvind Krishnamurthy, Thomas E. Anderson |
NSDI | 3 |
| 2008 | P4p: provider portal for applications
Haiyong Xie 0001, Yang Richard Yang, Arvind Krishnamurthy, Yanbin Grace Liu, Avi Silberschatz |
SIGCOMM | 3 |
| 2008 | Challenges and Directions for Monitoring P2P File Sharing Networks - or - Why My Printer Received a DMCA Takedown Notice
Michael Piatek, Tadayoshi Kohno, Arvind Krishnamurthy |
HotSec | 3 |
| 2008 | Privacy-Preserving Location Tracking of Lost or Stolen Devices: Cryptographic Techniques and Replacing Trusted Third Parties with DHTs
Thomas Ristenpart, Gabriel Maganis, Arvind Krishnamurthy, Tadayoshi Kohno |
USENIX Security Symposium | 3 |
| 2007 | Profiling a million user dhtabstractDistributed hash tables (DHTs) provide scalable, key-based lookup of objects in dynamic network environments. Although DHTs have been studied extensively from an analytical perspective, only recently have wide deployments enabled empirical examination. This paper reports measurements of the Azureus BitTorrent client's DHT, which is in active use by more than 1 million nodes on a daily basis. The Azureus DHT operates on untrusted, unreliable end-hosts, offering a glimpse into the implementation challenges associated with making structured overlays work in practice. Our measurements provide characterizations of churn, overhead, and performance in this environment. We leverage these measurements to drive the design of a modified DHT lookup algorithm that reduces median DHT lookup time by an order of magnitude for a nominal increase in overhead. Jarret Falkner, Michael Piatek, John P. John, Arvind Krishnamurthy, Thomas E. Anderson |
Internet Measurement Conference | 4 |
| 2007 | Do Incentives Build Robustness in BitTorrent? (Awarded Best Student Paper)
Michael Piatek, Tomas Isdal, Thomas E. Anderson, Arvind Krishnamurthy, Arun Venkataramani |
NSDI | 4 |
| 2007 | Leveraging BitTorrent for End Host Measurements
Tomas Isdal, Michael Piatek, Arvind Krishnamurthy, Thomas E. Anderson |
PAM | 3 |
| 2007 | Distributed Construction of a Multi-level Topology with Unpredictable Metric Values for Wireless Networks
Johannes Lessmann, Arvind Krishnamurthy |
WiMob | 2 |
| 2006 | Towards IP geolocation using delay and topology measurementsabstractWe present Topology-based Geolocation (TBG), a novel approach to estimating the geographic location of arbitrary Internet hosts. We motivate our work by showing that 1) existing approaches, based on end-to-end delay measurements from a set of landmarks, fail to outperform much simpler techniques, and 2) the error of these approaches is strongly determined by the distance to the nearest landmark, even when triangulation is used to combine estimates from different landmarks. Our approach improves on these earlier techniques by leveraging network topology, along with measurements of network delay, to constrain host position. We convert topology and delay data into a set of constraints, then solve for router and host locations simultaneously. This approach improves the consistency of location estimates, reducing the error substantially for structured networks in our experiments on Abilene and Sprint. For networks with insufficient structural constraints, our techniques integrate external hints that are validated using measurements before being trusted. Together, these techniques lower the median estimation error for our university-based dataset to 67 km vs. 228 km for the best previous approach. Ethan Katz-Bassett, John P. John, Arvind Krishnamurthy, David Wetherall, Thomas E. Anderson, Yatin Chawathe |
Internet Measurement Conference | 3 |
| 2006 | A structural approach to latency predictionabstractSeveral models have been recently proposed for predicting the latency of end to end Internet paths. These models treat the Internet as a black-box, ignoring its internal structure. While these models are simple, they can often fail systematically; for example, the most widely used models use metric embeddings that predict no benefit to detour routes even though half of all Internet routes can benefit from detours.In this paper, we adopt a structural approach that predicts path latency based on measurements of the Internet's routing topology, PoP connectivity, and routing policy. We find that our approach outperforms Vivaldi, the most widely used black-box model. Furthermore, unlike metric embeddings, our approach successfully predicts 65% of detour routes in the Internet. The number of measurements used in our approach is comparable with that required by black box techniques, but using traceroutes instead of pings. Harsha V. Madhyastha, Thomas E. Anderson, Arvind Krishnamurthy, Neil Spring, Arun Venkataramani |
Internet Measurement Conference | 3 |
| 2006 | Optimal Capacity Sharing of Networks with Multiple OverlaysabstractOverlay networks have emerged as a generic networking paradigm to improve network performance and construct new applications. Although many overlay algorithms have been proposed lately, they tend to focus on a single overlay, without considering how to share network capacity with other traffic and other overlays. In this paper, we study optimal capacity sharing of network with multiple overlays. We first formulate the problem of optimal capacity sharing of networks with multiple overlays as a nonlinear optimization problem. We show that traditional flow-level rate controllers result in sub-optimal sharing results between the different overlays. We design efficient and distributed overlay flows control algorithms and demonstrate the effectiveness of our design Yang Richard Yang, Arvind Krishnamurthy |
IWQoS | 4 |
| 2006 | PCP: Efficient Endpoint Congestion Control
Thomas E. Anderson, Andy Collins, Arvind Krishnamurthy, John Zahorjan |
NSDI | 3 |
| 2006 | iPlane: An Information Plane for Distributed Services
Harsha V. Madhyastha, Tomas Isdal, Michael Piatek, Colin Dixon, Thomas E. Anderson, Arvind Krishnamurthy, Arun Venkataramani |
OSDI | 6 |
| 2005 | Network localization in partially localizable networksabstractKnowing the positions of the nodes in a network is essential to many next generation pervasive and sensor network functionalities. Although many network localization systems have recently been proposed and evaluated, there has been no systematic study of partially localizable networks, i.e., networks in which there exist nodes whose positions cannot be uniquely determined. There is no existing study which correctly identifies precisely which nodes in a network are uniquely localizable and which are not. This absence of a sufficient uniqueness condition permits the computation of erroneous positions that may in turn lead applications to produce flawed results. In this paper, in addition to demonstrating the relevance of networks that may not be fully localizable, we design the first framework for two dimensional network localization with an efficient component to correctly determine which nodes are localizable and which are not. Implementing this system, we conduct comprehensive evaluations of network localizability, providing guidelines for both network design and deployment. Furthermore, we study an integration of traditional geographic routing with geographic routing over virtual coordinates in the partially localizable network setting. We show that this novel cross-layer integration yields good performance, and argue that such optimizations will be likely be necessary to ensure acceptable application performance in partially localizable networks. David Kiyoshi Goldenberg, Arvind Krishnamurthy, Wesley C. Maness, Yang Richard Yang, Anthony Young, A. Stephen Morse, Andreas Savvides, Brian D. O. Anderson |
INFOCOM | 2 |
| 2005 | Combining Flexibility and Scalability in a Peer-to-Peer Publish/Subscribe System
Chi Zhang 0070, Arvind Krishnamurthy, Randolph Y. Wang, Jaswinder Pal Singh |
Middleware | 2 |
| 2005 | A collision model for randomized routing in fat-tree networks
Volker Strumpen, Arvind Krishnamurthy |
J. Parallel Distributed Comput. | 2 |
| 2005 | Bridging the digital divide: storage media + postal network = generic high-bandwidth communicationabstractMaking high-bandwidth Internet access pervasively available to a large worldwide audience is a difficult challenge, especially in many developing regions. As we wait for the uncertain takeoff of technologies that promise to improve the situation, we propose to explore an approach that is potentially more easily realizable: the use of digital storage media transported by the postal system as a general digital communication mechanism. We shall call such a system a Postmanet . Compared to more conventional wide-area connectivity options, the Postmanet has several important advantages, including wide global reach, great bandwidth potential, low cost, and ease of incremental adoption. While the idea of sending digital content via the postal system is not a new one, none of the existing attempts have turned the postal system into a generic and transparent communication channel that not only can cater to a wide array of applications, but also effectively manage the many idiosyncrasies associated with using the postal system. In the proposed Postmanet, we see two recurring themes at many different levels of the system. One is the simultaneous exploitation of the Internet and the postal system so we can combine their latency and bandwidth advantages. The other is the exploitation of the abundant capacity and bandwidth of the Postmanet to improve its latency, cost, and reliability. Sumeet Sobti, Junwen Lai, Fengzhou Zheng, Kai Li 0001, Randolph Y. Wang, Arvind Krishnamurthy |
ACM Trans. Storage | 7 |
| 2004 | Segank: A Distributed Mobile Storage System
Sumeet Sobti, Fengzhou Zheng, Junwen Lai, Yilei Shao, Chi Zhang 0070, Elisha Ziskind, Arvind Krishnamurthy, Randolph Y. Wang |
FAST | 8 |
| 2004 | Highly Secure and Efficient RoutingabstractIn this paper, we consider the problem of routing in an adversarial environment, where a sophisticated adversary has penetrated arbitrary parts of the routing infrastructure and attempts to disrupt routing. We present protocols that are able to route packets as long as at least one nonfaulty path exists between the source and the destination. These protocols have low communication overhead, low processing requirements, low incremental cost, and fast fault detection. We also present extensions to the protocols that penalize adversarial routers by blocking their traffic. Ioannis C. Avramopoulos, Hisashi Kobayashi, Randy Wang, Arvind Krishnamurthy |
INFOCOM | 4 |
| 2004 | Overlay Mesh Construction Using Interleaved Spanning TreesabstractIn this paper we evaluate a method of using interleaved spanning trees to compose a resilient, high performance overlay mesh. Though spanning trees of arbitrary type could be used to construct an overlay mesh, we focus on a distributed algorithm that computes k minimum spanning trees on an arbitrary graph. The principal motivation behind this strategy is to provide applications with a k-redundant, high quality mesh suitable for demanding applications like A/V broadcast, video conferencing, data collection, multi-path routing, and file mirroring/transfer. We elaborate details of k-MST, pointing out advantages and potential problem points of the protocol, and then analyze its performance using a variety of metrics with simulation as well as a functional PlanetLab implementation. Anthony Young, Arvind Krishnamurthy, Larry L. Peterson, Randy Wang |
INFOCOM | 4 |
| 2004 | Network-Embedded Programmable Storage and Its Applications
Sumeet Sobti, Junwen Lai, Yilei Shao, Chi Zhang 0070, Ming Zhang 0005, Fengzhou Zheng, Arvind Krishnamurthy, Randolph Y. Wang |
NETWORKING | 8 |
| 2004 | Managing a portfolio of overlay pathsabstractIn recent years, several architectures have been proposed and developed for supporting streaming applications that take advantage of multiple paths through the network simultaneously. We consider the problem of computing a set of paths and the relative amounts of data conveyed through them in order to provide the desired level of performance for data streams. Given the expectation, variance, and covariance of an appropriate metric of interest for overlay links, we attempt to solve the underlying resource allocation problem by applying methods used in managing a finance portfolio. We observe that the flow allocation problem requires constrained application of these methods, and we discuss the tractability of enforcing the constraints. We finally present some simulation results to evaluate the effectiveness of our proposed techniques. Daria Antonova, Arvind Krishnamurthy, Ravi Sundaram |
NOSSDAV | 2 |
| 2004 | Load balancing and locality in range-queriable data structuresabstractWe describe a load-balancing mechanism for assigning elements to servers in a distributed data structure that supports range queries. The mechanism ensures both load-balancing with respect to an arbitrary load measure specified by the user and geographical locality, assigning elements with similar keys to the same server. Though our mechanism is specifically designed to improve the performance of skip graphs, it can be adapted to provide deterministic, locality-preserving load-balancing to any distributed data structure that orders machines in a ring or line. James Aspnes, Jonathan Kirsch, Arvind Krishnamurthy |
PODC | 3 |
| 2004 | Turning the postal system into a generic digital communication mechanismabstractThe phenomenon that rural residents and people with low incomes lag behind in Internet access is known as the "digital divide." This problem is particularly acute in developing countries, where most of the world's population lives. Bridging this digital divide, especially by attempting to increase the accessibility of broadband connectivity, can be challenging. The improvement of wide-area connectivity is constrained by factors such as how quickly we can dig ditches to bury fibers in the ground; and the cost of furnishing "last-mile" wiring can be prohibitively high.In this paper, we explore the use of digital storage media transported by the postal system as a general digital communication mechanism. While some companies have used the postal system to deliver software and movies, none of them has turned the postal system into a truly generic digital communication medium supporting a wide variety of applications. We call such a generic system a Postmanet. Compared to traditional wide-area connectivity options, the Postmanet has several important advantages, including wide global reach, great bandwidth potential and low cost.Manually preparing mobile storage devices for shipment may appear deceptively simple, but with many applications, communicating parties and messages, manual management becomes infeasible, and systems support at several levels becomes necessary. We explore the simultaneous exploitation of the Internet and the Postmanet, so we can combine their latency and bandwidth advantages to enable sophisticated bandwidth-intensive applications. Randolph Y. Wang, Sumeet Sobti, Elisha Ziskind, Junwen Lai, Arvind Krishnamurthy |
SIGCOMM | 6 |
| 2004 | A Transport Layer Approach for Improving End-to-End Performance and Robustness Using Redundant Paths
Ming Zhang 0005, Junwen Lai, Arvind Krishnamurthy, Larry L. Peterson, Randolph Y. Wang |
USENIX ATC, General Track | 3 |
| 2003 | Modeling Hard-Disk Power Consumption
John Zedlewski, Sumeet Sobti, Fengzhou Zheng, Arvind Krishnamurthy, Randolph Y. Wang |
FAST | 5 |
| 2003 | Approximation and collusion in multicast cost sharingabstractNo abstract available. Joan Feigenbaum, Arvind Krishnamurthy, Rahul Sami, Scott Shenker |
EC | 2 |
| 2003 | Hardness results for multicast cost sharing
Joan Feigenbaum, Arvind Krishnamurthy, Rahul Sami, Scott Shenker |
Theor. Comput. Sci. | 2 |
| 2002 | PersonalRAID: Mobile Storage for Distributed and Disconnected Computers
Sumeet Sobti, Chi Zhang 0070, Arvind Krishnamurthy, Randolph Y. Wang |
FAST | 5 |
| 2002 | Configuring and Scheduling an Eager-Writing Disk Array for a Transaction Processing Workload
Chi Zhang 0070, Arvind Krishnamurthy, Randolph Y. Wang |
FAST | 3 |
| 2002 | Hardness Results for Multicast Cost Sharing
Joan Feigenbaum, Arvind Krishnamurthy, Rahul Sami, Scott Shenker |
FSTTCS | 2 |
| 2002 | Probabilistic Packet Scheduling: Achieving Proportional Share Bandwidth Allocation for TCP FlowsabstractThis paper describes and evaluates a probabilistic packet scheduling (PPS) algorithm for providing different levels of service to TCP flows. With our approach, each router defines a local currency in terms of tickets and assigns tickets to its inputs based on contractual agreements with its upstream routers. A flow is tagged with tickets to represent the relative share of bandwidth it should receive at each link. When multiple flows share the same bottleneck, the bandwidth that each flow obtains is proportional to the relative tickets assigned to that flow. Simulations show that PPS does a better job of proportionally allocating bandwidth than DiffServ and weighted CSFQ. In addition, PPS accommodates flows that cross multiple currency domains. Ming Zhang 0005, Randolph Y. Wang, Larry L. Peterson, Arvind Krishnamurthy |
INFOCOM | 4 |
| 2001 | Approximation and collusion in multicast cost sharing (extended abstract)abstractArticle Share on Approximation and collusion in multicast cost sharing (extended abstract) Authors: J. Feigenbaum Yale University, New Haven, CT Yale University, New Haven, CTView Profile , A. Krishnamurthy Yale University, New Haven, CT Yale University, New Haven, CTView Profile , R. Sami Yale University, New Haven, CT Yale University, New Haven, CTView Profile , S. Shenker ACIRI/ICSI, Berkeley, CA ACIRI/ICSI, Berkeley, CAView Profile Authors Info & Claims EC '01: Proceedings of the 3rd ACM conference on Electronic CommerceOctober 2001 Pages 253–255https://doi.org/10.1145/501158.501190Online:14 October 2001Publication History 9citation174DownloadsMetricsTotal Citations9Total Downloads174Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Joan Feigenbaum, Arvind Krishnamurthy, Rahul Sami, Scott Shenker |
EC | 2 |
| 2000 | Trading Capacity for Performance in a Disk Array
Benjamin Gum, Yuqun Chen, Randolph Y. Wang, Kai Li 0001, Arvind Krishnamurthy, Thomas E. Anderson |
OSDI | 6 |
| 1998 | Modeling Communication Pipeline LatencyabstractIn this paper, we study how to minimize the latency of a message through a network that consists of a number of store-and-forward stages. This research is especially relevant for today's low overhead communication systems that employ dedicated processing elements for protocol processing. We develop an abstract pipeline model that reveals a crucial performance tradeoff involving the effects of the overhead of the bottleneck stage and the bandwidth of the remaining stages. We exploit this tradeoff to develop a suite of fragmentation algorithms designed to minimize message latency. We also provide an experimental methodology that enables the construction of customized pipeline algorithms that can adapt to the specific system characteristics and application workloads. By applying this methodology to the Myrinet-GAM system, we have improved its latency by up to 51%. Our theoretical framework is also applicable to pipelined systems beyond the context of high speed networks. Randolph Y. Wang, Arvind Krishnamurthy, Richard P. Martin, Thomas E. Anderson, David E. Culler |
SIGMETRICS | 2 |
| 1998 | Titanium: A High-performance Java DialectabstractTitanium is a language and system for high-performance parallel scientific computing. Titanium uses Java as its base, thereby leveraging the advantages of that language and allowing us to focus attention on parallel computing issues. The main additions to Java are immutable classes, multidimensional arrays, an explicitly parallel SPMD model of computation with a global address space, and zone-based memory management. We discuss these features and our design approach, and report progress on the development of Titanium, including our current driving application: a three-dimensional adaptive mesh refinement parallel Poisson solver. © 1998 John Wiley & Sons, Ltd. Katherine A. Yelick, Luigi Semenzato, Geoff Pike, Carleton Miyamoto, Ben Liblit, Arvind Krishnamurthy, Paul N. Hilfinger, Susan L. Graham, David Gay, Phillip Colella, Alex Aiken |
Concurr. Pract. Exp. | 6 |
| 1996 | Evaluation of Architectural Support for Global Address-Based Communication in Large-Scale Parallel MachinesabstractLarge-scale parallel machines are incorporating increasingly sophisticated architectural support for user-level messaging and global memory access. We provide a systematic evaluation of a broad spectrum of current design alternatives based on our implementations of a global address language on the Thinking Machines CM-5, Intel Paragon, Meiko CS-2, Cray T3D, and Berkeley NOW. This evaluation includes a range of compilation strategies that make varying use of the network processor; each is optimized for the target architecture and the particular strategy. We analyze a family of interacting issues that determine the performance trade-offs in each implementation, quantify the resulting latency, overhead, and bandwidth of the global access operations, and demonstrate the effects on application performance. Arvind Krishnamurthy, Klaus E. Schauser, Chris J. Scheiman, Randolph Y. Wang, David E. Culler, Katherine A. Yelick |
ASPLOS | 1 |
| 1996 | Analyses and Optimizations for Shared Address Space Programs
Arvind Krishnamurthy, Katherine A. Yelick |
J. Parallel Distributed Comput. | 1 |
| 1995 | Empirical Evaluation of the CRAY-T3D: A Compiler PerspectiveabstractMost recent MPP systems employ a fast microprocessor surrounded by a shell of communication and synchronization logic. The CRAY-T3D provides an elaborate shell to support global-memory access, prefetch, atomic operations, barriers, and block transfers. We provide a detailed empirical performance characterization of these primitives using micro-benchmarks and evaluate their utility in compiling for a parallel language. We have found that the raw performance of the machine is quite impressive and the most effective forms of communication are prefetch and write. Other shell provisions, such as the bulk transfer engine and the external Annex register set, are cumbersome and of little use. By evaluating the system in the context of a language implementation, we shed light on important trade-offs and pitfalls in the machine architecture. Remzi H. Arpaci-Dusseau, David E. Culler, Arvind Krishnamurthy, Steve G. Steinberg, Katherine A. Yelick |
ISCA | 3 |
| 1995 | Optimizing Parallel Programs with Explicit SynchronizationabstractWe present compiler analyses and optimizations for explicitly parallel programs that communicate through a shared address space. Any type of code motion on explicitly parallel programs requires a new kind of analysis to ensure that operations reordered on one processor cannot be observed by another. The analysis, based on work by Shasha and Snir, checks for cycles among interfering accesses. We improve the accuracy of their analysis by using additional information from post-wait synchronization, barriers, and locks. Arvind Krishnamurthy, Katherine A. Yelick |
PLDI | 1 |
| 1995 | Towards Modeling the Performance of a Fast Connected Components Algorithm on Parallel Machinesabstract: We present and analyze a portable, high-performance algorithm for finding connected components on modern distributed memory multiprocessors. The algorithm is a hybrid of the classic DFS on the subgraph local to each processor and a variant of the Shiloach-Vishkin PRAM algorithm on the global collection of subgraphs. We implement the algorithm in Split-C and measure performance on the the Cray T3D, the Meiko CS-2, and the Thinking Machines CM-5 using a class of graphs derived from cluster dynamics methods in computational physics. On a 256 processor Cray T3D, the implementation outperforms all previous solutions by an order of magnitude. A characterization of graph parameters allows us to select graphs that highlight key performance features. We study the effects of these parameters and machine characteristics on the balance of time between the local and global phases of the algorithm and find that edge density, surface-to-volume ratio, and relative communication cost dominate perform... Steven S. Lumetta, Arvind Krishnamurthy, David E. Culler |
SC | 2 |
| 1993 | Parallel programming in Split-CabstractNo abstract available. David E. Culler, Andrea C. Arpaci-Dusseau, Seth Copen Goldstein, Arvind Krishnamurthy, Steven S. Lumetta, Thorsten von Eicken, Katherine A. Yelick |
SC | 4 |