Ashish Raniwala

dblp:32/3555 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 3 first-authorSoftware engineering, systems software and programming languages · 4 · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2024 SmartOClock: Workload- and Risk-Aware Overclocking in the Cloud
abstract
Operating server components beyond their voltage and power design limit (i.e., overclocking) enables improving performance and lowering cost for cloud workloads. However, overclocking can significantly degrade component lifetime, increase power draw, and cause power capping events, eventually diminishing the performance benefits. In this paper, we characterize the impact of overclocking on cloud workloads by studying their profiles from production deployments. Based on the characterization insights, we propose SmartOClock, the first distributed overclocking management platform specifically designed for cloud environments. SmartOClock is a workload-aware scheme that relies on power predictions to heterogeneously distribute the power budgets across its servers based on their needs and then enforce budget compliance locally, per-server, in a decentralized manner. SmartOClock reduces the tail latency by 9%, application cost by 30% and total energy consumption by 10% for latencysensitive microservices on a 36-server deployment. Simulation analysis using production traces show that SmartOClock reduces the number of power capping events by up to 95% while increasing the overclocking success rate by up to 62%. We also describe lessons from building a first-of-its-kind overclockable cluster in Microsoft Azure for production experiments.
Jovan Stojkovic, Pulkit A. Misra, Íñigo Goiri, Sam Whitlock, Esha Choukse, Mayukh Das, Chetan Bansal, Zoey Sun, Haoran Qiu, Reed Zimmermann, Savyasachi Samal, Brijesh Warrier, Ashish Raniwala, Ricardo Bianchini
ISCA14
2023 Hyrax: Fail-in-Place Server Operation in Cloud Platforms
Jialun Lyu, Marisa You, Celine Irvene, Mark Jung, Tyler Narmore, Jacob Shapiro, Luke Marshall, Savyasachi Samal, Ioannis Manousakis, Lisa Hsu, Preetha Subbarayalu, Ashish Raniwala, Brijesh Warrier, Ricardo Bianchini, Bianca Schroeder, Daniel S. Berger
OSDI12
2021 Cost-Efficient Overclocking in Immersion-Cooled Datacenters
abstract
Cloud providers typically use air-based solutions for cooling servers in datacenters. However, increasing transistor counts and the end of Dennard scaling will result in chips with thermal design power that exceeds the capabilities of air cooling in the near future. Consequently, providers have started to explore liquid cooling solutions (e.g., cold plates, immersion cooling) for the most power-hungry workloads. By keeping the servers cooler, these new solutions enable providers to operate server components beyond the normal frequency range (i.e., overclocking them) all the time. Still, providers must tradeoff the increase in performance via overclocking with its higher power draw and any component reliability implications.In this paper, we argue that two-phase immersion cooling (2PIC) is the most promising technology, and build three prototype 2PIC tanks. Given the benefits of 2PIC, we characterize the impact of overclocking on performance, power, and reliability. Moreover, we propose several new scenarios for taking advantage of overclocking in cloud platforms, including oversubscribing servers and virtual machine (VM) auto-scaling. For the auto-scaling scenario, we build a system that leverages overclocking for either hiding the latency of VM creation or postponing the VM creations in the hopes of not needing them. Using realistic cloud workloads running on a tank prototype, we show that overclocking can improve performance by 20%, increase VM packing density by 20%, and improve tail latency in auto-scaling scenarios by 54%. The combination of 2PIC and overclocking can reduce platform cost by up to 13% compared to air cooling.
Majid Jalili 0004, Ioannis Manousakis, Íñigo Goiri, Pulkit A. Misra, Ashish Raniwala, Husam Alissa, Bharath Ramakrishnan, Phillip Tuma, Christian Belady, Marcus Fontoura, Ricardo Bianchini
ISCA5
2021 Analyzing and Mitigating Data Stalls in DNN Training
abstract
Training Deep Neural Networks (DNNs) is resource-intensive and time-consuming. While prior research has explored many different ways of reducing DNN training time, the impact of input data pipeline , i.e., fetching raw data items from storage and performing data pre-processing in memory, has been relatively unexplored. This paper makes the following contributions: (1) We present the first comprehensive analysis of how the input data pipeline affects the training time of widely-used computer vision and audio Deep Neural Networks (DNNs), that typically involve complex data pre-processing. We analyze nine different models across three tasks and four datasets while varying factors such as the amount of memory, number of CPU threads, storage device, GPU generation etc on servers that are a part of a large production cluster at Microsoft. We find that in many cases, DNN training time is dominated by data stall time : time spent waiting for data to be fetched and pre-processed. (2) We build a tool, DS-Analyzer to precisely measure data stalls using a differential technique, and perform predictive what-if analysis on data stalls. (3) Finally, based on the insights from our analysis, we design and implement three simple but effective techniques in a data-loading library, CoorDL, to mitigate data stalls. Our experiments on a range of DNN tasks, models, datasets, and hardware configs show that when PyTorch uses CoorDL instead of the state-of-the-art DALI data loading library, DNN training time is reduced significantly (by as much as 5X on a single server).
Jayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay Chidambaram
Proc. VLDB Endow.3
2010 FlumeJava: easy, efficient data-parallel pipelines
abstract
MapReduce and similar systems significantly ease the task of writing data-parallel code. However, many real-world computations require a pipeline of MapReduces, and programming and managing such pipelines can be difficult. We present FlumeJava, a Java library that makes it easy to develop, test, and run efficient data-parallel pipelines. At the core of the FlumeJava library are a couple of classes that represent immutable parallel collections, each supporting a modest number of operations for processing them in parallel. Parallel collections and their operations present a simple, high-level, uniform abstraction over different data representations and execution strategies. To enable parallel operations to run efficiently, FlumeJava defers their evaluation, instead internally constructing an execution plan dataflow graph. When the final results of the parallel operations are eventually needed, FlumeJava first optimizes the execution plan, and then executes the optimized operations on appropriate underlying primitives (e.g., MapReduces). The combination of high-level abstractions for parallel data and computation, deferred evaluation and optimization, and efficient parallel primitives yields an easy-to-use system that approaches the efficiency of hand-optimized pipelines. FlumeJava is in active use by hundreds of pipeline developers within Google.
Craig Chambers, Ashish Raniwala, Frances Perry, Stephen Adams 0001, Robert R. Henry 0001, Robert Bradshaw, Nathan Weizenbaum
PLDI2
2009 Globally fair radio resource allocation for wireless mesh networks
abstract
Network flows running on a wireless mesh network (WMN) may suffer from partial failures in the form of serious throughput degradation, sometimes to the extent of starvation, because of weaknesses in the underlying MAC protocol, dissimilar physical transmission rates or different degrees of local congestion. Most existing WMN transport protocols fail to take these factors into account. This paper describes the design, implementation and evaluation of a coordinated congestion control (C3L) algorithm that guarantees fair resource allocation under adverse scenarios and thus provides end-to-end max-min fairness among competing flows. The C3L algorithm features an advanced topology discovery mechanism that detects the inhibition of wireless communication links, and a general collision domain capacity re-estimation mechanism that effectively addresses such inhibition. A comprehensive ns-2-based simulation study as well as empirical measurements taken from an IEEE 802.11a-based multi-hop wireless testbed demonstrate that the C3L algorithm greatly improves inter-flow fairness, eliminates the starvation problem, and at the same time maintains high radio resource utilization efficiency.
Ashish Raniwala, Pradipta De, Srikant Sharma, Rupa Krishnan, Tzi-cker Chiueh
MASCOTS1
2008 Design of a Channel Characteristics-Aware Routing Protocol
abstract
Radio channel quality of real-world wireless networks tends to exhibit both short-term and long-term temporal variations that are in general difficult to model. To maximize the utilization efficiency of radio resources, it is critical that these temporal fluctuations in radio signal quality be incorporated into wireless routing decisions. In this paper, we explore the design considerations in leveraging accurate real-time radio channel quality information when making routing decisions. Specifically, we propose a channel characteristics-aware routing protocol (CARP) that (1) uses per-packet transmission time to estimate the effective residual capacity of a wireless link, (2) employs a bandwidth probability distribution model to better approximate a wireless path's capacity profile, and (3) applies multi-path routing to exploit diversity among alternative paths and deliver more robust throughputs despite temporal fluctuations in wireless link quality. We evaluated the performance gains of incorporating each of these mechanisms on a miniaturized multi-hop wireless network testbed- MiNT-m.
Rupa Krishnan, Ashish Raniwala, Tzi-cker Chiueh
INFOCOM2
2007 End-to-End Flow Fairness Over IEEE 802.11-Based Wireless Mesh Networks
abstract
Economies of scale make IEEE 802.11 an attractive technology for building wireless mesh networks (WMNs). However, the IEEE 802.11 protocol exhibits serious link-layer unfairness when used in multi-hop networks. Existing fairness solutions either do not address this problem, or require proprietary MAC protocol to provide fairness. In this paper, we argue that an ideal transport protocol should be able to achieve fairness even on top of an unfair MAC layer such as 802.11. Towards this end, we propose a co-ordinatedcongestioncontrolalgorithm that performs global bandwidth allocation and provides end-to-end flow-level max-min fairness despite weaknesses in the MAC layer. The proposed algorithm features an advanced topology discovery mechanism that detects the inhibition of wireless communication links, and a general collision domain capacity re-estimation mechanism that effectively addresses such inhibition. Through an ns-2-based simulation study we demonstrate that the proposed algorithm substantially improves the fairness across flows, eliminates starvation problem, and simultaneously maintains a high overall network throughput.
Ashish Raniwala, Pradipta De, Srikant Sharma, Rupa Krishnan, Tzi-cker Chiueh
INFOCOM1
2007 Evaluation of a Stateful Transport Protocol for Multi-channel Wireless Mesh Networks
abstract
An effective transport protocol for a wireless mesh network (WMN) must fairly and efficiently allocate the limited network resources among multiple flows sharing the network while minimizing the performance overhead it incurs. While many transport protocols have been proposed specifically for multi-hop wireless networks, most of them refrain from keeping state in the intermediate network nodes. In this paper, we focus on the other extreme of the design space:statefultransportprotocol, and study the research question of how much performance improvement is possible if intermediate network nodes could maintain as much state as needed. We present the design of a stateful transport protocol, namedlink-awarereliabletransportprotocol(LRTP), and examine how LRTP can fairly and efficiently allocate the network resources by accurately estimating the sending rate of each flow traversing the network using information about effective physical link capacity and the number of sharing flows. LRTP reduces the performance overhead associated with reliable packet delivery by leveraging the link-layer retransmission mechanism to eliminate per-packet end-to-end acknowledgments and unnecessary packet transmissions. Experiments conducted on anIEEE802.1la-basedmulti-channelwirelessmeshnetworktestbedas well asns-2simulationsdemonstrate that LRTP can achieve significant improvements in both overall network throughput and inter-flow fairness, especially on wireless networks with channel errors, when compared with the de facto Internet transport protocol TCP, and state-of-the-art MANET transport protocols such as ATP.
Ashish Raniwala, Srikant Sharma, Pradipta De, Rupa Krishnan, Tzi-cker Chiueh
IWQoS1
2006 MiNT-m: an autonomous mobile wireless experimentation platform
abstract
Limited fidelity of software-based wireless network simulations has prompted many researchers to build testbeds for developing and evaluating their wireless protocols and mobile applications. Since most testbeds are tailored to the needs of specific research projects, they cannot be easily reused for other research projects that may have different requirements on physical topology, radio channel characteristics or mobility pattern. In this paper, we describe the design, implementation and evaluation of MiNT-m, an experimentation platform devised specifically to support arbitrary experiments for mobile multi-hop wireless network protocols. In addition to inheriting the miniaturization feature from its predecessor MiNT [9], MiNT-m enables flexible testbed reconfiguration on an experiment-by-experiment basis by putting each testbed node on a centrally controlled untethered mobile robot. To support mobility and reconfiguration of testbed nodes, MiNT-m includes a scalable mobile robot navigation control subsystem, which in turn consists of a vision-based robot positioning module and a collision avoidance-based trajectory planning module. Further, MiNT-m provides a comprehensive network/experiment management subsystem that affords a user full interactive control over the testbed as well as real-time visualization of the testbed activities. Finally, because MiNT-m is designed to be a shared research infrastructure that supports 24x7 operation, it incorporates a novel automatic battery recharging capability that enables testbed robots to operate without human intervention for weeks.
Pradipta De, Ashish Raniwala, Rupa Krishnan, Krishna Tatavarthi, Jatan Modi, Nadeem Ahmed Syed, Srikant Sharma, Tzi-cker Chiueh
MobiSys2
2005 MiNT: a miniaturized network testbed for mobile wireless research
abstract
Most mobile wireless networking research today relies on simulations. However, fidelity of simulation results has always been a concern, especially when the protocols being studied are affected by the propagation and interference characteristics of the radio channels. Inherent difficulty in faithfully modeling the wireless channel characteristics has encouraged several researchers to build wireless network testbeds. A full-fledged wireless testbed is spread over a large physical space because of the wide coverage area of radio signals. This makes a large-scale testbed difficult and expensive to set up, configure, and manage. This paper describes a miniaturized 802.11b-based, multi-hop wireless network testbed called MiNT. MiNT occupies a significantly small space, and dramatically reduces the efforts required in setting up a multi-hop wireless network used for wireless application/protocol testing and evaluation. MiNT is also a hybrid simulation platform that can execute ns-2 simulation scripts with the link, MAC and physical layer in the simulator replaced by real hardware. We demonstrate the fidelity of MiNT by comparing experimental results on it with similar experiments conducted on a non-miniaturized testbed. We also compare the results of experiments conducted using hybrid simulation on MiNT with those obtained using pure simulation. Finally, using a case study we show the usefulness of MiNT in wireless application testing and evaluation.
Pradipta De, Ashish Raniwala, Srikant Sharma, Tzi-cker Chiueh
INFOCOM2
2005 Architecture and algorithms for an IEEE 802.11-based multi-channel wireless mesh network
abstract
Even though multiple non-overlapped channels exist in the 2.4 GHz and 5 GHz spectrum, most IEEE 802.11-based multi-hop ad hoc networks today use only a single channel. As a result, these networks rarely can fully exploit the aggregate bandwidth available in the radio spectrum provisioned by the standards. This prevents them from being used as an ISP's wireless last-mile access network or as a wireless enterprise backbone network. In this paper, we propose a multi-channel wireless mesh network (WMN) architecture (called Hyacinth) that equips each mesh network node with multiple 802.11 network interface cards (NICs). The central design issues of this multi-channel WMN architecture are channel assignment and routing. We show that intelligent channel assignment is critical to Hyacinth's performance, present distributed algorithms that utilize only local traffic load information to dynamically assign channels and to route packets, and compare their performance against a centralized algorithm that performs the same functions. Through an extensive simulation study, we show that even with just 2 NICs on each node, it is possible to improve the network throughput by a factor of 6 to 7 when compared with the conventional single-channel ad hoc network architecture. We also describe and evaluate a 9-node Hyacinth prototype that Is built using commodity PCs each equipped with two 802.11a NICs.
Ashish Raniwala, Tzi-cker Chiueh
INFOCOM1
1999 Phoenix: a low-power fault-tolerant real-time network-attached storage device
abstract
Phoenix is a real-time network-attached storage device (NASD) that guarantees real-time data delivery to network clients even across single disk failure. The service interfaces that Phoenix provides are best-effort/real-time reads/writes based on unique object identifiers and block offsets. Data retrieval from Phoenix can be serviced in server push or client pull modes. Phoenix's real-time disk subsystem performance results from a standard cycle-based scan-order disk scheduling mechanism. However, the disk I/O cycle of Phoenix is either completely active or completely idle. This on-off disk scheduling model effectively reduces the power consumption of the disk subsystem, without increasing the buffer size requirement. Phoenix also exploits unused disk storage space and maintains additional redundancy beyond the generic RAID5-style parity. This extra redundancy, typically in the form of block replication, reduces the time to reconstruct the data on the failed disk. This paper describes the design, implementation, and evaluation of Phoenix, one of the first, if not the first, NASDs that support fault-tolerant, real-time, and low-power network storage service.
Anindya Neogi, Ashish Raniwala, Tzi-cker Chiueh
ACM Multimedia (1)2