A. L. Narasimha Reddy

dblp:r/ALNarasimhaReddy · also Narasimha Reddy Annapareddy · DBLP profile ↗
← Back
89ranked-venue papers
15as first author
10since 2021 · last 2025
0000-0003-4625-8819ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 50 · 12 first-author · 8 since 2021Computer networks · 30 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-author · 1 since 2021Security and privacy · 2Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Storage access optimization for efficient GPU-centric information retrieval
abstract
Abstract The rapid growth of AI/ML workloads has outpaced the capabilities of CPU-centric architectures to deliver the required data throughput and compute efficiency. This paper introduces a GPU-centric architecture leveraging GPUDirect Storage (GDS) to transfer data directly from SSDs to GPU memory, bypassing CPU bottlenecks and enabling high-throughput data paths. We propose Embedding from Storage Pipelined Network (ESPN) and its extension, ESPN-LIVE, which employ optimizations like data prefetching and on-demand embedding generation to align storage latency with GPU throughput. Experiments show ESPN reduces query latency by up to $$3.9\times$$ 3.9 × , cuts memory usage by up to $$16\times$$ 16 × , and improves throughput by up to 68%. ESPN-LIVE eliminates the need to store multi-vector embeddings by dynamically computing document representations, reducing storage costs by up to $$16\times$$ 16 × , and making it particularly effective for single-query systems. These results highlight the potential of SSD-GPU integration for scalable, high-performance AI/ML workloads in information retrieval and LLM applications.
Susav Shrestha, Aayush Gautam, A. L. Narasimha Reddy
J. Supercomput.3
2024 ESPN: Memory-Efficient Multi-vector Information Retrieval
abstract
Recent advances in large language models have demonstrated remarkable effectiveness in information retrieval (IR) tasks. While many neural IR systems encode queries and documents into single-vector representations, multi-vector models elevate the retrieval quality by producing multi-vector representations and facilitating similarity searches at the granularity of individual tokens. However, these models significantly amplify memory requirements for retrieval indices by an order of magnitude. This escalation in index size renders the scalability of multi-vector IR models progressively challenging due to their substantial memory demands. We introduce Embedding from Storage Pipelined Network (ESPN) where we offload the entire re-ranking embedding tables to SSDs and reduce the memory requirements by 5−16×. We design a flexible software prefetcher applicable to any hierarchical clustering based search, achieving hit rates exceeding 90%. ESPN improves SSD based retrieval up to 6.4× and end-to-end throughput by 68% to maintain near-memory levels of query latency even for large query batch sizes. The code is available at https://github.com/susavlsh10/ESPN-v1.
Susav Shrestha, A. L. Narasimha Reddy, Zongwang Li
ISMM2
2024 WannaLaugh: A Configurable Ransomware Emulator - Learning to Mimic Malicious Storage Traces
abstract
Ransomware, a fearsome and an evolving cybersecurity threat, continues to inflict severe consequences on individuals and organizations worldwide. Traditional detection methods, reliant on static signatures and application behavioral patterns, are challenged by the dynamic nature of these threats. This paper introduces two primary contributions to address this challenge. First, we introduce the WannaLaugh ransomware emulator. This tool is designed to safely mimic ransomware attacks without causing actual harm or spreading malware, making it a unique solution for studying ransomware behavior. Second, we show how this emulator can be used to mimic the I/O behavior of existing ransomware. Experimental results show that WannaLaugh can mimic six real ransomware with high accuracy. Both the emulator and its mimicking application aim to represent significant steps forward in ransomware detection in the era of machine-learning-driven cybersecurity.
Dionysios Diamantopoulos, Roman A. Pletka, Slavisa Sarafijanovic, A. L. Narasimha Reddy, Haralampos Pozidis
SYSTOR4
2023 KVRangeDB: Range Queries for a Hash-based Key-Value Device
abstract
Key–value (KV) software has proven useful to a wide variety of applications including analytics, time-series databases, and distributed file systems. To satisfy the requirements of diverse workloads, KV stores have been carefully tailored to best match the performance characteristics of underlying solid-state block devices. Emerging KV storage device is a promising technology for both simplifying the KV software stack and improving the performance of persistent storage-based applications. However, while providing fast, predictable put and get operations, existing KV storage devices do not natively support range queries that are critical to all three types of applications described above. In this article, we present KVRangeDB, a software layer that enables processing range queries for existing hash-based KV solid-state disks (KVSSDs). As an effort to adapt to the performance characteristics of emerging KVSSDs, KVRangeDB implements log-structured merge tree key index that reduces compaction I/O, merges keys when possible, and provides separate caches for indexes and values. We evaluated the KVRangeDB under a set of representative workloads, and compared its performance with two existing database solutions: a Rocksdb variant ported to work with the KVSSD, and Wisckey, a key–value database that is carefully tuned for conventional block devices. On filesystem aging workloads, KVRangeDB outperforms Wisckey by 23.7× in terms of throughput and reduce CPU usage and external write amplifications by 14.3× and 9.8×, respectively.
Qing Zheng, Jason Lee 0004, Bradley W. Settlemyer, Fei Wen 0003, A. L. Narasimha Reddy, Paul Gratz
ACM Trans. Storage6
2022 Reducing Minor Page Fault Overheads through Enhanced Page Walker
abstract
Application virtual memory footprints are growing rapidly in all systems from servers down to smartphones. To address this growing demand, system integrators are incorporating ever larger amounts of main memory, warranting rethinking of memory management. In current systems, applications produce page fault exceptions whenever they access virtual memory regions that are not backed by a physical page. As application memory footprints grow, they induce more and more minor page faults. Handling of each minor page fault can take a few thousands of CPU cycles and blocks the application till the OS kernel finds a free physical frame. These page faults can be detrimental to the performance when their frequency of occurrence is high and spread across application runtime. Specifically, lazy allocation-induced minor page faults are increasingly impacting application performance. Our evaluation of several workloads indicates an overhead due to minor page faults as high as 29% of execution time. In this article, we propose to mitigate this problem through a hardware, software co-design approach. Specifically, we first propose to parallelize portions of the kernel page allocation to run ahead of fault time in a separate thread. Then we propose the Minor Fault Offload Engine (MFOE), a per-core hardware accelerator for minor fault handling. MFOE is equipped with a pre-allocated page frame table that it uses to service a page fault. On a page fault, MFOE quickly picks a pre-allocated page frame from this table, makes an entry for it in the TLB, and updates the page table entry to satisfy the page fault. The pre-allocation frame tables are periodically refreshed by a background kernel thread, which also updates the data structures in the kernel to account for the handled page faults. We evaluate this system in the gem5 architectural simulator with a modified Linux kernel running on top of simulated hardware containing the MFOE accelerator. Our results show that MFOE improves the average critical path fault handling latency by 33× and tail critical path latency by 51×. Among the evaluated applications, we observed an improvement of runtime by an average of 6.6%.
Chandrahas Tirumalasetty, Chih-Chieh Chou, A. L. Narasimha Reddy, Paul Gratz, Ayman Abouelwafa
ACM Trans. Archit. Code Optim.3
2022 Software Hint-Driven Data Management for Hybrid Memory in Mobile Systems
abstract
Hybrid memory systems, comprised of emerging non-volatile memory (NVM) and DRAM, have been proposed to address the growing memory demand of current mobile applications. Recently emerging NVM technologies, such as phase-change memories (PCM), memristor, and 3D XPoint, have higher capacity density, minimal static power consumption and lower cost per GB. However, NVM has longer access latency and limited write endurance as opposed to DRAM. The different characteristics of distinct memory classes render a new challenge for memory system design. Ideally, pages should be placed or migrated between the two types of memories according to the data objects’ access properties. Prior system software approaches exploit the program information from OS but at the cost of high software latency incurred by related kernel processes. Hardware approaches can avoid these latencies, however, hardware’s vision is constrained to a short time window of recent memory requests, due to the limited on-chip resources. In this work, we propose OpenMem: a hardware-software cooperative approach that combines the execution time advantages of pure hardware approaches with the data object properties in a global scope. First, we built a hardware-based memory manager unit (HMMU) that can learn the short-term access patterns by online profiling, and execute data migration efficiently. Then, we built a heap memory manager for the heterogeneous memory systems that allows the programmer to directly customize each data object’s allocation to a favorable memory device within the presumed object life cycle. With the programmer’s hints guiding the data placement at allocation time, data objects with similar properties will be congregated to reduce unnecessary page migrations. We implemented the whole system on the FPGA board with embedded ARM processors. In testing under a set of benchmark applications from SPEC 2017 and PARSEC, experimental results show that OpenMem reduces 44.6% energy consumption with only a 16% performance degradation compared to the all-DRAM memory system. The amount of writes to the NVM is reduced by 14% versus the HMMU-only, extending the NVM device lifetime.
Fei Wen 0003, Paul Gratz, A. L. Narasimha Reddy
ACM Trans. Embed. Comput. Syst.4
2021 OpenMem: Hardware/Software Cooperative Management for Mobile Memory System
abstract
Hybrid memory systems, comprised of emerging non-volatile memory (NVM) and DRAM, have been proposed to address the growing memory demand of current mobile applications. NVM technologies have higher capacity density, minimal static power consumption, but longer access latency and limited write endurance compared to DRAM. The different characteristics of these two memory classes, however, pose new challenges for memory system design. Ideally, pages shall be placed or migrated between the two types of memories according to the data objects’ access properties. Prior works use the OS for placement and migration in these systems, but at the cost of high software latency incurred by related kernel processes. Hardware approaches can avoid these latencies, however, hardware’s vision is constrained to a short time window of recently memory request, due to the limited on-chip resources.In this work, we propose OpenMem: a hardware-software cooperative approach to address placement and migration within hybrid memory systems, that combines the execution time advantages of pure hardware approaches with the data object properties in a global scope. We emulate OpenMem on an FPGA board with embedded ARM CPU, and run a set of benchmark applications from SPEC 2017 and PARSEC. Experimental results show that OpenMem reduces energy consumption by 44.6% with only a 16% performance degradation compared to an all-DRAM memory system. Further, writes to the NVM are reduced by 14% versus a hardware-only approach, extending the NVM device lifetime.
Fei Wen 0003, Paul Gratz, A. L. Narasimha Reddy
DAC4
2021 An FPGA-based Hybrid Memory Emulation System
abstract
Hybrid memory systems, comprised of emerging non-volatile memory (NVM) and DRAM, have been proposed to address the growing memory demand of applications. Emerging NVM technologies, such as phase-change memories (PCM), memristor, and 3D XPoint, have higher capacity density, minimal static power consumption and lower cost per GB. However, NVM has longer access latency and limited write endurance as opposed to DRAM. The different characteristics of two memory classes point towards the design of hybrid memory systems containing multiple classes of main memory.In the iterative and incremental development of new architectures, the timeliness of simulation completion is critical to project progression. Hence, a highly efficient simulation method is needed to evaluate the performance of different hybrid memory system designs. Design exploration for hybrid memory systems is challenging, because it requires emulation of the full system stack, including the OS, memory controller, and interconnect. Moreover, benchmark applications for memory performance tests typically have much larger working sets, thus taking an even longer simulation warm-up period.In this paper, we propose an FPGA-based hybrid memory system emulation platform. We target the mobile computing system, which is sensitive to energy consumption and is likely to adopt NVM for its power efficiency. The focus of our platform is on the design of hybrid memory system, so we leverage the on-board hard IP ARM processors to enhance simulation performance while improving the accuracy of results. Thus, users can implement their data placement/migration policies with the FPGA logic elements and evaluate new designs quickly and effectively. Results show that our emulation platform provides a speedup of 9280x in simulation time compared to the software counterpart gem5.
Fei Wen 0003, Paul Gratz, A. L. Narasimha Reddy
FPL4
2021 KVRAID: high performance, write efficient, update friendly erasure coding scheme for KV-SSDs
abstract
Key-value (KV) stores have been widely deployed in a variety of scale-out enterprise applications such as online retail, big data analytics, social networks, etc. Key-Value SSDs (KVSSDs) provide a key-value interface directly from the device aiming at lowering software overhead and reducing I/O amplification for such applications.
A. L. Narasimha Reddy, Paul Gratz, Rekha Pitchumani, Yang-Seok Ki
SYSTOR2
2021 A Survey of Cybersecurity of Digital Manufacturing
abstract
The Industry 4.0 concept promotes a digital manufacturing (DM) paradigm that can enhance quality and productivity, which reduces inventory and the lead time for delivering custom, batch-of-one products based on achieving convergence of additive, subtractive, and hybrid manufacturing machines, automation and robotic systems, sensors, computing, and communication networks, artificial intelligence, and big data. A DM system consists of embedded electronics, sensors, actuators, control software, and interconnectivity to enable the machines and the components within them to exchange data with other machines, components therein, the plant operators, the inventory managers, and customers. This article presents the cybersecurity risks in the emerging DM context, assesses the impact on manufacturing, and identifies approaches to secure DM.
Priyanka Mahesh, Akash Tiwari, Chenglu Jin, P. R. Kumar 0001, A. L. Narasimha Reddy, Satish T. S. Bukkapatnam, Nikhil Gupta 0002, Ramesh Karri
Proc. IEEE5
2020 A Generic FPGA Accelerator for Minimum Storage Regenerating Codes
abstract
Erasure coding is widely used in storage systems to achieve fault tolerance while minimizing the storage overhead. Recently, Minimum Storage Regenerating (MSR) codes are emerging to minimize repair bandwidth while maintaining the storage efficiency. Traditionally, erasure coding is implemented in the storage software stacks, which hinders normal operations and blocks resources that could be serving other user needs due to poor cache performance and costs high CPU and memory utilizations. In this paper, we propose a generic FPGA accelerator for MSR codes encoding/decoding which maximizes the computation parallelism and minimizes the data movement between off-chip DRAM and the on-chip SRAM buffers. To demonstrate the efficiency of our proposed accelerator, we implemented the encoding/decoding algorithms for a specific MSR code called Zigzag code on Xilinx VCU1525 acceleration card. Our evaluation shows our proposed accelerator can achieve ~2.4-3.1x better throughput and ~4.2-5.7x better power efficiency compared to the state-of-art multi-core CPU implementation and ~2.8-3.3x better throughput and ~4.2-5.3x better power efficiency compared to a modern GPU accelerator.
Joo Hwan Lee, Rekha Pitchumani, Yang-Seok Ki, A. L. Narasimha Reddy, Paul Gratz
ASP-DAC5
2020 Virtualize and share non-volatile memories in user space
Chih-Chieh Chou, Jaemin Jung, A. L. Narasimha Reddy, Paul Gratz, Doug Voigt
CCF Trans. High Perform. Comput.3
2020 Hardware Memory Management for Future Mobile Hybrid Memory Systems
abstract
The current mobile applications have rapidly growing memory footprints, posing a great challenge for memory system design. Insufficient DRAM main memory will incur frequent data swaps between memory and storage, a process that hurts performance, consumes energy, and deteriorates the write endurance of typical flash storage devices. Alternately, a larger DRAM has higher leakage power and drains the battery faster. Furthermore, DRAM scaling trends make further growth of DRAM in the mobile space prohibitive due to cost. Emerging nonvolatile memory (NVM) has the potential to alleviate these issues due to its higher capacity per cost than DRAM and minimal static power. Recently, a wide spectrum of NVM technologies, including phase-change memories (PCMs), memristor, and 3-D XPoint has emerged. Despite the mentioned advantages, NVM has longer access latency compared to DRAM and NVM writes can incur higher latencies and wear costs. Therefore, the integration of these new memory technologies in the memory hierarchy requires a fundamental rearchitecting of traditional system designs. In this work, we propose a hardware-accelerated memory manager (HMMU) that addresses in a flat address space, with a small partition of the DRAM reserved for subpage block-level management. We design a set of data placement and data migration policies within this memory manager such that we may exploit the advantages of each memory technology. By augmenting the system with this HMMU, we reduce the overall memory latency while also reducing writes to the NVM. The experimental results show that our design achieves a 39% reduction in energy consumption with only a 12% performance degradation versus an all-DRAM baseline that is likely untenable in the future.
Fei Wen 0003, Paul Gratz, A. L. Narasimha Reddy
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2019 vNVML: An Efficient User Space Library for Virtualizing and Sharing Non-Volatile Memories
abstract
The emerging non-volatile memory (NVM) has attractive characteristics such as DRAM-like, low-latency together with the non-volatility of storage devices. Recently, byte-addressable, memory bus-attached NVM has become available. This paper addresses the problem of combining a smaller, faster byte-addressable NVM with a larger, slower storage device, like SSD, to create the impression of a larger and faster byte-addressable NVM which can be shared across many applications. In this paper, we propose vNVML, a user space library for virtualizing and sharing NVM. vNVML provides for applications transaction like memory semantics that ensures write ordering and persistency guarantees across system failures. vNVML exploits DRAM for read caching, to enable improvements in performance and potentially to reduce the number of writes to NVM, extending the NVM lifetime. vNVML is implemented and evaluated with realistic workloads to show that our library allows applications to share NVM, both in a single O/S and when docker like containers are employed. The results from the evaluation show that vNVML incurs less than 10% overhead while providing the benefits of an expanded virtualized NVM space to the applications, allowing applications to safely share the virtual NVM.
Chih-Chieh Chou, Jaemin Jung, A. L. Narasimha Reddy, Paul Gratz, Doug Voigt
MSST3
2018 IoTAegis: A Scalable Framework to Secure the Internet of Things
abstract
The infamous Mirai attack which hijacked nearly half a million Internet connected devices demonstrated the widespread security vulnerabilities of the Internet-of-Things (IoT). This study employs a set of active and passive observation methods to discover the security vulnerabilities of IoT devices within a university campus. We show that (a) the number of non-compute devices dominates the number of compute devices with open ports in a campus network; (b) 58.9% or more devices do not keep up-to-date firmware and 51.3% or more do not have a user defined password; and (c) the number of devices together with the diversity of device ages and vendors make the protection of IoT devices a difficult problem. We further develop IoTAegis framework which offers device-level protection to automatically manage device configurations and security updates. Our solution is shown to be effective, scalable, lightweight, and deployable in different forms and network types.
Allen Webb, A. L. Narasimha Reddy, Riccardo Bettati
ICCCN3
2017 Safeguarding Building Automation Networks: THE-Driven Anomaly Detector Based on Traffic Analysis
abstract
Building Automation Systems (BAS) are distributed networks of hardware and software that monitor and control heating, ventilation, and air-conditioning (HVAC), as well as lighting and security of smart buildings. BACnet is a standard data communication protocol designed to operate across many types of BAS field panels and controllers. This paper studies BACnet traffic in a real-world BAS from various vantage points and develops an anomaly detector for BAS networks. Our analysis of BACnet traffic through several measures reveals that BACnet traffic is neither strictly periodic as expected of control traffic nor exhibits diurnal patterns of IP network traffic. BACnet traffic is a combination of multiple flow-service streams that belong to "THE-driven'" categories: Time-driven, Human-driven, and Event-driven. Time-driven traffic follows periodic patterns, regular patterns, or on/off models. Human-driven and event- driven traffic present non-periodic patterns. We construct flow-service models for time-driven traffic and develop THE-Driven Anomaly Detector which adopts different mechanisms for each category of traffic. We evaluate the anomaly detector using k-fold cross validation and synthetic attacks. The proposed THE-Driven Anomaly Detector is shown to be able to effectively detect suspicious traffic in BAS networks with small false alarm rate.
A. L. Narasimha Reddy
ICCCN2
2016 Path confidence based lookahead prefetching
abstract
Designing prefetchers to maximize system performance often requires a delicate balance between coverage and accuracy. Achieving both high coverage and accuracy is particularly challenging in workloads with complex address patterns, which may require large amounts of history to accurately predict future addresses. This paper describes the Signature Path Prefetcher (SPP), which offers effective solutions for three classic challenges in prefetcher design. First, SPP uses a compressed history based scheme that accurately predicts complex address patterns. Second, unlike other history based algorithms, which miss out on many prefetching opportunities when address patterns make a transition between physical pages, SPP tracks complex patterns across physical page boundaries and continues prefetching as soon as they move to new pages. Finally, SPP uses the confidence it has in its predictions to adaptively throttle itself on a per-prefetch stream basis. In our analysis, we find that SPP improves performance by 27.2% over a no-prefetching baseline, and outperforms the state-of-the-art Best Offset prefetcher by 6.4%. SPP does this with minimal overhead, operating strictly in the physical address space, and without requiring any additional processor core state, such as the PC.
Jinchun Kim, Seth H. Pugsley, Paul Gratz, A. L. Narasimha Reddy, Chris Wilkerson, Zeshan Chishti
MICRO4
2016 Adaptive policies for balancing performance and lifetime of mixed SSD arrays through workload sampling
abstract
Solid-state drives (SSDs) have become promising storage components to serve large I/O demands in modern storage systems. Enterprise class (high-end) SSDs are faster and more resilient than client class (low-end) SSDs but they are expensive to be deployed in large scale storage systems. It is an attractive and practical alternative to exploit the high-end SSDs as a cache and low-end SSDs as main storage. This paper explores how to optimize a mixed SSD array in terms of performance and lifetime. This paper shows that simple integration of different classes of SSDs in traditional caching policies results in poor reliability. This paper also reveals that caching policies with static workload distribution are not always efficient. In this paper, we propose a sampling based adaptive approach that achieves fair workload distribution across the cache and the storage. The proposed algorithm enables fine-grained control of the workload distribution which minimizes latency over lifetime of mixed SSD arrays. We show that our adaptive algorithm is very effective in improving the latency over lifetime metric, on an average, by up to 2.36 times over LRU, across a number of workloads.
Sangwhan Moon, A. L. Narasimha Reddy
MSST2
2016 Does RAID Improve Lifetime of SSD Arrays?
abstract
Parity protection at the system level is typically employed to compose reliable storage systems. However, careful consideration is required when SSD-based systems employ parity protection. First, additional writes are required for parity updates. Second, parity consumes space on the device, which results in write amplification from less efficient garbage collection at higher space utilization. This article analyzes the effectiveness of SSD-based RAID and discusses the potential benefits and drawbacks in terms of reliability. A Markov model is presented to estimate the lifetime of SSD-based RAID systems in different environments. In a small array, our results show that parity protection provides benefit only with considerably low space utilizations and low data access rates. However, in a large system, RAID improves data lifetime even when we take write amplification into account.
Sangwhan Moon, A. L. Narasimha Reddy
ACM Trans. Storage2
2013 Don't Let RAID Raid the Lifetime of Your SSD Array
Sangwhan Moon, A. L. Narasimha Reddy
HotStorage2
2013 Multi Path PERT
abstract
This paper presents a new multipath delay based algorithm, MPPERT (Multipath Probabilistic Early response TCP), which provides high throughput and efficient load balancing. In all-PERT environment, MPPERT suffers no packet loss and maintains much smaller queue sizes compared to existing MPTCP, making it suitable for real time data transfer. MPPERT is suitable for incremental deployment in a heterogeneous environment. PERT, being a delay based TCP protocol, has continuous information about the state of the bottleneck queue along its path. This information is valuable in enabling MPPERT to detect subflows sharing a common bottleneck and obtain a smaller set of disjoint subflows. This information can even be used to switch from coupled (a set of subflows having interdependent increase/decrease of congestion windows) to uncoupled (independent increase/decrease of congestion windows) subflows, yielding higher throughput when best single-path TCP constraint is relaxed. The ns-2 simulations support MPPERT as a highly competitive multipath approach, suitable for real time data transfer, which is capable of offering higher throughput and improved reliability.
A. L. Narasimha Reddy
ICCCN2
2013 NVMFS: A hybrid file system for improving random write in nand-flash SSD
abstract
In this paper, we design a storage system consisting of Nonvolatile DIMMs (as NVRAM) and NAND-flash SSD. We propose a file system NVMFS to exploit the unique characteristics of these devices which simplifies and speeds up file system operations. We use the higher performance NVRAM as both a cache and permanent space for data. Hot data can be permanently stored on NVRAM without writing back to SSD, while relatively cold data can be temporarily cached by NVRAM with another copy on SSD. We also reduce the erase overhead of SSD by reorganizing writes on NVRAM before flushing to SSD. We have implemented a prototype NVMFS within a Linux Kernel and compared with several modern file systems such as ext3, btrfs and NILFS2. We also compared with another hybrid file system Conquest, which originally was designed for NVRAM and HDD. The experimental results show that NVMFS improves IO throughput by an average of 98.9 % when segment cleaning is not active, while improves throughput by an average of 19.6% under high disk utilization (over 85%) compared to other file systems. We also show that our file system can reduce the erase operations and overheads at SSD.
Sheng Qiu, A. L. Narasimha Reddy
MSST2
2013 ARI: Adaptive LLC-memory traffic management
abstract
Decreasing the traffic from the CPU LLC to main memory is a very important issue in modern systems. Recent work focuses on cache misses, overlooking the impact of writebacks on the total memory traffic, energy consumption, IPC, and so forth. Policies that foster a balanced approach, between reducing write traffic to memory and improving miss rates, can increase overall performance and improve energy efficiency and memory system lifetime for NVM memory technology, such as phase-change memory (PCM). We propose Adaptive Replacement and Insertion (ARI), an adaptive approach to last-level CPU cache management, optimizing the two parameters (miss rate and writeback rate) simultaneously. Our specific focus is to reduce writebacks as much as possible while maintaining or improving the miss rate relative to conventional LRU replacement policy. ARI reduces LLC writebacks by 33%, on average, while also decreasing misses by 4.7%, on average. In a typical system, this boosts IPC by 4.9%, on average, while decreasing energy consumption by 8.9%. These results are achieved with minimal hardware overheads.
Viacheslav V. Fedorov, Sheng Qiu, A. L. Narasimha Reddy, Paul Gratz
ACM Trans. Archit. Code Optim.3
2013 SCMFS: A File System for Storage Class Memory and its Extensions
abstract
Modern computer systems have been built around the assumption that persistent storage is accessed via a slow, block-based interface. However, emerging nonvolatile memory technologies (sometimes referred to as storage class memory (SCM)), are poised to revolutionize storage systems. The SCM devices can be attached directly to the memory bus and offer fast, fine-grained access to persistent storage. In this article, we propose a new file system---SCMFS, which is specially designed for Storage Class Memory. SCMFS is implemented on the virtual address space and utilizes the existing memory management module of the operating system to help mange the file system space. As a result, we largely simplified the file system operations of SCMFS, which allowed us a better exploration of performance gain from SCM. We have implemented a prototype in Linux and evaluated its performance through multiple benchmarks. The experimental results show that SCMFS outperforms other memory resident file systems, tmpfs, ramfs and ext2 on ramdisk, and achieves about 70% of memory bandwidth for file read/write operations.
Xiaojian Wu, Sheng Qiu, A. L. Narasimha Reddy
ACM Trans. Storage3
2012 Write amplification due to ECC on flash memory or leave those bit errors alone
abstract
While flash memory is receiving significant attention because of many attractive properties, concerns about write endurance delay the wider deployment of the flash memory. This paper analyzes the effectiveness of protection schemes designed for flash memory, such as ECC and scrubbing. The bit error rate of flash memory is a function of the number of program-erase cycles the cell has gone through, making the reliability dependent on time and workload. Moreover, some of the protection schemes require additional write operations, which degrade flash memory's reliability. These issues make it more complex to reveal the relationship between the protection schemes and flash memory's lifetime. In this paper, a Markov model based analysis of the protection schemes is presented. Our model considers the time varying reliability of flash memory as well as write amplification of various protection schemes such as ECC. Our study shows that write amplification from these various sources can significantly affect the benefits of these schemes in improving the lifetime. Based on the results from our analysis, we propose that bit errors within a page be left uncorrected until a threshold of errors are accumulated. We show that such an approach can significantly improve lifetimes by up to 40%.
Sangwhan Moon, A. L. Narasimha Reddy
MSST2
2012 Exploiting superpages in a nonvolatile memory file system
abstract
Emerging nonvolatile memory technologies (sometimes referred as Storage Class Memory (SCM)), are poised to close the enormous performance gap between persistent storage and main memory. The SCM devices can be attached directly to memory bus and accessed like normal DRAM. It becomes then possible to exploit memory management hardware resources to improve file system performance. However, in this case, SCM may share critical system resources such as the TLB, page table with DRAM which can potentially impact SCM's performance. In this paper, we propose to solve this problem by employing superpages to reduce the pressure on memory management resources such as the TLB. As a result, the file system performance is further improved. We also analyze the space utilization efficiency of superpages. We improve space efficiency of the file system by allocating normal pages (4KB) for small files while allocating super pages (2MB on ×86) for large files. We show that it is possible to achieve better performance without loss of space utilization efficiency of nonvolatile memory.
Sheng Qiu, A. L. Narasimha Reddy
MSST2
2012 Constructing disjoint paths for failure recovery and multipath routing
Yong Oh Lee, A. L. Narasimha Reddy
Comput. Networks2
2012 A Large-Scale Empirical Study of Conficker
abstract
Conficker is the most recent widespread, well-known worm/bot. According to several reports, it has infected about 7 million to 15 million hosts and the victims are still increasing even now. In this paper, we analyze Conficker infections at a large scale, about 25 million victims, and study various interesting aspects about this state-of-the-art malware. By analyzing Conficker, we intend to understand current and new trends in malware propagation, which could be very helpful in predicting future malware trends and providing insights for future malware defense. We observe that Conficker has some very different victim distribution patterns compared to many previous generation worms/botnets, suggesting that new malware spreading models and defense strategies are likely needed. We measure the potential power of Conficker to estimate its effects on the networks/hosts when it performs malicious operations. Furthermore, we intend to determine how well a reputation-based blacklisting approach can perform when faced with new malware threats such as Conficker. We cross-check several DNS blacklists and IP/AS reputation data from Dshield and FIRE and our evaluation shows that unlike a previous study which shows that a blacklist-based approach can detect most bots, these reputation-based approaches did relatively poorly for Conficker. This raises a question of how we can improve and complement existing reputation-based techniques to prepare for future malware defense? Based on this, we look into some insights for defenders. We show that neighborhood watch is a surprisingly effective approach in the case of Conficker. This suggests that security alert sharing/correlation (particularly among neighborhood networks) could be a promising approach and play a more important role for future malware defense.
Seungwon Shin 0001, Guofei Gu, A. L. Narasimha Reddy, Christopher P. Lee 0001
IEEE Trans. Inf. Forensics Secur.3
2012 Detecting Algorithmically Generated Domain-Flux Attacks With DNS Traffic Analysis
abstract
Recent botnets such as Conficker, Kraken, and Torpig have used DNS-based “domain fluxing” for command-and-control, where each Bot queries for existence of a series of domain names and the owner has to register only one such domain name. In this paper, we develop a methodology to detect such “domain fluxes” in DNS traffic by looking for patterns inherent to domain names that are generated algorithmically, in contrast to those generated by humans. In particular, we look at distribution of alphanumeric characters as well as bigrams in all domains that are mapped to the same set of IP addresses. We present and compare the performance of several distance metrics, including K-L distance, Edit distance, and Jaccard measure. We train by using a good dataset of domains obtained via a crawl of domains mapped to all IPv4 address space and modeling bad datasets based on behaviors seen so far and expected. We also apply our methodology to packet traces collected at a Tier-1 ISP and show we can automatically detect domain fluxing as used by Conficker botnet with minimal false positives, in addition to discovering a new botnet within the ISP trace. We also analyze a campus DNS trace to detect another unknown botnet exhibiting advanced domain-name generation technique.
Sandeep Yadav, Ashwath Kumar Krishna Reddy, A. L. Narasimha Reddy, Supranamaya Ranjan
IEEE/ACM Trans. Netw.3
2011 SCMFS: a file system for storage class memory
abstract
This paper considers the problem of how to implement a file system on Storage Class Memory (SCM), that is directly connected to the memory bus, byte addressable and is also non-volatile. In this paper, we propose a new file system, called SCMFS, which is implemented on the virtual address space. In SCMFS, we utilize the existing memory management module in the operating system to do the block management and keep the space always contiguous for each file. The simplicity of SCMFS not only makes it easy to implement, but also improves the performance. We have implemented a prototype in Linux and evaluated its performance through multiple benchmarks.
Xiaojian Wu, A. L. Narasimha Reddy
SC2
2011 Winning with DNS Failures: Strategies for Faster Botnet Detection
Sandeep Yadav, A. L. Narasimha Reddy
SecureComm2
2010 Disjoint Multi-Path Routing and Failure Recovery
abstract
Applications such as Voice over IP and video delivery require continuous network service, requiring fast failure recovery mechanisms. Proactive Fast failure recovery mechanisms have been recently proposed to improve network performance during the failure transients. The proposed mechanisms need extra infrastructural support in the form of routing table entries, extra addresses etc. In this paper, we study if the extra infrastructure support can be exploited to build disjoint paths in those frameworks, while keeping the recovery path lengths close to the primary paths. Our evaluations show that it is possible to extend the proactive recovery mechanisms to provide support for nearly-disjoint paths.
Yong Oh Lee, A. L. Narasimha Reddy
ICC2
2010 Detecting algorithmically generated malicious domain names
abstract
Recent Botnets such as Conficker, Kraken and Torpig have used DNS based "domain fluxing" for command-and-control, where each Bot queries for existence of a series of domain names and the owner has to register only one such domain name. In this paper, we develop a methodology to detect such "domain fluxes" in DNS traffic by looking for patterns inherent to domain names that are generated algorithmically, in contrast to those generated by humans. In particular, we look at distribution of alphanumeric characters as well as bigrams in all domains that are mapped to the same set of IP-addresses. We present and compare the performance of several distance metrics, including KL-distance, Edit distance and Jaccard measure. We train by using a good data set of domains obtained via a crawl of domains mapped to all IPv4 address space and modeling bad data sets based on behaviors seen so far and expected. We also apply our methodology to packet traces collected at a Tier-1 ISP and show we can automatically detect domain fluxing as used by Conficker botnet with minimal false positives.
Sandeep Yadav, Ashwath Kumar Krishna Reddy, A. L. Narasimha Reddy, Supranamaya Ranjan
Internet Measurement Conference3
2010 Performance of Quantized Congestion Notification in TCP Incast Scenarios of Data Centers
abstract
This paper analyzes the performance of Ethernet layer congestion control mechanism Quantized Congestion Notification (QCN) during data access from clustered servers in data centers. We analyze the reasons why QCN does not perform adequately in these situations and propose several modifications to the protocol to improve its performance in these scenarios. We trace the causes of QCN performance degradation to flow rate variability, and show that adaptive sampling at the switch and adaptive self-increase of flow rates at the rate limiter improve performance in a TCP In cast setup significantly. We compare the performance of QCN against TCP modifications in a heterogeneous environment, and show that modifications to QCN yield better performance.
Prajjwal Devkota, A. L. Narasimha Reddy
MASCOTS2
2010 Exploiting Concurrency to Improve Latency and throughput in a Hybrid Storage System
abstract
This paper considers the problem of how to improve the performance of hybrid storage system employing solid state disks and hard disk drives. We utilize both initial block allocation as well as migration to reach “Wardrop equilibrium”, in which the response times of different devices equalize. We show that such a policy allows adaptive load balancing across devices of different performance. We also show that such a policy exploits parallelism in the storage system effectively to improve throughput and latency simultaneously. We implemented a prototype in Linux and evaluated it in multiple workloads and multiple configurations. The results show that the proposed approach improved both the latency of requests and the throughput significantly, and it adapted to different configurations of the system under different workloads.
Xiaojian Wu, A. L. Narasimha Reddy
MASCOTS2
2009 Managing storage space in a flash and disk hybrid storage system
abstract
This paper considers the problem of efficiently managing storage space in a hybrid storage system employing flash and disk drives. The flash and disk drives exhibit different performance characteristics of read and write behavior. We propose a technique for balancing the workload properties across flash and disk drives in such a hybrid storage system. The presented approach automatically and transparently manages migration of data blocks among flash and disk drives based on their access patterns. This paper presents the design and an evaluation of the proposed approach on a Linux testbed through realistic experiments.
Xiaojian Wu, A. L. Narasimha Reddy
MASCOTS2
2009 Umbrella file system: Storage management across heterogeneous devices
abstract
With the advent of and recent developments in Flash storage, device characteristic diversity is becoming both more prevalent and more distinct. In this article, we describe the Umbrella File System (UmbrellaFS), a stackable file system designed to provide flexibility in matching diversity of file access characteristics to diversity of device characteristics through a user or system administrator specified policy. We present the design and results from a prototype implementation of UmbrellaFS on both Linux 2.4 and 2.6. The results show that UmbrellaFS has little overhead for most file system operations while providing an ability better to utilize the differences in Flash and traditional hard drives. With appropriate use of rules, we have shown improvements of up to 44% in certain situations.
John A. Garrison, A. L. Narasimha Reddy
ACM Trans. Storage2
2008 User-centric data migration in networked storage systems
abstract
This paper considers the problem of balancing locality and load in networked storage systems with multiple storage devices (or bricks). Data distribution affects locality and load balance across the devices in a networked storage system. This paper proposes a user-centric data migration scheme which tries to balance locality and load in such networked storage systems. The presented approach automatically and transparently manages migration of data blocks among disks as data access patterns and loads change over time. We implemented a prototype system, embodying our ideas, on PCs running Linux. This paper presents the design of user-centric migration and an evaluation of it through realistic experiments.
Sukwoo Kang, A. L. Narasimha Reddy
IPDPS2
2008 Making a Delay-Based Protocol Adaptive to Heterogeneous Environments
abstract
This paper investigates the issues in making a delay-based protocol adaptive to heterogeneous environments. We address how a delay-based protocol can compete with a loss- based protocol such as TCP. We investigate if potential noise and variability in delay measurements in environments such as cable and ADSL access networks impact the protocol behavior significantly. We investigate these issues in the context of incremental deployment of a new delay-based protocol, PERT. We propose design modifications to PERT to compete with SACK. We show that PERT experiences lower drop rates than SACK and leads to lower overall drop rates with different mixes of PERT and SACK protocols. Second, we show that a single PERT flow can fully utilize a high-speed, high-delay link. The results from ns-2 simulations indicate that PERT can adapt to heterogeneous networks and can operate well in an environment of heterogeneous protocols. We also show that proposed changes retain the desirable properties of PERT such as low loss rates and fairness, when operating alone. The protocol has also been implemented in the Linux kernel and tested through experiments on live networks, by measuring the throughput and losses between nodes in our lab at TAMU and different machines on the planet-lab.
Kiran Kotla, A. L. Narasimha Reddy
IWQoS2
2008 Statistical techniques for detecting traffic anomalies through packet header data
Seong Soo Kim, A. L. Narasimha Reddy
IEEE/ACM Trans. Netw.2
2007 Emulating AQM from end hosts
abstract
In this paper, we show that end-host based congestion prediction is more accurate than previously characterized. However, it may not be possible to entirely eliminate the uncertainties in congestion prediction. To address these uncertainties, we propose Probabilistic Early Response TCP (PERT). PERT emulates the behavior of AQM/ECN, in the congestion response function of end-hosts. We present fluid-flow analysis of PERT/RED and PERT/PI, versions of PERT that emulate router-based RED and PI controllers. Our analysis shows that PERT/RED has better stability behavior than router-based RED. We also present results from ns-2 simulations to show the practical feasibility of PERT. The scheme presented here is general and can be used for emulating other AQM algorithms.
Sumitha Bhandarkar, A. L. Narasimha Reddy, Yueping Zhang, Dmitri Loguinov
SIGCOMM2
2007 Multihoming route control among a group of multihomed stub networks
Yong Liu 0014, A. L. Narasimha Reddy
Comput. Commun.2
2006 A Service Providers Approach for Improving Performance of Aggregate Voice-over-IP Traffic
abstract
The emerging popularity and interest in Voice-over-IP (VoIP) has been accompanied by customer concerns about voice quality over these networks. The lack of an appropriate real-time capable infrastructure in packet networks along with the threats of denial-of service (DoS) attacks can deteriorate the service that these voice calls receive. Traditionally, each voice call employs its own end-to-end forward-error-correction (FEC) mechanisms. In this paper, we show that when VoIP calls are aggregated over a network link or path, the provider can employ a suitable linear-time encoding for the aggregated voice traffic, resulting in considerable quality improvement with little redundancy. We show that it is possible to achieve rates closer to link capacity as more calls are combined with very small output loss rates even in the presence of significant packet loss rates in the network. The advantages of the proposed scheme exceed similar or other techniques applied to individual voice calls.
Camelia Al-Najjar, A. L. Narasimha Reddy
IWQoS2
2006 Image-Based Anomaly Detection Technique: Algorithm, Implementation and Effectiveness
abstract
The frequent and large-scale network attacks have led to an increased need for developing techniques for analyzing network traffic. This paper presents NetViewer, a network measurement approach that can simultaneously detect, identify, and visualize attacks and anomalous traffic in real-time by passively monitoring packet headers. We propose to represent samples of network packet header data as frames or images. With such a formulation, a series of samples can be seen as a sequence of frames or video, revealing certain kinds of attacks to the human eye. This enables techniques from image processing and video compression to be applied to the packet header data to reveal interesting properties of traffic. We show that "scene change analysis" can reveal sudden changes in traffic behavior or anomalies. We also show that "motion prediction" techniques can be employed to understand the patterns of some of the attacks. We show that it may be feasible to represent multiple pieces of data as different colors of an image enabling a uniform treatment of multidimensional packet header data. We compare the effectiveness of NetViewer with classical detection theory-based Neyman-Pearson test
Seong Soo Kim, A. L. Narasimha Reddy
IEEE J. Sel. Areas Commun.2
2006 An approach to virtual allocation in storage systems
abstract
This article presents virtual allocation , a scheme for flexible storage allocation. Virtual allocation separates storage allocation from the file system. It employs an allocate-on-write strategy which lets applications fit into the actual usage of storage space, without regard to the configured file system size. This improves flexibility by allowing storage space to be shared across different file systems. This article presents the design of virtual allocation and its evaluation through benchmarks. To illustrate our approach, we implemented a prototype system on PCs running Linux. We present the results from the prototype implementation and its evaluation.
Sukwoo Kang, A. L. Narasimha Reddy
ACM Trans. Storage2
2005 Route optimization among a group of multihomed stub networks
abstract
Multihoming is used by stub networks to improve the reliability of their Internet connectivity. In recent years, commercial "intelligent route control" devices are used by multihomed stub networks to optimize the routing of their Internet traffic. In this work, we first conduct measurement of qualities of alternate paths through multihoming. In the remaining part, we study the route control of traffic among a group of multihomed stub networks which may belong to an organization and exchange data regularly. This type of route control is special because the access links of such networks may not be over-provisioned and are shared by traffic controlled by multiple route control devices. Such shared bottlenecks may affect the effectiveness of uncoordinated "intelligent route control". We propose a global optimization based approach to coordinate the route control among such a group of networks. Our approach can avoid oscillations which may be caused by uncoordinated route control. Simulation results show that our approach has advantages over equal-splitting based static load-balancing.
Yong Liu 0014, A. L. Narasimha Reddy
GLOBECOM2
2005 Modeling network traffic as images
abstract
The paper presents a network measurement approach to represent samples of network packet header data as frames or images. With such a formulation, a series of samples can be seen as a sequence of frames or video. This enables techniques from image processing and video compression to be applied to the analysis of packet header data to reveal interesting traffic properties. We show that traffic images can reveal sudden changes in traffic behavior or anomalies. Using a combination of visual modeling and trace-driven simulation, we evaluate how the design factors impact the representation of dynamic network traffic. In particular, we study the impact of sampling rate and retained DCT coefficients on the network traffic data representation.
Seong Soo Kim, A. L. Narasimha Reddy
ICC2
2005 Real-time detection and containment of network attacks using QoS regulation
abstract
In this paper, we present a network measurement mechanism that can detect and mitigate attacks and anomalous traffic in real-time using QoS regulation. The detection method rapidly pursues the dynamics of the network on the basis of correlation properties of the network protocols. By observing the proportion occupied by each traffic protocol and correlating it to that of previous states of traffic, it can be possible to determine whether the current traffic is behaving normally. When abnormalities are detected, our mechanism allows aggregated resource regulation of each protocol's traffic. The trace-driven results show that the rate-based regulation of traffic characterized by protocol classes is a feasible vehicle for mitigating the impact of network attacks on end servers.
Seong Soo Kim, A. L. Narasimha Reddy
ICC2
2005 A study of analyzing network traffic as images in real-time
abstract
This paper presents NetViewer, a network measurement approach that can simultaneously detect, identify and visualize attacks and anomalous traffic in real-time by passively monitoring packet headers. We propose to represent samples of network packet header data as frames or images. With such a formulation, a series of samples can be seen as a sequence of frames or video. This enables techniques from image processing and video compression to be applied to the packet header data to reveal interesting properties of traffic. We show that "scene change analysis" can reveal sudden changes in traffic behavior or anomalies. We also show that "motion prediction" techniques can be employed to understand the patterns of some of the attacks. We show that it may be feasible to represent multiple pieces of data as different colors of an image enabling a uniform treatment of multidimensional packet header data. We compare NetViewer with classical detection theory based Neyman-Pearson test and an IDS tool.
Seong Soo Kim, A. L. Narasimha Reddy
INFOCOM2
2005 NetViewer: A Network Traffic Visualization and Analysis Tool
Seong Soo Kim, A. L. Narasimha Reddy
LISA2
2005 TCP-DCR: A Novel Protocol for Tolerating Wireless Channel Errors
abstract
This paper presents TCP-DCR, a set of simple modifications to the TCP protocol to improve its robustness to channel errors in wireless networks. TCP-DCR is based on the simple idea of allowing the link-level mechanism to recover the packets lost, due to channel errors, thereby limiting the response of the transport protocol to mostly congestion losses. This is done by delaying the triggering of congestion response algorithms for a small bounded period of time /spl tau/ to allow the link-level retransmissions to recover the loss due to channel errors. If at the end of the delay /spl tau/ the packet is not recovered, then it is treated as a packet lost due to congestion. We analyze TCP-DCR to show that the delay in congestion response does not impact the fairness towards the native implementations of TCP that respond to congestion immediately after receiving three dupacks. We evaluate TCP-DCR through simulations to show that it offers significantly better performance when channel errors contribute more towards packet losses in the network with no or minimal impact on the performance when congestion is the primary cause for packet loss. We also present an analysis to show that the number of flows in the network significantly influences protocol evaluation in the wireless networks.
Sumitha Bhandarkar, Nauzad Erach Sadry, A. L. Narasimha Reddy, Nitin H. Vaidya
IEEE Trans. Mob. Comput.3
2005 Disk scheduling in a multimedia I/O system
abstract
This article provides a retrospective of our original paper by the same title in the Proceedings of the First ACM Conference on Multimedia, published in 1993. This article examines the problem of disk scheduling in a multimedia I/O system. In a multimedia server, the disk requests may have constant data rate requirements and need guaranteed service. We propose a new scheduling algorithm, SCAN-EDF, that combines the features of SCAN type of seek optimizing algorithm with an Earliest Deadline First (EDF) type of real-time scheduling algorithm. We compare SCAN-EDF with other scheduling strategies and show that SCAN-EDF combines the best features of both SCAN and EDF. We also investigate the impact of buffer space on the maximum number of video streams that can be supported.We show that by making the deadlines larger than the request periods, a larger number of streams can be supported.We also describe how we extended the SCAN-EDF algorithm in the PRISM multimedia architecture. PRISM is an integrated multimedia server, designed to satisfy the QOS requirements of multiple classes of requests. Our experience in implementing the extended SCAN-EDF algorithm in a generic operating system is discussed and performance metrics and results are presented to illustrate how the SCAN-EDF extensions and implementation strategies have succeeded in meeting the QOS requirements of different classes of requests.
A. L. Narasimha Reddy, James C. Wyllie, Ravi Wijayaratne
ACM Trans. Multim. Comput. Commun. Appl.1
2004 Design and evaluation of a partial state router
abstract
We present the design and evaluation of a partial state router. A partial state router maintains a fixed amount of state irrespective of the number of flows served at the router. We show the practical feasibility of partial state routers by implementing a novel partial state scheme, LRU-FQ, on the Linux platform. We report on our experience in employing the developed LRU-FQ router in several realistic experiments. Our results show the effectiveness of LRU-FQ in controlling high-bandwidth traffic and providing better response times for Web traffic. We also present a detailed evaluation of the developed router to demonstrate the feasibility and scalability of partial state schemes.
Phani Gopal V. Achanta, A. L. Narasimha Reddy
ICC2
2004 Impact of bandwidth-delay product and non-responsive flows on the performance of queue management schemes
abstract
In this paper, we study the impact-of bandwidth-delay products and non-responsive flows on queue management schemes. Our focus is to understand the aggregate performance of various classes of traffic under different queue management schemes. Our work is motivated by the expected trends of increasing link capacities and increasing amounts of non-responsive traffic. In this paper, the impact on the performance of RED, RED with ECN enabled and droptail routers are investigated. Our study considers the aggregate bandwidth of different classes of traffic and the delays observed at the router.
Zhili Zhao, A. L. Narasimha Reddy
ICC2
2004 A Fast Rerouting Scheme for OSPF/IS-IS Networks
abstract
Most current backbone networks use link-state protocol, OSPF or IS-IS, as their intra-domain routing protocol. Link-state protocols perform global routing table update to route around the failures. It usually takes seconds. As real-time applications like VoIP emerge in recent years, there is a requirement for a fast rerouting mechanism to route around failures before all routers on the network update their routing tables. In addition, fast rerouting is more appropriate than global routing table update when failures are transient. We propose such a fast rerouting extension for link-state protocols. In our approach, when a link fails, the affected traffic is rerouted along a pre-computed rerouting path. In case rerouting cannot be done locally, the local router signals the minimal number of upstream routers to setup the rerouting path for rerouting. We propose algorithms that simplify the rerouting operation and the rerouting path setup. With a simple extension to the current link state protocols, our scheme can route around failures faster and involves minimal number of routers for rerouting.
Yong Liu 0014, A. L. Narasimha Reddy
ICCCN2
2004 TCP-DCR: Making TCP Robust to Non-congestion Events
Sumitha Bhandarkar, A. L. Narasimha Reddy
NETWORKING2
2004 Detecting Traffic Anomalies through Aggregate Analysis of Packet Header Data
Seong Soo Kim, A. L. Narasimha Reddy, Marina Vannucci
NETWORKING2
2004 A method for estimating the proportion of nonresponsive traffic at a router
abstract
In this paper, a scheme for estimating the proportion of the incoming traffic that is not responsive to congestion at a router is presented. The idea of the proposed scheme is that if the observed queue length and packet drop probability do not match the predictions from a model of responsive (TCP) traffic, then the error must come from nonresponsive traffic; it can then be used for estimating the proportion of nonresponsive traffic. The proposed scheme is based on the queue length history, packet drop history, and expected TCP and queue dynamics. The effectiveness of the proposed scheme over a wide range of traffic scenarios is corroborated using ns-2-based simulations. Potential applications of the proposed algorithms in traffic engineering and control are discussed.
Zhili Zhao, Swaroop Darbha, A. L. Narasimha Reddy
IEEE/ACM Trans. Netw.3
2003 MVSS: An Active Storage Architecture
abstract
This paper presents MVSS, a storage system for active storage devices MVSS offers a single framework for supporting various services at the device level. It provides a flexible interface for associating services to a file through multiple views of the file Similar to views of a database in a multiview database system, views in MVSS are generated dynamically and are not stored on physical storage devices. MVSS represents each view of an underlying file through a separate entry in the file system namespace. MVSS separates the deployment of services from file system implementations and, thus, allows services to be migrated to storage devices. The paper presents the design of MVSS and how different services can be supported in MVSS at the device level. To illustrate our approach, we implemented a prototype system on PCs running Linux. We present results from example applications implemented on the prototype and discuss a variety of architectural issues including mixed workloads.
Xiaonan Ma, A. L. Narasimha Reddy
IEEE Trans. Parallel Distributed Syst.2
2002 A method for estimating non-responsive traffic at a router
abstract
In this paper, we propose a scheme for estimating the proportion of the incoming traffic that is not responding to congestion at a router. The idea of the proposed scheme is that if the observed queue length and packet drop probability do not match with the predicted results from the TCP model, then the error must come from the non-responsive traffic; it can then be used for estimating non-responsive traffic. The proposed scheme utilizes queue length history, packet drop history, expected TCP and queue dynamics to estimate the proportion. We show that the proposed scheme is effective over a wide range of traffic scenarios through simulations.
Zhili Zhao, Jayesh Ametha, Swaroop Darbha, A. L. Narasimha Reddy
SIGMETRICS4
2001 Adaptive marking for aggregated flows
abstract
The differentiated services architecture is receiving wide attention as a framework for providing different levels of service according to a service profile in the Internet. The current architecture allows aggregation of flows sharing a service profile. This paper looks at the problem of achieving specific QoS goals of individual flows by flexibly managing resources available to an aggregated source. We derive a simple analytic model for the relationship between per-session behavior, aggregate packet marking and packet differentiation within a DiffServ network. The paper presents an adaptive marker based on a TCP performance model within a DiffServ network. The paper shows that an aggregated marker can maintain the state of individual flows at the edge of the network and utilize this state effectively in adaptively marking packets of individual flows to meet their QoS goals.
Ikjun Yeom, A. L. Narasimha Reddy
GLOBECOM2
2001 Identifying Long-Term High-Bandwidth Flows at a Router
Smitha, Inkoo Kim, A. L. Narasimha Reddy
HiPC3
2001 MVSS: Multi-view Storage System
abstract
Presents MVSS, a storage system for active storage devices. MVSS offers a single framework for supporting various services at the device level. It provides a flexible interface for associating services to a file through multiple views of the file. Similar to views of a database in a multi-view database system, views in MVSS are generated dynamically and are not stored on physical storage devices. MVSS represents each view of an underlying file through a separate entry in the file system namespace. MVSS separates the deployment of services from file system implementations and thus allows services to be migrated to the storage devices. The paper presents the design of MVSS and shows how different services can be supported in MVSS at the device level. To illustrate our approach, we implemented a prototype system on PCs running Linux. We present results from the prototype implementation to demonstrate the effectiveness of our approach.
Xiaonan Ma, A. L. Narasimha Reddy
ICDCS2
2001 System support for providing integrated services from networked multimedia storage servers
abstract
In this paper, we describe our experience in building an integrated multimedia storage system, Prism. Our current Linux-based implementation of Prism provides three levels of service: deadline guarantees for periodic applications, best-effort better response times for interactive applications and starvation-free throughput guarantees for aperiodic applications. Prism separates resource allocation from resource scheduling. Resource allocation is controlled across the service classes by a system-wide policy and service class specific admission controllers. Resource scheduling is done at the resources. This separation allows Prism to be deployed even when the storage system is separated on a network from the file system.We report on the important aspects of Prism architecture, innovations required to build Prism on top of Linux and lessons learned during the implementation and testing of Prism. We present experimental results to show that Prism achieves its goals in supporting multiple service classes within a single system. We compare Prism against standard Linux operating system to show the impact of our approach.
Ravi Wijayaratne, A. L. Narasimha Reddy
ACM Multimedia2
2001 SACRIO: an active buffer management scheme for differentiated service networks
abstract
In this paper, we propose an active resource management approach for D ifferentiated Services networks. The proposed approach, SACRIO, employs caching and localized packet remarking within a router. It is shown that SACRIO is simple and can be implemented transparently within the diff-serv architecture. It is shown that the packet handling cost remains O(1) with SACRIO. SACRIO is shown to be quite effective and scalable through both ns-based simulations and real-world trace-driven simulations.
Saikrishnan Gopalkrishnan, A. L. Narasimha Reddy
NOSSDAV2
2001 Marking for QoS improvement
Ikjun Yeom, A. L. Narasimha Reddy
Comput. Commun.2
2001 ENDE: An End-to-end Network Delay Emulator Tool for Multimedia Protocol Development
Ikjun Yeom, A. L. Narasimha Reddy
Multim. Tools Appl.2
2001 Modeling TCP behavior in a differentiated services network
abstract
The differentiated services architecture has been proposed for providing different levels of services and has received wide attention. A packet in a diff-serv domain is classified into a class of service according to its contract profile and treated differently by its class. While many studies have addressed issues on the diff-serv architecture (e.g., dropper, marker, classifier and shaper), there have been few attempts to analytically understand a flow's behavior in a diff-serv network. We propose simple models of TCP behavior in a diff-serv network. Our models quantitatively characterize TCP throughput as functions of the contract rate, the packet-drop rate and the round-trip time in either two-drop precedence or three-drop precedence network. We also extend our models to aggregated flows. The models are validated through a number of simulations.
Ikjun Yeom, A. L. Narasimha Reddy
IEEE/ACM Trans. Netw.2
2000 Providing QOS Guarantees for Disk I/O
Ravi Wijayaratne, A. L. Narasimha Reddy
Multim. Syst.2
2000 Performance Evaluation of Storage Systems Based on Network-Attached Disks
abstract
The emergence of network-attached disks provides the possibility of transferring data between the storage system and the client directly. This offers new possibilities in building a distributed storage system. In this paper, we examine different storage organizations based on network-attached disks and compare the performance of these systems to a traditional system. Trace-driven simulations are used to measure the average response times of the client requests in two different workloads. We show that the advantages of distributing the server's network processing load to disks can be offset by losses in cache hits in a system based on network-attached disks. The two workloads we considered highlight this impact. We show that it is possible to offload significant load from the server by utilizing network-attached disks and point out that specific applications may be able to better exploit the features of network-attached disks.
Gang Ma 0012, Adnan Khaleel, A. L. Narasimha Reddy
IEEE Trans. Parallel Distributed Syst.3
1999 Evaluation of Data and Request Distribution Policies in Clustered Servers
Adnan Khaleel, A. L. Narasimha Reddy
HiPC2
1999 A Client Oriented, IP Level Redirection Mechanism
abstract
This paper introduces a new approach for implementing transparent client access to network services. Ever increasing load on the Internet has made it essential to design services that are fast, reliable, easily manageable, transparent to access, and that can scale gracefully with load. A common way of achieving this has been replicating services across multiple servers and redirecting clients to different servers depending upon various criteria. Existing schemes are either entirely server or network based. This scheme involves the client network layer actively in redirection. The paper describes the redirection protocol in detail and a FreeBSD based implementation of a testbed. The performance of the mechanism is measured by experiments on the testbed and analyzed. The advantages and disadvantages of client based network level redirection are discussed and some useful applications that it enables are described.
A. L. Narasimha Reddy
INFOCOM2
1999 Real-Time Communication Scheduling in a Multicomputer Video Server
A. L. Narasimha Reddy, Eli Upfal
J. Parallel Distributed Comput.1
1998 An Evaluation of Storage Systems based on Network-attached Disks
abstract
The emergence of network-attached disks provides the possibility of transferring data between the storage system and the client directly. This offers new possibilities in building a distributed storage system. We examine different storage organizations based on network-attached disks and compare the performance of these systems to a traditional system. Trace driven simulations are used to measure the average response times of the client requests in two different workloads. The results indicate that it is possible to reduce the workload on the file manager and improve performance in some workloads. However, in other workloads, reduced cache hit ratios may offset the gains of distributing the network processing workload.
Gang Ma 0012, A. L. Narasimha Reddy
ICPP2
1997 Effectiveness of caching policies for a Web server
abstract
We look at a number of policies for managing a file system cache in a Web server. Traces from NCSA World Wide Web (WWW) server are used in this study. We study several caching policies for improving the overall performance of such a server and show that request response time can be improved by some of the replacement policies that take size of the request into account. It is shown that least frequently used (LFU) performs well when cache sizes are small (<16 MB) and SpacexAge replacement policy performs consistently better than LRU. We show that previously proposed cache sharing policies improve performance significantly in clustered Web servers.
A. L. Narasimha Reddy
HiPC1
1993 Failure Evaluation of Disk Array Organizations
abstract
The authors present an evaluation of some of the disk array organizations proposed in the literature. They evaluate three alternatives for sparing, hot sparing, distributed sparing, and parity sparing, and two options for data layout, regular RAID5 and block designs, and systems based on combinations of these data layout and sparing alternatives. The performance of these organizations is evaluated with different reconstruction strategies. It is shown that parity sparing and distributed sparing have better performance and shorter reconstruction times than hot sparing. It is shown that both block designs as a data layout policy and distributed sparing as a sparing policy reduce the reconstruction time after a failure. The impact of reconstruction strategies is studied, and it is shown that, at higher workloads, choice of reconstruction strategy has a significant impact on the performance of the systems.>
John A. Chandy, A. L. Narasimha Reddy
ICDCS2
1993 Disk Scheduling in a Multimedia I/O System
abstract
In this paper, we look at the problem of disk scheduling in a multimedia I/O system.In a multimedia server, the disk requests may have constant data rate requirements and need guaranteed service.We propose a new scheduling algorithm, SCAN-EDF, that combines the features of SCAN type of seek optimizing algorithms with Earliest Deadline First (EDF) type of real-time scheduling algorithms.We compare SCAN-EDF with other scheduling strategies and show that SCAN-EDF combines the best features of both SCAN and EDF.We also investigate the impact of buffer space on the maximum number of video streams that can be supported.We show that by making the deadlines larger than the request periods, a larger number of streams can be supported.
A. L. Narasimha Reddy, James C. Wyllie
ACM Multimedia1
1993 Design and Evaluation of Gracefully Degradable Disk Arrays
A. L. Narasimha Reddy, John A. Chandy, Prithviraj Banerjee
J. Parallel Distributed Comput.1
1992 A Study of I/O System Organizations
abstract
With the increasing processing speeds, it has become important to design powerful and efficient I/O systems. In this paper, we look at several design options in designing an I/O system and study their impact on the performance. Specifically, we use trace driven simulations to study a disk system with a nonvolatile cache. Some of the considered design parameters include the cache block size, the fetch size, the cache size and the disk access policy. We show that decoupling the fetch size and the cache block size results in significant performance improvements. A new write-back policy is presented that is shown to offer significant performance benefits. We show that optimal block size in a two-level memory hierarchy is dependent only on the latency, data rate product of the second level as previously conjectured. We also present results showing the effect of a split access operation of a disk read/write head.
A. L. Narasimha Reddy
ISCA1
1991 Compiler Support for Parallel I/O Operations
A. L. Narasimha Reddy, Prithviraj Banerjee, D. K. Chen
ICPP (2)1
1990 A Study of I/O Behavior of Perfect Benchmarks on a Multiprocessor
abstract
The I/O behavior of some scientific applications, a subset of Perfect benchmarks, executing on a multiprocessor is studied. The aim of this study is to explore the various patterns of I/O access of large scientific applications and to understand the impact of this observed behavior on the I/O subsystem architecture. I/O behavior of the program is characterized by the demands it imposes on the I/O subsystem. It is observed that implicit I/O or paging is not a major problem for the applications considered and the I/O problem is mainly manifest in the explicit I/O done in the program. Various characteristics of I/O accesses are studied and their impact on architecture design is discussed.
A. L. Narasimha Reddy, Prithviraj Banerjee
ISCA1
1990 Algorithms-Based Fault Detection for Signal Processing Applications
abstract
The increasing demands for high-performance signal processing along with the availability of inexpensive high-performance processors have results in numerous proposals for special-purpose array processors for signal processing applications. A functional-level concurrent error-detection scheme is presented for such VLSI signal processing architectures as those proposed for the FFT and QR factorization. Some basic properties involved in such computations are used to check the correctness of the computed output values. This fault-detection scheme is shown to be applicable to a class of problems rather than a particular problem, unlike the earlier algorithm-based error-detection techniques. The effects of roundoff/truncation errors due to finite-precision arithmetic are evaluated. It is shown that the error coverage is high with large word sizes.>
A. L. Narasimha Reddy, Prithviraj Banerjee
IEEE Trans. Computers1
1990 Design, Analysis, and Simulation of I/O Architectures for Hypercube
abstract
Several issues concerning the design of an I/O (input/output) system for a multiprocessor such as a hypercube are examined. A methodology is proposed for connecting the I/O processors to such a system for efficient I/O access. The effect of I/O communication on the multiprocessor network is analyzed. Different disk organizations that can be employed within such a system are evaluated to see which organization has a better performance. It is observed that parallelism in serving an I/O request plays a dominant role in the scientific workload. The problem of mapping specific data structures such as matrices onto the disks so that the data can be accessed efficiently is considered.>
A. L. Narasimha Reddy, Prithviraj Banerjee
IEEE Trans. Parallel Distributed Syst.1
1989 Performance Evaluation of Multiple-Disk I/O Systems
A. L. Narasimha Reddy, Prithviraj Banerjee
ICPP (1)1
1989 I/O issues for hypercubes
abstract
In this paper, we look at several issues concerning the design of a disk system for a multiprocessor such as a hypercube. We propose a methodology for connecting the I/O processors to such a system for efficient I/O access. An analysis is presented to see the effect of I/O communication on the network of the multiprocessor. We evaluate different disk organizations that can be employed within such a system to see which organization has a better performance. Then we consider the problem of mapping specific data structures such as matrices onto the disks such that the data can be accessed efficiently.
A. L. Narasimha Reddy, Prithviraj Banerjee
ICS1
1989 An Evaluation of Multiple-Disk I/O Systems
abstract
Alternative ways of configuring an I/O subsystem with multiple disks to improve the I/O performance are considered. Specifically, the author consider disk synchronization, data declustering/disk striping, and a combination of both these approaches. They evaluate many different organizations that have not been considered before. The effects of block size and other parameters of the system are examined. Two different workloads are considered for the evaluation: a file/transaction system workload and a scientific applications workload. Through simulations it is shown that synchronized organizations perform better than other organizations at very low request rates; that there is a tradeoff in the amount of declustering/synchronization to be used in a system; and that systems with higher parallelism in reading a file perform better in a scientific workload.>
A. L. Narasimha Reddy, Prithviraj Banerjee
IEEE Trans. Computers1
1988 I/O Embedding in Hypercubes
A. L. Narasimha Reddy, Prithviraj Banerjee
ICPP (1)1
1987 A Fault Secure Dictionary Machine
abstract
A fault-secure dictionary machine is presented in this paper. The symmetry of the binary tree architecture for a dictionary machine is exploited to obtain fault-secureness with little overhead. The proposed design utilizes only one extra processor and with some other modifications to the structure of the processors, can detect a single failure of a processor or a link. The proposed design keeps two copies of each record and whenever a record is extracted from the machine the two copies are compared to detect if any fault has occurred. The low overhead is a result of observing the fact that at any given time all the processors of the machine need not be active and these potentially-idle processors are used to do redundant processing to enable detecting single faults.
A. L. Narasimha Reddy, Prithviraj Banerjee
ICDE1
1986 An Implementation of Mixed-Radix Conversion for Residue Number Applications
abstract
A method of residue number system (RNS) conversion to mixed-radix (MR) representation is presented. This method is found to be cost-effective and efficient, particularly for moduli size 4/5 bits. A comparison of conversion times and hardware necessary for RNS conversion to MR digits based on different methods is also presented.
N. B. Chakraborti, John S. Soundararajan, A. L. Narasimha Reddy
IEEE Trans. Computers3