VLDB 2026 Research / reviewers in the wild / expert
Chita R. Das
dblp:d/ChitaRDas
· DBLP profile ↗
239ranked-venue papers
6as first author
30since 2021 · last 2025
0000-0002-4746-7578ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 190 · 5 first-author · 22 since 2021Software engineering, systems software and programming languages · 36 · 6 since 2021Computer networks · 24 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 5 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Security and privacy · 6Databases, data management, data science and information retrieval · 6 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Salient Store: Enabling Smart Storage for Continuous Learning Edge ServersabstractGlobal IP-video traffic is projected to exceed 4.8 Zettabytes annually by 2025, driven largely by edge applications like autonomous driving and urban mobility that generate hundreds of terabytes per device. While significant research has optimized edge inference and training architectures, the critical challenges of data archival and storage have been largely overlooked. Our analysis reveals that archival operations, not inference, dominate system resources, consuming up to 80% of memory bandwidth and one-third of CPU cycles in edge servers. We present Salient Store, a novel computational storage architecture that transforms passive storage devices into active participants in the video analytics pipeline. Salient Store integrates motion-aware layered neural compression directly within storage FPGAs, reusing feature maps from inference to eliminate redundant computation while employing anchor-delta temporal encoding to maximize compression efficiency. By executing these operations entirely within the storage plane through peer-to-peer communication, our system bypasses host memory bottlenecks that plague traditional architectures. Comprehensive evaluation across five real-world datasets demonstrates that Salient Store reduces end-to-end archival latency by $6.18 \times$, decreases host-side data movement by $5.63 \times$, and maintains up to 47 dB PSNR while reducing system power consumption by 65%. This work fundamentally re-imagines storage for continuous learning systems, transforming it from a passive bottleneck into an acceleration layer that operates symbiotically with neural inference and training. Cyan Subhra Mishra, Deeksha Chaudhary, Mahmut T. Kandemir, Chita R. Das |
PACT | 4 |
| 2025 | Load and MLP-Aware Thread Orchestration for Recommendation Systems Inference on CPUs
Teyuh Chou, Onur Kayiran, John Kalamatianos, Gabriel H. Loh, Mahmut T. Kandemir, Chita R. Das |
ASPLOS (2) | 7 |
| 2025 | Pirate: No Compromise Low-Bandwidth VR Streaming for Edge DevicesabstractDue to the limited compute power and storage capabilities of edge platforms, ''streaming'' often provides a better VR experience compared to ''rendering''. Yet, achieving high-quality VR streaming faces two significant challenges, namely, bandwidth limitations and the need for real-time operation with high frames per second (FPS). Previous efforts have tended to prioritize either conserving bandwidth without real-time performance or ensuring real-time operation without substantial bandwidth savings. In this work, we incorporate the concept of ''stereo similarity'' to develop a novel real-time stereo video compression framework for streaming, called Pirate. Unlike the previously proposed approaches that rely on large machine learning-based models for synthesizing stereo pairs from both eyes with disparity maps (which can be impractical for most edge platforms due to their high computational cost), Pirate iteratively synthesizes the target eye view using only a single eye view and its corresponding disparity and optical flow information, with alternating left or right eye transmission. This enables us to generate target view at an extremely low computational cost, even under bandwidth constraints as low as 0.1 bits per pixel (bpp), while maintaining a high frame rate of 90 FPS. Our evaluations also reveal that, the proposed approach not only achieves real-time VR streaming with a 20%-40% reduction in bandwidth usage, but also maintains similar superior quality standards. Yingtian Zhang, Ziyu Ying 0001, Wanhang Lu, Sijie Lan, Huijuan Xu 0001, Kiwan Maeng, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das |
ASPLOS (2) | 10 |
| 2025 | FLEXI: Phase-Aware Function Resizing for Heterogeneous Serverless GPU Workloads
Shruti Mohanty, Vivek M. Bhasi, Jashwant Raj Gunasekaran, Prashanth Thinakaran, Mahmut T. Kandemir, Chita R. Das |
IEEE Big Data | 6 |
| 2025 | Dally: A Network-Placement Sensitive Cluster Scheduler for Deep Learning
Aakash Sharma, Vivek M. Bhasi, Sonali Singh, Mahmut T. Kandemir, George Kesidis, Chita R. Das |
IEEE Big Data | 6 |
| 2025 | NExUME: Adaptive Training and Inference for DNNs under Intermittent Power EnvironmentsabstractThe deployment of Deep Neural Networks (DNNs) in energy-constrained environments, such as Energy Harvesting Wireless Sensor Networks (EH-WSNs), introduces significant challenges due to the intermittent nature of power availability. This study introduces NExUME, a novel training methodology designed specifically for DNNs operating under such constraints. We propose a dynamic adjustment of training parameters—dropout rates and quantization levels—that adapt in real-time to the available energy, which varies in energy harvesting scenarios.
This approach utilizes a model that integrates the characteristics of the network architecture and the specific energy harvesting profile. It dynamically adjusts training strategies, such as the intensity and timing of dropout and quantization, based on predictions of energy availability. This method not only conserves energy but also enhances the network’s adaptability, ensuring robust learning and inference capabilities even under stringent power constraints. Our results show a 6% to 22% improvement in accuracy over current methods, with an increase of less than 5% in computational overhead. This paper details the development of the adaptive training framework, describes the integration of energy profiles with dropout and quantization adjustments, and presents a comprehensive evaluation using real-world data. Additionally, we introduce a novel dataset aimed at furthering the application of energy harvesting in computational settings. Cyan Subhra Mishra, Deeksha Chaudhary, Jack Sampson, Mahmut T. Kandemir, Chita R. Das |
ICLR | 5 |
| 2025 | GSCoder: Enabling Fast and Efficient Encoding for Game Streaming ApplicationsabstractThe recent proliferation of cloud gaming (also referred to as game streaming), enabling high-fidelity gaming quality on edge devices without high-end hardware, promises a transformative and democratized gaming experience across diverse populations. Yet, streaming high definition (4K UHD) game frames to thin-client devices, especially mobile, requires substantially higher bandwidth than traditional video streaming and often results in frame drops, degrading user experience. We identified that this high bandwidth demand arises from inefficient compression of game frames with real-time compression requirement (60 frames per second (FPS)) because the rapid, irregular motion in game frames violates the predictability assumptions baked into standard video motion estimation algorithms.To address this issue, we propose and evaluate GSCoder, a realtime efficient game frame compression framework that utilizes motion cues from the game’s rendering pipeline to directly acquire and optimize encoding-compliant motion vectors precisely, bypassing the costly motion estimation step inherent in standard encoders. Our evaluation, conducted in five open-source games with varying motion complexities, demonstrates that GSCoder achieves on average 49% and 19% higher compression efficiency than state-of-the-art (SOTA) game and video encoders, respectively, while maintaining real-time performance and delivering high-quality streams. Moreover, GSCoder delivers at least $3.5 \times$ encoding speedup compared to SOTA video encoders. Sandeepa Bhuyan, Ziyu Ying 0001, Vivek M. Bhasi, Mahmut T. Kandemir, Chita R. Das |
MASCOTS | 5 |
| 2025 | FPGA-based accelerator for adaptive banded event alignment in nanopore sequencing data analysisabstractAdaptive Banded Event Alignment (ABEA) stands as a critical algorithmic component in sequence polishing and DNA methylation detection, employing dynamic programming to align raw Nanopore signal with reference reads. Motivated by the observation that, compared to CPUs and GPUs, cutting-edge FPGAs demonstrate—in certain cases—superior performance at a reduced cost and energy consumption, this paper presents an efficient FPGA-based accelerator for ABEA, leveraging the inherent high parallelism and sequential access pattern within ABEA. Our proposed FPGA-based ABEA accelerator significantly enhances ABEA performance compared to the original CPU-based implementation in Nanopolish as well as the state-of-art acceleration on GPU and FPGA platforms. Specifically, targeting Xilinx VU9P, our accelerator achieves an average throughput speedup of 10.05 $$\times$$ over the CPU-only implementation, an average 1.81 $$\times$$ speedup over the state-of-art GPU acceleration with only 7.2% of the energy, and a speedup of 10.11 $$\times$$ compared to an existing FPGA accelerator. Our work demonstrates that intensive genome analysis can benefit significantly from cutting-edge FPGAs, offering improvements in both performance and energy consumption. Yilin Feng, Gulsum Gudukbay Akbulut, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Chita R. Das |
BMC Bioinform. | 6 |
| 2024 | FAAStloop: Optimizing Loop-Based Applications for Serverless ComputingabstractServerless Computing has garnered significant interest for executing High-Performance Computing (HPC) applications in recent years, attracting attention for its elastic scalability, reduced entry barriers, and pay-per-use pricing model. Specifically, highly parallel HPC apps can be divided and offloaded to multiple Serverless Functions (SFs) that execute their respective tasks concurrently and, finally, their results are stored/aggregated. While state-of-the-art userside serverless frameworks have attempted to fine-tune task division amongst the SFs to optimize for performance and/or cost, they have either used static task division parameters or have only focused on minimizing the number of SFs through task packing. However, these methods treat the HPC code as a black-box and usually require significant manual intervention to find the optimal task division. Since a significant portion of the HPC applications have a loop structure, in this work, we try to answer the following two questions: (i) Can modifying the loop structure in the HPC code, originally optimized for monolithic (non-serverless) frameworks, enhance performance and reduce costs in a serverless architecture?, and (ii) Can we develop a framework that allows for an efficient transition of monolithic code to serverless, with minimum user input? Shruti Mohanty, Vivek M. Bhasi, Myungjun Son, Mahmut T. Kandemir, Chita R. Das |
SoCC | 5 |
| 2024 | Usas: A Sustainable Continuous-Learning' Framework for Edge ServersabstractEdge servers have recently become very popular for performing localized analytics, especially on video, as they reduce data traffic and protect privacy. However, due to their resource constraints, these servers often employ compressed models, which are typically prone to data drift. Consequently, for edge servers to provide cloud-comparable quality, they must also perform continuous learning to mitigate this drift. However, at expected deployment scales, performing continuous training on every edge server is not sustainable due to their aggregate power demands on grid supply and associated sustainability footprints. To address these challenges, we propose Us.as,´ an approach combining algorithmic adjustments, hardware-software co-design, and morphable acceleration hardware to enable the training of workloads on these edge servers to be powered by renewable, but intermittent, solar power that can sustainably scale alongside data sources. Our evaluation of Us.as on a real-world´ traffic dataset indicates that our continuous learning approach simultaneously improves both accuracy and efficiency: Us.as´ offers a 4.96% greater mean accuracy than prior approaches while our morphable accelerator that adapts to solar variance can save up to {234.95kWH, 2.63MWH}/year/edge-server compared to a {DNN accelerator, data center scale GPU}, respectively. Cyan Subhra Mishra, Jack Sampson, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Chita R. Das |
HPCA | 5 |
| 2024 | Foveated HDR: Efficient HDR Content Generation on Edge Devices Leveraging User's Visual AttentionabstractIn recent years, high dynamic range (HDR) content has become increasingly popular for its ability to represent a broader brightness range, enhancing the realism and immersion in applications like augmented reality/virtual reality (AR/VR) on edge devices. While DNN-based solutions are effective for reconstructing high-fidelity HDR content, due to their high computational demands and memory usage, the DNN-based HDR reconstruction takes up to several seconds to generate one HDR image, making it very challenging to deploy such techniques onto the edge devices. Ziyu Ying 0001, Sandeepa Bhuyan, Yingtian Zhang, Mahmut T. Kandemir, Chita R. Das |
ICCAD | 6 |
| 2024 | Paldia: Enabling SLO-Compliant and Cost-Effective Serverless Computing on Heterogeneous HardwareabstractAmong the variety of applications (apps) being deployed on serverless platforms, apps such as Machine Learning (ML) inference serving can achieve better performance from leveraging accelerators like GPUs. Yet, major serverless providers, despite having GPU-equipped servers, do not offer GPU support for their serverless functions. Given that serverless functions are deployed on various generations of CPUs already, extending this to various (typically more expensive) GPU generations can offer providers a greater range of hardware to serve incoming requests according to the functions and request traffic. Here, providers are faced with the challenge of selecting hardware to reach a well-proportioned trade-off point between cost and performance. While recent works have attempted to address this, they often fail to do so as they overlook optimization opportunities arising from intelligently leveraging existing GPU sharing mechanisms. To address this point, we devise a heterogeneous serverless framework, PALDIA, which uses a prudent Hardware selection policy to acquire capable, cost-effective hardware and perform intelligent request scheduling on it to yield high performance and cost savings. Specifically, our scheduling algorithm employs hybrid spatio-temporal GPU sharing that intelligently trades off job queueing delays and interference to allow the chosen cost-effective hardware to also be highly performant. We extensively evaluate PALDIA using 16 ML inference workloads with real-world traces on a 6 node heterogeneous cluster. Our results show that PALDIA significantly outperforms state-of-the-art works in terms of Service Level Objective (SLO) compliance (up to 13.3% more) and tail latency (up to ∼50% less), with cost savings up to 86%. Vivek M. Bhasi, Aakash Sharma, Shruti Mohanty, Mahmut T. Kandemir, Chita R. Das |
IPDPS | 5 |
| 2024 | GameStreamSR: Enabling Neural-Augmented Game Streaming on Commodity Mobile PlatformsabstractCloud gaming (also referred to as Game Streaming) is a rapidly emerging application that is changing the way people enjoy video games. However, if the user demands a high-resolution (e.g., 2 K or 4 K) stream, the game frames require high bandwidth and the stream often suffers from a significant number of frame drops due to network congestion degrading the Quality of Experience (QoE). Recently, the DNN-based Super Resolution (SR) technique has gained prominence as a practical alternative for streaming low-resolution frames and upscaling them at the client for enhanced video quality. However, performing such DNN-based tasks on resource-constrained and battery-operated mobile platforms is very expensive and also fails to meet the real-time requirement (60 frames per second (FPS)). Unlike traditional video streaming, where the frames can be downloaded and buffered, and then upscaled by their playback turn, Game Streaming is real-time and interactive, where the frames are generated on the fly and cannot tolerate high latency/lags for frame upscaling. Thus, state-of-the-art (SOTA) DNN-based SR cannot satisfy the mobile Game Streaming requirements. Towards this, we propose GameStreamSR, a framework for enabling real-time Super Resolution for Game Streaming applications on mobile platforms. We take visual perception nature into consideration and propose to only apply DNN-based SR to the regions with high visual importance and upscale the remaining regions using traditional solutions such as bilinear interpolation. Especially, we leverage the depth data from the game rendering pipeline to intelligently localize the important regions, called regions of importance (RoI), in the rendered game frames. Our evaluation of ten popular games on commodity mobile platforms shows that our proposal can enable realtime (60 FPS) neurally-augmented SR. Our design achieves a $13 \times$ frame rate speedup (and $\approx 4 \times$ Motion-to-Photon latency improvement) for the reference frames and a $1.6 \times$ frame rate speedup for the non-reference frames, which translates to, on average $2 \times$ FPS performance improvement and 26-33% energy savings over the SOTA DNN-based SR execution, while achieving about 2dB PSNR gain and better perceptual quality than the current SOTA. Sandeepa Bhuyan, Ziyu Ying 0001, Mahmut T. Kandemir, Mahanth Gowda, Chita R. Das |
ISCA | 5 |
| 2024 | Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUsabstractPersonalized recommendation is a ubiquitous appli-cation on the internet, with many industries and hyperscalers extensively leveraging Deep Learning Recommendation Models (DLRMs) for their personalization needs (like ad serving or movie suggestions). With growing model and dataset sizes pushing computation and memory requirements, GPUs are being increasingly preferred for executing DLRM inference. However, serving newer DLRMs, while meeting acceptable latencies, continues to remain challenging, making traditional deployments increasingly more GPU-hungry, resulting in higher inference serving costs. In this paper, we show that the embedding stage continues to be the primary bottleneck in the GPU inference pipeline, leading up to a 3.2 x embedding-only performance slowdown. To thoroughly grasp the problem, we conduct a detailed microarchitecture characterization and highlight the presence of low occupancy in the standard embedding kernels. By leveraging direct compiler optimizations, we achieve optimal occupancy, pushing the performance by up to 53 %. Yet, long memory latency stalls continue to exist. To tackle this challenge, we propose spe-cialized plug-and-play-based software prefetching and L2 pinning techniques, which help in hiding and decreasing the latencies. Further, we propose combining them, as they complement each other. Experimental evaluations using AI00 GPUs with large models and datasets show that our proposed techniques improve performance by up to 103% for the embedding stage, and up to 77 % for the overall D LRM inference pipeline. Vivek M. Bhasi, Adwait Jog, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das |
MICRO | 6 |
| 2024 | Towards SLO-Compliant and Cost-Effective Serverless Computing on Emerging GPU ArchitecturesabstractServerless platforms are supporting an increasing variety of applications (apps). Among these, apps such as Machine Learning (ML) inference serving can benefit significantly from leveraging accelerators like GPUs. Yet, major serverless providers, despite having GPU-equipped servers, do not offer GPU support for their serverless functions. While recent works have attempted to bridge this gap, they are agnostic to the capabilities of new-generation GPUs, thereby, overlooking several performance optimization opportunities. Vivek M. Bhasi, Aakash Sharma, Jashwant Raj Gunasekaran, Ashutosh Pattnaik, Mahmut T. Kandemir, Chita R. Das |
Middleware | 7 |
| 2023 | Stash: A Comprehensive Stall-Centric Characterization of Public Cloud VMs for Distributed Deep LearningabstractDeep neural networks (DNNs) are increasingly popular owing to their ability to solve complex problems such as image recognition, autonomous driving, and natural language processing. Their growing complexity coupled with the use of larger volumes of training data (to achieve acceptable accuracy) has warranted the use of GPUs and other accelerators. Such accelerators are typically expensive, with users having to pay a high upfront cost to acquire them. For infrequent use, users can, instead, leverage the public cloud to mitigate the high acquisition cost. However, with the wide diversity of hardware instances (particularly GPU instances) available in public cloud, it becomes challenging for a user to make an appropriate choice from a cost/performance standpoint. In this work, we try to address this problem by (i) introducing a comprehensive distributed deep learning (DDL) profiler Stash, which determines the various execution stalls that DDL suffers from, and (ii) using Stash to extensively characterize various public cloud GPU instances by running popular DNN models on them. Specifically, it estimates two types of communication stalls, namely, interconnect and network stalls, that play a dominant role in DDL execution time. Stash is implemented on top of prior work, DS-analyzer, that computes only the CPU and disk stalls. Using our detailed stall characterization, we list the advantages and shortcomings of public cloud GPU instances for users to help them make an informed decision(s). Our characterization results indicate that the more expensive GPU instances may not be the most performant for all DNN models and that AWS can sometimes sub-optimally allocate hardware interconnect resources. Specifically, the intra-machine interconnect can introduce communication overheads of up to 90% of DNN training time and the network-connected instances can suffer from up to 5× slowdown compared to training on a single instance. Furthermore, (iii) we also model the impact of DNN macroscopic features such as the number of layers and the number of gradients on communication stalls, and finally, (iv) we briefly discuss a cost comparison with existing work. Aakash Sharma, Vivek M. Bhasi, Sonali Singh, Jashwant Raj Gunasekaran, Subrata Mitra, Mahmut T. Kandemir, George Kesidis, Chita R. Das |
ICDCS | 9 |
| 2023 | EdgePC: Efficient Deep Learning Analytics for Point Clouds on Edge DevicesabstractRecently, point cloud (PC) has gained popularity in modeling various 3D objects (including both synthetic and real-life) and has been extensively utilized in a wide range of applications such as AR/VR, 3D reconstruction, and autonomous driving. For such applications, it is critical to analyze/understand the surrounding scenes properly. To achieve this, deep learning based methods (e.g., convolutional neural networks (CNNs)) have been widely employed for higher accuracy. Unlike the deep learning on conventional 2D images/videos, where the feature computation (matrix multiplication) is the major bottleneck, in point cloud-based CNNs, the sample and neighbor search stages are the primary bottlenecks, and collectively contribute to 54% (up to 80%) of the overall execution latency on a typical edge device. While prior efforts have attempted to solve this issue by designing custom ASICs or pipelining the neighbor search with other stages, to our knowledge, none of them has tried to "structurize" the unstructured PC data for improving computational efficiency. Ziyu Ying 0001, Sandeepa Bhuyan, Yingtian Zhang, Mahmut T. Kandemir, Chita R. Das |
ISCA | 6 |
| 2023 | Optimizing CPU Performance for Recommendation Systems At-ScaleabstractDeep Learning Recommendation Models (DLRMs) are very popular in personalized recommendation systems and are a major contributor to the data-center AI cycles. Due to the high computational and memory bandwidth needs of DLRMs, specifically the embedding stage in DLRM inferences, both CPUs and GPUs are used for hosting such workloads. This is primarily because of the heavy irregular memory accesses in the embedding stage of computation that leads to significant stalls in the CPU pipeline. As the model and parameter sizes keep increasing with newer recommendation models, the computational dominance of the embedding stage also grows, thereby, bringing into question the suitability of CPUs for inference. In this paper, we first quantify the cause of irregular accesses and their impact on caches and observe that off-chip memory access is the main contributor to high latency. Therefore, we exploit two well-known techniques: (1) Software prefetching, to hide the memory access latency suffered by the demand loads and (2) Overlapping computation and memory accesses, to reduce CPU stalls via hyperthreading to minimize the overall execution time. We evaluate our work on a single-core and 24-core configuration with the latest recommendation models and recently released production traces. Our integrated techniques speed up the inference by up to 1.59x, and on average by 1.4x. Scott Cheng, Vishwas Kalagi, Vrushabh Sanghavi, Samvit Kaul, Meena Arunachalam, Kiwan Maeng, Adwait Jog, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das |
ISCA | 11 |
| 2022 | SandPiper: A Cost-Efficient Adaptive Framework for Online Recommender SystemsabstractOnline recommender systems have proven to have ubiquitous applications in various domains. To provide accurate recommendations in real time it is imperative to constantly train and deploy models with the latest data samples. This retraining involves adjusting the model weights by incorporating newly-arrived streaming data into the model to bridge the accuracy gap. To provision resources for the retraining, typically the compute is hosted on VMs, however, due to the dynamic nature of the data arrival patterns, stateless functions would be an ideal alternative over VMs, as they can instantaneously scale on demand. However, it is non-trivial to statically configure the stateless functions because the model retraining exhibits varying resource needs during different phases of retraining. Therefore, it is crucial to dynamically configure the functions to meet the resource requirements, while bridging the accuracy gap. In this paper, we propose Sandpiper, an adaptive framework that leverages stateless functions to deliver accurate predictions at low cost for online recommender systems. The three main ideas in Sandpiper are (i) we design a data-drift monitor that automatically triggers model retraining at required time intervals to bridge the accuracy gap due to incoming data drifts; (ii) we develop an online configuration model that selects the appropriate function configurations while maintaining the model serving accuracy within the latency and cost budget; and (iii) we propose a dynamic synchronization policy for stateless functions to speed up the distributed model retraining leading to cloud cost minimization. A prototype implementation on AWS shows that Sandpiper maintains the average accuracy above 90%, while 3.8× less expensive than the traditional VM-based schemes. Prashanth Thinakaran, Kanak Mahadik, Jashwant Raj Gunasekaran, Mahmut T. Kandemir, Chita R. Das |
IEEE Big Data | 5 |
| 2022 | Cypress: input size-sensitive container provisioning and request scheduling for serverless platformsabstractThe growing popularity of the serverless platform has seen an increase in the number and variety of applications (apps) being deployed on it. The majority of these apps process user-provided input to produce the desired results. Existing work in the area of input-sensitive profiling has empirically shown that many such apps have input size-dependent execution times which can be determined through modelling techniques. Nevertheless, existing serverless resource management frameworks are agnostic to the input size-sensitive nature of these apps. We demonstrate in this paper that this can potentially lead to container over-provisioning and/or end-to-end Service Level Objective (SLO) violations. To address this, we propose Cypress, an input size-sensitive resource management framework, that minimizes the containers provisioned for apps, while ensuring a high degree of SLO compliance. We perform an extensive evaluation of Cypress on top of a Kubernetes-managed cluster using 5 apps from the AWS Serverless Application Repository and/or Open-FaaS Function Store with real-world traces and varied input size distributions. Our experimental results show that Cypress spawns up to 66% fewer containers, thereby, improving container utilization and saving cluster-wide energy by up to 2.95X and 23%, respectively, versus state-of-the-art frameworks, while remaining highly SLO-compliant (up to 99.99%). Vivek M. Bhasi, Jashwant Raj Gunasekaran, Aakash Sharma, Mahmut T. Kandemir, Chita R. Das |
SoCC | 5 |
| 2022 | Exploiting Frame Similarity for Efficient Inference on Edge DevicesabstractDeep neural networks (DNNs) are being widely used in various computer vision tasks as they can achieve very high accuracy. However, the large number of parameters employed in DNNs can result in long inference times for vision tasks, thus making it even more challenging to deploy them in the compute- and memory-constrained mobile/edge devices. To boost the inference of DNNs, some existing works employ compression (model pruning or quantization) or enhanced hardware. How-ever, most prior works focus on improving model structure and implementing custom accelerators. As opposed to the prior work, in this paper, we target the video data that are processed by edge devices, and study the similarity between frames. Based on that, we propose two runtime approaches to boost the performance of the inference process, while achieving high accuracy.Specifically, considering the similarities between successive video frames, we propose a frame-level compute reuse algorithm based on the motion vectors of each frame. With frame-level reuse, we are able to skip 53% of frames in inference with negligible overhead and remain within less than 1% mAP (accuracy) drop for the object detection task. Additionally, we implement a partial inference scheme to enable region/tile-level reuse. Our experiments on a representative mobile device (Pixel 3 Phone) show that the proposed partial inference scheme achieves 2 × speedup over the baseline approach that performs full inference on every frame. We integrate these two data reuse algorithms to accelerate the neural network inference and improve its energy efficiency. More specifically, for each frame in the video, we can dynamically select between (i) performing a full inference, (ii) performing a partial inference, or (iii) skipping the inference altogether. Our experimental evaluations using six different videos reveal that the proposed schemes are up to 80% (56% on average) energy efficient and 2.2× performance efficient compared to the conventional scheme, which performs full inference, while losing less than 2% accuracy. Additionally, the experimental analysis indicates that our approach outperforms the state-of-the-art work with respect to accuracy and/or performance/energy savings. Ziyu Ying 0001, Shulin Zhao 0001, Haibo Zhang 0005, Cyan Subhra Mishra, Sandeepa Bhuyan, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das |
ICDCS | 8 |
| 2022 | Pushing Point Cloud Compression to the EdgeabstractAs Point Clouds (PCs) gain popularity in processing millions of data points for 3D rendering in many applications, efficient data compression becomes a critical issue. This is because compression is the primary bottleneck in minimizing the latency and energy consumption of existing PC pipelines. Data compression becomes even more critical as PC processing is pushed to edge devices with limited compute and power budgets. In this paper, we propose and evaluate two complementary schemes, intra-frame compression and inter-frame compression, to speed up the PC compression, without losing much quality or compression efficiency. Unlike existing techniques that use sequential algorithms, our first design, intra-frame compression, exploits parallelism for boosting the performance of both geometry and attribute compression. The proposed parallelism brings around $43.7 \times$ performance improvement and 96.6% energy savings at a cost of $1.01 \times$ larger compressed data size. To further improve the compression efficiency, our second scheme, inter-frame compression, considers the temporal similarity among the video frames and reuses the attribute data from the previous frame for the current frame. We implement our designs on an NVIDIA Jetson AGX Xavier edge GPU board. Experimental results with six videos show that the combined compression schemes provide $34.0 \times$ speedup compared to a state-of-the-art scheme, with minimal impact on quality and compression ratio. Ziyu Ying 0001, Shulin Zhao 0001, Sandeepa Bhuyan, Cyan Subhra Mishra, Mahmut T. Kandemir, Chita R. Das |
MICRO | 6 |
| 2022 | Skipper: Enabling efficient SNN training through activation-checkpointing and time-skippingabstractSpiking neural networks (SNNs) are a highly efficient signal processing mechanism in biological systems that have inspired a plethora of research efforts aimed at translating their energy efficiency to computational platforms. Efficient training approaches are critical for the successful deployment of SNNs. Compared to mainstream deep neural networks (ANNs), training SNNs is far more challenging due to complex neural dynamics that evolve with time and their discrete, binary computing paradigm. Back-propagation-through-time (BPTT) with surrogate gradients has recently emerged as an effective technique to train deep SNNs directly. SNN-BPTT, however, has a major drawback in that it has a high memory requirement that increases with the number of timesteps. SNNs generally result from the discretization of Ordinary Differential Equations, due to which the sequence length must be typically longer than RNNs, compounding the time dependence problem. It, therefore, becomes hard to train deep SNNs on a single or multi-GPU setup with sufficiently large batch sizes or timesteps, and extended periods of training are required to achieve reasonable network performance. In this work, we reduce the memory requirements of BPTT in SNNs to enable the training of deeper SNNs with more timesteps (T). For this, we leverage the notion of activation re-computation in the context of SNN training that enables the GPU memory to scale sub-linearly with increasing time-steps. We observe that naively deploying the re-computation based approach leads to a considerable computational overhead. To solve this, we propose a time-skipped BPTT approximation technique, called Skipper, for SNNs, that not only alleviates this computation overhead, but also lowers memory consumption further with little to no loss of accuracy. We show the efficacy of our proposed technique by comparing it against a popular method for memory footprint reduction during training. Our evaluations on 5 state-of-the-art networks and 4 datasets show that for a constant batch size and time-steps, skipper reduces memory usage by 3.3× to 8.4× (6.7× on average) over baseline SNN-BPTT. It also achieves a speedup of 29% to 70% over the checkpointed approach and of 4% to 40% over the baseline approach. For a constant memory budget, skipper can scale to an order of magnitude higher timesteps compared to baseline SNN-BPTT. Sonali Singh, Anup Sarma, Sen Lu, Abhronil Sengupta, Mahmut T. Kandemir, Emre Neftci, Narayanan Vijaykrishnan, Chita R. Das |
MICRO | 8 |
| 2022 | Cocktail: A Multidimensional Optimization for Model Serving in Cloud
Jashwant Raj Gunasekaran, Cyan Subhra Mishra, Prashanth Thinakaran, Bikash Sharma, Mahmut T. Kandemir, Chita R. Das |
NSDI | 6 |
| 2021 | CASH: A Credit Aware Scheduling for Public Cloud PlatformsabstractDistributed data processing frameworks such as Hadoop, Tez, Spark, and Flink are exclusively used by public cloud tenants for executing large scale data analytics applications in various domains including but not limited to content management, financial sector, healthcare etc. These frameworks slice a job into a number of smaller tasks, which are then executed by a job scheduler on a multi-node compute cluster. While making scheduling decisions, the State-of-art schedulers employed in these frameworks assume hardware resources such as CPU, disk I/O and network I/O to offer a fixed service rate. However, in a public cloud environment, many of these resources are associated with burstable service rates. More specifically, the resources offer a guaranteed baseline service rate with an option to burst above their baseline rate by expending accumulated burst credits. Being unaware about this underlying hardware burstability, schedulers tend to make sub-optimal task placement decisions, thereby adversely affecting the job completion times, leading to higher deployment costs.In this paper, we propose CASH, a burst credit aware scheduler, which is cognizant about the burst credits associated with the individual hardware resources in the public cloud cluster. Through coarse grained task annotations depicting the burst credit demand of individual tasks and dynamically monitoring the credits for the underlying resources, CASH performs optimal task placement decisions. We prototype CASH on YARN, Hadoop, and Tez, and extensively evaluate it using both batch and streaming workloads. Our experimental results with CASH show CPU-credit based instances, like AWS T3, are a viable cost effective alternative when compared to self-managed offerings like Amazon EMR, for running large scale batch workloads. Furthermore, we demonstrate that CASH can accelerate streaming SQL queries on a large Hive database by up to 39.4% , leading to public cloud cost savings by up to 22%. Aakash Sharma, Saravanan Dhakshinamurthy, George Kesidis, Chita R. Das |
CCGRID | 4 |
| 2021 | Kraken: Adaptive Container Provisioning for Deploying Dynamic DAGs in Serverless PlatformsabstractThe growing popularity of microservices has led to the proliferation of online cloud service-based applications, which are typically modelled as Directed Acyclic Graphs (DAGs) comprising of tens to hundreds of microservices. The vast majority of these applications are user-facing, and hence, have stringent SLO requirements. Serverless functions, having short resource provisioning times and instant scalability, are suitable candidates for developing such latency-critical applications. However, existing serverless providers are unaware of the workflow characteristics of application DAGs, leading to container over-provisioning in many cases. This is further exacerbated in the case of dynamic DAGs, where the function chain for an application is not known a priori. Motivated by these observations, we propose Kraken, a workflow-aware resource management framework that minimizes the number of containers provisioned for an application DAG while ensuring SLO-compliance. We design and implement Kraken on OpenFaaS and evaluate it on a multi-node Kubernetes-managed cluster. Our extensive experimental evaluation using DeathStarbench workload suite and real-world traces demonstrates that Kraken spawns up to 76% fewer containers, thereby improving container utilization and saving cluster-wide energy by up to 4x and 48%, respectively, when compared to state-of-the art schedulers employed in serverless platforms. Vivek M. Bhasi, Jashwant Raj Gunasekaran, Prashanth Thinakaran, Cyan Subhra Mishra, Mahmut T. Kandemir, Chita R. Das |
SoCC | 6 |
| 2021 | GSSA: A Resource Allocation Scheme Customized for 3D NAND SSDsabstractThe high density of 3D NAND-based SSDs comes with longer write latencies due to the increasing program complexity. To address this write performance degradation issue, NAND flash manufacturers implement a 3D NAND-specific full-sequence program (FSP) operation. The FSP can program multiple-bit information into a cell simultaneously with the same latency as the baseline program operation, thereby dramatically boosting the write performance. However, directly adopting the (large granularity) FSP operation in SSD firmware can result in a lifetime degradation problem, where small writes are amplified to large granularities with a significant fraction of empty data. This problem cannot completely be mitigated by the DRAM buffer in the SSDs since the “sync” commands from the host prevent the DRAM buffer from accumulating enough written data. To solve this FSP-induced performance/lifetime dilemma, in this work, we propose and evaluate GSSA (Generalized and Specialized Scramble Allocation), a novel written-data allocation scheme in SSD firmware, which considers both various 3D NAND program operations and the internal 3D NAND flash architecture. By adopting GSSA, SSDs can enjoy the performance benefits brought by the FSP without excessively consuming the lifetime. Our experimental evaluations reveal that GSSA can achieve the throughput and the spent-lifetime of the best-performance and best-lifetime single granularity schemes, respectively. Chun-Yi Liu 0002, Yunju Lee, Wonil Choi, Myoungsoo Jung, Mahmut T. Kandemir, Chita R. Das |
HPCA | 6 |
| 2021 | Gesture-SNN: Co-optimizing accuracy, latency and energy of SNNs for neuromorphic vision sensorsabstractAs originally published figures in the document were missing. A corrected replacement file was provided by the authors.Spiking neural networks (SNNs) are recently gaining popularity due to their low-power, spatio-temporal computing paradigm as opposed to more conventional deep learning approaches that mainly focus on spatial characteristics of data. When paired with biologically-inspired asynchronous event sensors, they can create energy-efficient near-sensor systems that are ideal for mobile, resource-constrained and embedded-computing scenarios. Training deep SNNs, however, is challenging due to their discrete nature. The most successful method so far involves training deep artificial neural networks (ANNs) using Gradient-Descent and then converting them to SNNs. The ANN-to-SNN conversion technique has mostly been evaluated on standard static image datasets using rate-based encoding of spikes. In this work, we find that a direct application of the ANN-to-SNN conversion technique to process event data via SNNs leads to arbitrary accuracy losses. Through insights gained from theoretical analyses as well as empirical observations, we propose three novel techniques to restore the conversion accuracy on event data and show proof-of-concept results, comparable to the state-of-the-art, on the IBM DVS Gesture dataset. Further exploration of the SNN design space reveals additional insights to fine-tune the accuracy-latency-peak power trade-off. Finally, we evaluate our proposed schemes on an existing neuromorphic accelerator and show that our best-performing model is $\sim 38$% more accurate with $\sim 35$% lower energy and $\sim 55$% lower EDP compared to its traditional SNN counterpart. Sonali Singh, Anup Sarma, Sen Lu, Abhronil Sengupta, Narayanan Vijaykrishnan, Chita R. Das |
ISLPED | 6 |
| 2021 | HoloAR: On-the-fly Optimization of 3D Holographic Processing for Augmented RealityabstractHologram processing is the primary bottleneck and contributes to more than 50% of energy consumption in battery-operated augmented reality (AR) headsets. Thus, improving the computational efficiency of the holographic pipeline is critical. The objective of this paper is to maximize its energy efficiency without jeopardizing the hologram quality for AR applications. Towards this, we take the approach of analyzing the workloads to identify approximation opportunities. We show that, by considering various parameters like region of interest and depth of view, we can approximate the rendering of the virtual object to minimize the amount of computation without affecting the user experience. Furthermore, by optimizing the software design flow, we propose HoloAR, which intelligently renders the most important object in sight to the clearest detail, while approximating the computations for the others, thereby significantly reducing the amount of computation, saving energy, and gaining performance at the same time. We implement our design in an edge GPU platform to demonstrate the real-world applicability of our research. Our experimental results show that, compared to the baseline, HoloAR achieves, on average, 2.7 × speedup and 73% energy savings. Shulin Zhao 0001, Haibo Zhang 0005, Cyan Subhra Mishra, Sandeepa Bhuyan, Ziyu Ying 0001, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das |
MICRO | 8 |
| 2021 | Structured in Space, Randomized in Time: Leveraging Dropout in RNNs for Efficient TrainingabstractRecurrent Neural Networks (RNNs), more specifically their Long Short-Term Memory (LSTM) variants, have been widely used as a deep learning tool for tackling sequence-based learning tasks in text and speech. Training of such LSTM applications is computationally intensive due to the recurrent nature of hidden state computation that repeats for each time step. While sparsity in Deep Neural Nets has been widely seen as an opportunity for reducing computation time in both training and inference phases, the usage of non-ReLU activation in LSTM RNNs renders the opportunities for such dynamic sparsity associated with neuron activation and gradient values to be limited or non-existent. In this work, we identify dropout induced sparsity for LSTMs as a suitable mode of computation reduction. Dropout is a widely used regularization mechanism, which randomly drops computed neuron values during each iteration of training. We propose to structure dropout patterns, by dropping out the same set of physical neurons within a batch, resulting in column (row) level hidden state sparsity, which are well amenable to computation reduction at run-time in general-purpose SIMD hardware as well as systolic arrays. We provide a detailed analysis of how the dropout-induced sparsity propagates through the different stages of network training and how it can be leveraged in each stage. More importantly, our proposed approach works as a direct replacement for existing dropout-based application settings. We conduct our experiments for three representative NLP tasks: language modelling on the PTB dataset, OpenNMT based machine translation using the IWSLT De-En and En-Vi datasets, and named entity recognition sequence labelling using the CoNLL-2003 shared task. We demonstrate that our proposed approach can be used to translate dropout-based computation reduction into reduced training time, with improvement ranging from 1.23$\times$ to 1.64$\times$, without sacrificing the target metric. Anup Sarma, Sonali Singh, Huaipan Jiang, Rui Zhang 0037, Mahmut T. Kandemir, Chita R. Das |
NeurIPS | 6 |
| 2020 | Multiverse: Dynamic VM Provisioning for Virtualized High Performance Computing ClustersabstractTraditionally, HPC workloads have been deployed in bare-metal clusters; but the advances in virtualization have led the pathway for these workloads to be deployed in virtualized clusters. However, HPC cluster administrators/providers still face challenges in terms of resource elasticity and virtual machine (VM) provisioning at large-scale, due to the lack of coordination between a traditional HPC scheduler and the VM hypervisor (resource management layer). This lack of interaction leads to low cluster utilization and job completion throughput. Furthermore, the VM provisioning delays directly impact the overall performance of jobs in the cluster. Hence, there is a need for effectively provisioning virtualized HPC clusters, which can best-utilize the physical hardware with minimal provisioning overheads.Towards this, we propose Multiverse, a VM provisioning framework, which can dynamically spawn VMs for incoming jobs in a virtualized HPC cluster, by integrating the HPC scheduler along with VM resource manager. We have implemented this framework on the Slurm scheduler along with the vSphere VM resource manager. In order to reduce the VM provisioning overheads, we use instant cloning which shares both the disk and memory with the parent VM, when compared to full VM cloning which has to boot-up a new VM from scratch. Measurements with real-world HPC workloads demonstrate that, instant cloning is 2.5× faster than full cloning in terms of VM provisioning time. Further, it improves resource utilization by up to 40%, and cluster throughput by up to 1.5×, when compared to full clone for bursty job arrival scenarios. Jashwant Raj Gunasekaran, Michael Cui, Prashanth Thinakaran, Josh Simons, Mahmut T. Kandemir, Chita R. Das |
CCGRID | 6 |
| 2020 | Characterizing Bottlenecks in Scheduling Microservices on Serverless PlatformsabstractDatacenters are witnessing an increasing trend in adopting microservice-based architecture for application design, which consists of a combination of different microservices. Typically these applications are short-lived and are administered with strict Service Level Objective (SLO) requirements. Traditional virtual machine (VM) based provisioning for such applications not only suffers from long latency when provisioning resources (as VMs tend to take a few minutes to start up), but also places an additional overhead of server management and provisioning on the users. This led to the adoption of serverless functions, where applications are composed as functions and hosted in containers. However, state-of-the-art schedulers employed in serverless platforms tend to look at microservice-based applications similar to conventional monolithic black-box applications. To detect all the inefficiencies, we characterize the end-to-end life cycle of these microservice-based applications in this work. Our findings show that the applications suffer from poor scheduling of microservices due to reactive container provisioning during workload fluctuations, thereby resulting in either in SLO violations or colossal container over-provisioning, in turn leading to poor resource utilization. We also find that there is an ample amount of slack available at each stage of application execution, which can potentially be leveraged to improve the overall application performance. Jashwant Raj Gunasekaran, Prashanth Thinakaran, Nachiappan Chidambaram Nachiappan, Ram Srivatsa Kannan, Mahmut T. Kandemir, Chita R. Das |
ICDCS | 6 |
| 2020 | NEBULA: A Neuromorphic Spin-Based Ultra-Low Power Architecture for SNNs and ANNsabstractBrain-inspired cognitive computing has so far followed two major approaches - one uses multi-layered artificial neural networks (ANNs) to perform pattern-recognition-related tasks, whereas the other uses spiking neural networks (SNNs) to emulate biological neurons in an attempt to be as efficient and fault-tolerant as the brain. While there has been considerable progress in the former area due to a combination of effective training algorithms and acceleration platforms, the latter is still in its infancy due to the lack of both. SNNs have a distinct advantage over their ANN counterparts in that they are capable of operating in an event-driven manner, thus consuming very low power. Several recent efforts have proposed various SNN hardware design alternatives, however, these designs still incur considerable energy overheads.In this context, this paper proposes a comprehensive design spanning across the device, circuit, architecture and algorithm levels to build an ultra low-power architecture for SNN and ANN inference. For this, we use spintronics-based magnetic tunnel junction (MTJ) devices that have been shown to function as both neuro-synaptic crossbars as well as thresholding neurons and can operate at ultra low voltage and current levels. Using this MTJ-based neuron model and synaptic connections, we design a low power chip that has the flexibility to be deployed for inference of SNNs, ANNs as well as a combination of SNN-ANN hybrid networks - a distinct advantage compared to prior works. We demonstrate the competitive performance and energy efficiency of the SNNs as well as hybrid models on a suite of workloads. Our evaluations show that the proposed design, NEBULA, is up to 7.9× more energy efficient than a state-of-the-art design, ISAAC, in the ANN mode. In the SNN mode, our design is about 45× more energy-efficient than a contemporary SNN architecture, INXS. Power comparison between NEBULA ANN and SNN modes indicates that the latter is at least 6.25× more power-efficient for the observed benchmarks. Sonali Singh, Anup Sarma, Nicholas Jao, Ashutosh Pattnaik, Sen Lu, Kezhou Yang, Abhronil Sengupta, Narayanan Vijaykrishnan, Chita R. Das |
ISCA | 9 |
| 2020 | Déjà View: Spatio-Temporal Compute Reuse for' Energy-Efficient 360° VR Video StreamingabstractThe emergence of virtual reality (VR) and augmented reality (AR) has revolutionized our lives by enabling a 360° artificial sensory stimulation across diverse domains, including, but not limited to, sports, media, healthcare, and gaming. Unlike the conventional planar video processing, where memory access is the main bottleneck, in 360° VR videos the compute is the primary bottleneck and contributes to more than 50% energy consumption in battery-operated VR headsets. Thus, improving the computational efficiency of the video processing pipeline in a VR is critical. While prior efforts have attempted to address this problem through acceleration using a GPU or FPGA, none of them has analyzed the 360° VR pipeline to examine if there is any scope to optimize the computation with known techniques such as memoization.Thus, in this paper, we analyze the VR computation pipeline and observe that there is significant scope to skip computations by leveraging the temporal and spatial locality in head orientation and eye correlations, respectively, resulting in computation reduction and energy efficiency. The proposed Déjà View design takes advantage of temporal reuse by memoizing head orientation and spatial reuse by establishing a relationship between left and right eye projection, and can be implemented either on a GPU or an FPGA. We propose both software modifications for existing compute pipeline and microarchitectural additions for further enhancement. We evaluate our design by implementing the software enhancements on an NVIDIA Jetson TX2 GPU board and our microarchitectural additions on a Xilinx Zynq-7000 FPGA model using five video workloads. Experimental results show that Déjà View can provide 34% computation reduction and 17% energy saving, compared to the state-of-the-art design. Shulin Zhao 0001, Haibo Zhang 0005, Sandeepa Bhuyan, Cyan Subhra Mishra, Ziyu Ying 0001, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das |
ISCA | 8 |
| 2020 | Fifer: Tackling Resource Underutilization in the Serverless EraabstractDatacenters are witnessing a rapid surge in the adoption of serverless functions for microservices-based applications. A vast majority of these microservices typically span less than a second, have strict SLO requirements, and are chained together as per the requirements of an application. The aforementioned characteristics introduce a new set of challenges, especially in terms of container provisioning and management, as the state-of-the-art resource management frameworks, employed in serverless platforms, tend to look at microservice-based applications similar to conventional monolithic applications. Hence, these frameworks suffer from microservice agnostic scheduling and colossal container over-provisioning, especially during workload fluctuations, thereby resulting in poor resource utilization. Jashwant Raj Gunasekaran, Prashanth Thinakaran, Nachiappan Chidambaram Nachiappan, Mahmut T. Kandemir, Chita R. Das |
Middleware | 5 |
| 2020 | Selective Caching: Avoiding Performance Valleys in Massively Parallel ArchitecturesabstractEmerging general purpose graphics processing units (GPGPU) make use of a memory hierarchy very similar to that of modern multi-core processors - they typically have multiple levels of on-chip caches and a DDR-like off-chip main memory. In such massively parallel architectures, caches are expected to reduce the average data access latency by reducing the number of off-chip memory accesses; however, our extensive experimental studies confirm that not all applications utilize the on-chip caches in an efficient manner. Even though GPGPUs are adopted to run a wide range of general purpose applications, the conventional cache management policies are incapable of achieving the optimal performance over different memory characteristics of the applications. This paper first investigates the underlying reasons for inefficiency of common cache management policies in GPGPUs. To address and resolve those issues, we then propose (i) a characterization mechanism to analyze each kernel at runtime and, (ii) a selective caching policy to manage the flow of cache accesses. Evaluation results of the studied platform show that our proposed dynamically reconfigurable cache hierarchy improves the system performance by up to 105% (average of 27%) over a wide range of modern GPGPU applications, which is within 10% of the optimal improvement. Amin Jadidi, Mahmut T. Kandemir, Chita R. Das |
PDP | 3 |
| 2019 | Spock: Exploiting Serverless Functions for SLO and Cost Aware Resource Procurement in Public CloudabstractWe are witnessing the emergence of elastic web services which are hosted in public cloud infrastructures. For reasons of cost-effectiveness, it is crucial for the elasticity of these web services to match the dynamically-evolving user demand. Traditional approaches employ clusters of virtual machines (VMs) to dynamically scale resources based on application demand. However, they still face challenges such as higher cost due to over-provisioning or incur service level objective (SLO) violations due to under-provisioning. Motivated by this observation, we propose Spock, a new scalable and elastic control system that exploits both VMs and serverless functions to reduce cost and ensure SLO for elastic web services. We show that under two different scaling policies, Spock reduces SLO violations of queries by up to 74% when compared to VM-based resource procurement schemes. Further, Spock yields significant cost savings, by up to 33% compared to traditional approaches which use only VMs. Jashwant Raj Gunasekaran, Prashanth Thinakaran, Mahmut T. Kandemir, Bhuvan Urgaonkar, George Kesidis, Chita R. Das |
CLOUD | 6 |
| 2019 | SOML Read: Rethinking the Read Operation Granularity of 3D NAND SSDsabstractNAND-based solid-state disks (SSDs) are known for their superior random read/write performance due to the high degrees of multi-chip parallelism they exhibit. Currently, as the chip density increases dramatically, fewer 3D NAND chips are needed to build an SSD compared to the previous generation chips. As a result, SSDs can be made more compact. However, this decrease in the number of chips also results in reduced overall throughput, and prevents 3D NAND high density SSDs from being widely-adopted. We analyzed 600 storage workloads, and our analysis revealed that the small read operations suffer significant performance degradation due to reduced chip-level parallelism in newer 3D NAND SSDs. The main question is whether some of the inter-chip parallelism lost in these new SSDs (due to the reduced chip count) can be won back by enhancing intra-chip parallelism. Motivated by this question, we propose a novel SOML (Single-Operation-Multiple-Location) read operation, which can perform several small intra-chip read operations to different locations simultaneously, so that multiple requests can be serviced in parallel, thereby mitigating the parallelism-related bottlenecks. A corresponding SOML read scheduling algorithm is also proposed to fully utilize the SOML read. Our experimental results with various storage workloads indicate that, the SOML read-based SSD with 8 chips can outperform the baseline SSD with 16 chips. Chun-Yi Liu 0002, Jagadish Kotra, Myoungsoo Jung, Mahmut T. Kandemir, Chita R. Das |
ASPLOS | 5 |
| 2019 | Kube-Knots: Resource Harvesting through Dynamic Container Orchestration in GPU-based DatacentersabstractCompute heterogeneity is increasingly gaining prominence in modern datacenters due to the addition of accelerators like GPUs and FPGAs. We observe that datacenter schedulers are agnostic of these emerging accelerators, especially their resource utilization footprints, and thus, not well equipped to dynamically provision them based on the application needs. We observe that the state-of-the-art datacenter schedulers fail to provide fine-grained resource guarantees for latency-sensitive tasks that are GPU-bound. Specifically for GPUs, this results in resource fragmentation and interference leading to poor utilization of allocated GPU resources. Furthermore, GPUs exhibit highly linear energy efficiency with respect to utilization and hence proactive management of these resources is essential to keep the operational costs low while ensuring the end-to-end Quality of Service (QoS) in case of user-facing queries.Towards addressing the GPU orchestration problem, we build Knots, a GPU-aware resource orchestration layer and integrate it with the Kubernetes container orchestrator to build Kube- Knots. Kube-Knots can dynamically harvest spare compute cycles through dynamic container orchestration enabling co-location of latency-critical and batch workloads together while improving the overall resource utilization. We design and evaluate two GPU-based scheduling techniques to schedule datacenter-scale workloads through Kube-Knots on a ten node GPU cluster. Our proposed Correlation Based Prediction (CBP) and Peak Prediction (PP) schemes together improves both average and 99thpercentile cluster-wide GPU utilization by up to 80% in case of HPC workloads. In addition, CBP+PP improves the average job completion times (JCT) of deep learning workloads by up to 36% when compared to state-of-the-art schedulers. This leads to 33% cluster-wide energy savings on an average for three different workloads compared to state-of-the-art GPU-agnostic schedulers. Further, the proposed PP scheduler guarantees the end-to-end QoS for latency-critical queries by reducing QoS violations by up to 53% when compared to state-of-the-art GPU schedulers. Prashanth Thinakaran, Jashwant Raj Gunasekaran, Bikash Sharma, Mahmut T. Kandemir, Chita R. Das |
CLUSTER | 5 |
| 2019 | Understanding Energy Efficiency in IoT App ExecutionsabstractBillions of Internet-of-Things (IoT) devices such as sensors, actuators, computing units, etc., are connected to form IoT platforms. However, it is observed that such hardware platforms today spend a significant proportion of their energy in communication between the CPU and the sensors (which are controlled by micro-controller unit (MCU)). Motivated by this observation, two simple, yet effective, optimizations are proposed to minimize the energy consumption. The first optimization, called Batching, interrupts the CPU after collecting multiple sensor data points at the MCU (instead of only 1), and thus, minimizes the interrupt overheads. The second optimization, called Computation Offloading to MCU (COM), offloads app-specific computations to the MCU to minimize data transfer overheads, and makes use of the relatively low energy footprint, and low-compute capabilities of the MCU in place of the CPU in the hub. However, questions such as why these two schemes are needed, where the energy benefit comes from, which IoT apps are suitable for these optimizations, etc., remain unclear. To better understand the Batching and COM approaches towards energy efficiency in IoT app executions, we characterize ten representative workloads on a Raspberry Pi and ESP8266 MCU platform, and evaluate the energy savings using these two optimizations and illustrate that for light-weight workloads (where COM is applicable), Batching and COM reduce the energy consumption by 52% and 85%, respectively when compared to the baseline. And for heavy-weight apps (where COM is not possible due to limited capacity of MCU), by offloading the light-weight apps and batching for the heavy-weight, Batching + COM (BCOM) benefits 10% energy savings compared to the baseline. Shulin Zhao 0001, Prasanna Venkatesh Rengasamy, Haibo Zhang 0005, Sandeepa Bhuyan, Nachiappan Chidambaram Nachiappan, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das |
ICDCS | 8 |
| 2019 | Opportunistic computing in GPU architecturesabstractData transfer overhead between computing cores and memory hierarchy has been a persistent issue for von Neumann architectures and the problem has only become more challenging with the emergence of manycore systems. A conceptually powerful approach to mitigate this overhead is to bring the computation closer to data, known as Near Data Computing (NDC). Recently, NDC has been investigated in different flavors for CPU-based multicores, while the GPU domain has received little attention. In this paper, we present a novel NDC solution for GPU architectures with the objective of minimizing on-chip data transfer between the computing cores and Last-Level Cache (LLC). To achieve this, we first identify frequently occurring Load-Compute-Store instruction chains in GPU applications. These chains, when offloaded to a compute unit closer to where the data resides, can significantly reduce data movement. We develop two offloading techniques, called LLC-Compute and Omni-Compute. The first technique, LLC-Compute, augments the LLCs with computational hardware for handling the computation offloaded to them. The second technique (Omni-Compute) employs simple bookkeeping hardware to enable GPU cores to compute instructions offloaded by other GPU cores. Our experimental evaluations on nine GPGPU workloads indicate that the LLC-Compute technique provides, on an average, 19% performance improvement (IPC), 11% performance/watt improvement, and 29% reduction in on-chip data movement compared to the baseline GPU design. The Omni-Compute design boosts these benefits to 31%, 16% and 44%, respectively. Ashutosh Pattnaik, Xulong Tang, Onur Kayiran, Adwait Jog, Asit K. Mishra, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das |
ISCA | 8 |
| 2019 | Distilling the Essence of Raw Video to Reduce Memory Usage and Energy at Edge DevicesabstractVideo broadcast and streaming are among the most widely used applications for edge devices. Roughly 82% of the mobile internet traffic is made up of video data. This is likely to worsen with the advent of 5G that will open up new opportunities for high resolution videos, virtual and augmented reality-based applications. The raw video data produced and consumed by edge devices is considerably higher than what is transmitted out of them. This leads to huge memory bandwidth and energy requirements from such edge devices. Therefore, optimizing the memory bandwidth and energy consumption needs is imperative for further improvements in energy efficiency of such edge devices. In this paper, we propose two mechanisms for on-the-fly compression and approximation of raw video data that is generated by the image sensors. The first mechanism, MidVB, performs lossless compression of the video frames coming out of the sensors and stores the compressed format into the memory. The second mechanism, Distill, builds on top of MidVB and further reduces memory consumption by approximating the video frame data. On an average, across 20 raw videos, MidVB and Distill are able to reduce the memory bandwidth by 43% and 72%, respectively, over the raw representation. They outperform a well known memory saving mechanism by 7% and 36%, respectively. Furthermore, MidVB and Distill reduce the energy consumption by 40% and 67%, respectively, over the baseline. Haibo Zhang 0005, Shulin Zhao 0001, Ashutosh Pattnaik, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das |
MICRO | 6 |
| 2018 | The Curious Case of Container Orchestration and Scheduling in GPU-based DatacentersabstractModern data centers are increasingly being provisioned with compute accelerators such as GPUs, FPGAs and ASIC's to catch up with the workload performance demands and reduce the total cost of ownership (TCO). By 2021, traffic within hyperscale datacenters is expected to quadruple with 94% of workloads moving to cloud-based datacenters according to Cisco's global cloud index. A majority of these workloads include data mining, image processing, speech recognition and gaming which uses GPUs for high throughput computing. This trend is evident as public cloud operators like Amazon and Microsoft have started to offer GPU-based infrastructure services in the recent times. Prashanth Thinakaran, Jashwant Raj Gunasekaran, Bikash Sharma, Mahmut T. Kandemir, Chita R. Das |
SoCC | 5 |
| 2018 | FLOSS: FLOw sensitive scheduling on mobile platformsabstractToday's mobile platforms have grown in sophistication to run a wide variety of frame-based applications. To deliver better QoS and energy efficiency, these applications utilize multi-flow execution, which exploits hardware-level parallelism across participating accelerators in the SoC. Our study shows that multi-flow execution increases memory pressure, and motivates us to propose a rate-based memory-scheduling scheme, called FLOSS, that considers a flow, individual frames of a flow, and any sharing of IPs across concurrent flows to schedule memory requests. Experimental results indicate that FLOSS provides 12% QoS improvement over baseline FR-FCFS scheme, and outperforms two QoS-aware schemes in multi-flow execution scenarios. Haibo Zhang 0005, Prasanna Venkatesh Rengasamy, Nachiappan Chidambaram Nachiappan, Shulin Zhao 0001, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das |
DAC | 7 |
| 2018 | Parallelizing garbage collection with I/O to improve flash resource utilizationabstractGarbage Collection (GC) has been a critical optimization target for improving the performance of flash-based Solid State Drives (SSDs); the long-lasting GC process occupies the flash resources, thereby blocking normal I/O requests and increasing response times. This is a well-documented problem, and a wide range of prior works successfully hide the negative impact of GC on the 1/O response times. In this paper, however, we unveil another serious side-effect of GC, called the plane under-utilization problem. More specifically, while a plane is busy doing GC, the other plane(s) in the same die remain idle, as all the planes in a die share a single command and address path that is dedicated to the GC. We also note that most of the state-of-the-art proposals attacking the GC impact on I/O response times are not able to resolve the plane under-utilization problem, and in turn, miss a great potential to further improve the SSD performance. Thus, we next propose a scheduling technique, I/O-parallelized GC, which leverages the idle planes during GC to serve the blocked I/O requests. As a result, flash resources (planes) can be active during the most of GC time and the blocked I/O requests can get serviced quickly, and in turn, an improved SSD performance can be achieved. Using simulation-based evaluations over a wide variety of workloads, we show that the proposed I/O-parallelized GC scheme can improve the response times of the GC-affected I/O requests by 83% (reads) and 70% (writes), by increasing the average plane utilization from the (two planes-per-die) baseline 50% to 74.4% during GC. The I/O-parallelized GC is orthogonal to prior proposals that hide GC overheads; so, they can be combined for further SSD performance improvement. Wonil Choi, Myoungsoo Jung, Mahmut T. Kandemir, Chita R. Das |
HPDC | 4 |
| 2018 | Content Popularity-Based Selective Replication for Read Redirection in SSDsabstractDespite high degrees of parallelism in terms of the number of chips and channels on state-of-the-art SSDs, resource contention continues to be a big impediment to boosting their performance for both read and write requests. This is particularly significant in the delays due to queueing for service from individual NAND-flash chips that can take dozens/hundreds of microseconds to perform the read/write operations. Owing to the no-write-in-place policy that is employed in flash chips, writes are inherently suited to be redirected to chips with lower load, in case their original destination chip is overloaded. However, to date, there has been no work to redirect read requests, since they cannot be serviced by other chips, which do not have the data. While blindly replicating all the data everywhere seems very promising from a read redirection perspective, doing so results in high space overheads, high write/replication overheads and lower endurance. This paper presents a novel approach to selective replication, wherein the popularity of data is used to figure out the "what", "how much", "where" and "when" questions for replication. Leveraging value locality/popularity, that is often observed in practice, popular data is replicated across multiple chips to provide more opportunities for dynamic read redirection to less loaded flash chips. Using extensive workload traces running over weeks from real systems, we show that our Read Redirected SSD (RR-SSD) can provide up to 45% improvement in read performance, with average improvement of 23.9%, and up to 40% improvement when considering both read and write requests, with 16% improvement on average. Nima Elyasi, Mohammad Arjomand, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das |
MASCOTS | 5 |
| 2018 | Tolerating Write Disturbance Errors in PCM: Experimental Characterization, Analysis, and MechanismsabstractConstant technology scaling has enabled modern computing systems to achieve high degrees of thread-level parallelism, making the design of a highly scalable and dense main memory subsystem a major challenge, especially for data-intensive workloads. While during the past three decades DRAM has been widely used as the dominant technology to build main memory, it faces serious scalability and power consumption problems at sub-micron scales. Phase Change Memory (PCM) has been proposed as one of the most promising technologies to replace DRAM in future computing systems, because of its short access latency and high scalability. However, PCM has some reliability problems below 20nm technology. As the cell size scales down, the thermal-disturbance between the cells is exacerbated. More precisely, the generated heat during the write operation can disseminate to the adjacent cells and potentially change their values. This phenomenon is known as the write disturbance problem. In this work, we study the impact of write disturbance along the word-line and bit-line in a PCM-based main memory. We then propose two low-overhead mechanisms to address this problem in different dimensions. Our proposed schemes are based on a rank subsetting memory architecture where each chip can be accessed separately. To mitigate the write disturbance problem along the word-line, we use a combination of differential write (DW) and verify-and-correct (VnC) schemes only if the chance of having cascaded verification steps is low. Otherwise, we write into all the vulnerable cells within the target chip to avoid performance loss. To tolerate write disturbance along the bit-line, we use data compression to compact the data and place the adjacent memory lines in a non-overlapping fashion. We further use the extra space within a compressed memory line to store BCH code, in order to protect read-intensive memory lines against write-intensive addresses. Our proposed schemes guarantee reliable write operations in a super dense PCM-based main memory, and improve the system performance by 17% on average, compared to the state-of-the-art mechanism in this domain. Amin Jadidi, Mahmut T. Kandemir, Chita R. Das |
MASCOTS | 3 |
| 2018 | CritICs Critiquing Criticality in Mobile AppsabstractIn this paper, we conduct a systematic analysis to show that existing CPU optimizations targeting scientific/server workloads are not always well suited for mobile apps. In particular, we observe that the well-known and very important concept of identifying and accelerating individual critical instructions in workloads such as SPEC, are not as effective for mobile apps. Several differences in mobile app characteristics including (i) dependencies between critical instructions interspersed with non-critical instructions in the dependence chain, (ii) temporal proximity of the critical instructions in the dynamic stream, and (iii) the bottleneck shifting to the front from the rear of the datapath pipeline, are key contributors to the ineffectiveness of traditional criticality based optimizations. Instead, we propose the concept of Critical Instruction Chains (CritICs) - which are short, critical and self contained sequences of instructions, for aggregate level optimization. With motivating results, we show that an offline profiler/analysis framework can easily identify these CritICs, and we propose a very simple software mechanism in the compiler that exploits ARM's 16-bit ISA format to nearly double the fetch bandwidth of these instructions. We have implemented this entire framework - both profiler and compiler passes, and evaluated its effectiveness for 10 popular apps from the Play Store. Experimental evaluations show that our approach is much more effective than two previously studied criticality optimizations, yielding a speedup of 12.65%, and energy savings of 15% in the CPU (translating to a system wide energy savings of 4.6%), requiring very little additional hardware support. Prasanna Venkatesh Rengasamy, Haibo Zhang 0005, Shulin Zhao 0001, Nachiappan Chidambaram Nachiappan, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das |
MICRO | 7 |
| 2018 | Performance and Power-Efficient Design of Dense Non-Volatile Cache in CMPsabstractIn this paper, we present a novel cache design based on Multi-Level Cell Spin-Transfer Torque RAM (MLC STT-RAM) that can dynamically adjust the set capacity and associativity to efficiently use the full potential of MLC STT-RAM technology. We exploit the asymmetric nature of the MLC storage scheme to build cache lines featuring heterogeneous performances, that is, half of the cache lines are read-friendly, while the other half are write-friendly. Furthermore, we propose to opportunistically deactivate cache ways in underutilized sets to convert MLC to Single-Level Cell (SLC) mode, which features overall better performance and lifetime. Our ultimate goal is to build a cache architecture that combines the capacity advantages of MLC and performance/energy advantages of SLC. Our experimental evaluations show an average improvement of 43 percent in total numbers of conflict misses, 27 percent in memory access latency, 12 percent in system performance, and 26 percent in L3 access energy, with a slight degradation in lifetime (about 7 percent) compared to an SLC cache. Amin Jadidi, Mohammad Arjomand, Mahmut T. Kandemir, Chita R. Das |
IEEE Trans. Computers | 4 |
| 2018 | Stochastic Modeling and Optimization of StragglersabstractMapReduce framework is widely used to parallelize batch jobs since it exploits a high degree of multi-tasking to process them. However, it has been observed that when the number of servers increases, the map phase can take much longer than expected. This paper analytically shows that the stochastic behavior of the servers has a negative effect on the completion time of a MapReduce job, and continuously increasing the number of servers without accurate scheduling can degrade the overall performance. We analytically model the map phase in terms of hardware, system, and application parameters to capture the effects of stragglers on the performance. Mean sojourn time (MST), the time needed to sync the completed tasks at a reducer, is introduced as a performance metric and mathematically formulated. Following that, we stochastically investigate the optimal task scheduling which leads to an equilibrium property in a datacenter with different types of servers. Our experimental results show the performance of the different types of schedulers targeting MapReduce applications. We also show that, in the case of mixed deterministic and stochastic schedulers, there is an optimal scheduler that can always achieve the lowest MST. Farshid Farhat, Diman Zad Tootaghaj, Yuxiong He, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das |
IEEE Trans. Cloud Comput. | 6 |
| 2017 | Exploiting Intra-Request Slack to Improve SSD PerformanceabstractWith Solid State Disks (SSDs) offering high degrees of parallelism, SSD controllers place data and direct requests to exploit the maximum offered hardware parallelism. In the quest to maximize parallelism and utilization, sub-requests of a request that are directed to different flash chips by the scheduler can experience differential wait times since their individual queues are not coordinated and load balanced at all times. Since the macro request is considered complete only when its last sub-request completes, some of its sub-requests that complete earlier have to necessarily wait for this last sub-request. This paper opens the door to a new class of schedulers to leverage such slack between sub-requests in order to improve response times. Specifically, the paper presents the design and implementation of a slack-enabled re-ordering scheduler, called Slacker, for sub-requests issued to each flash chip. Layered under a modern SSD request scheduler, Slacker estimates the slack of each incoming sub-request to a flash chip and allows them to jump ahead of existing sub-requests with sufficient slack so as to not detrimentally impact their response times. Slacker is simple to implement and imposes only marginal additions to the hardware. Using a spectrum of 21 workloads with diverse read-write characteristics, we show that Slacker provides as much as 19.5%, 13% and 14.5% improvement in response times, with average improvements of 12%, 6.5% and 8.5%, for write-intensive, read-intensive and read-write balanced workloads, respectively. Nima Elyasi, Mohammad Arjomand, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das, Myoungsoo Jung |
ASPLOS | 5 |
| 2017 | Co-training of Feature Extraction and Classification using Partitioned Convolutional Neural NetworksabstractThere are an increasing number of neuromorphic hardware platforms designed to efficiently support neural network inference tasks. However, many applications contain structured processing in addition to classification. Being able to map both neural network classification and structured computation onto the same platform is appealing from a system design perspective. In this paper, we perform a case study on mapping the feature extraction stage of pedestrian detection using Histogram of Oriented Gradients (HoG) onto a neuromophic platform. We consider three implementations: one that approximates HoG using neuromorphic intrinsics, one that emulates HoG outputs using a trained network, and one that allows feature extraction to be absorbed into classification. The proposed feature extraction methods are implemented and evaluated on neuromorphic hardware (IBM Neurosynaptic System). Our study shows that both a designed approximation and a "parroted" emulation can achieve similar accuracy, and that the latter appears to better capitalize on limited training and resource budgets, compared to the absorbed approach, while also being more power efficient than the programmed approach by a factor of 6.5x-208x. Wei-Yu Tsai, Jinhang Choi, Tulika Parija, Priyanka Gomatam, Chita R. Das, Jack Sampson, Narayanan Vijaykrishnan |
DAC | 5 |
| 2017 | Exploring the Potential for Collaborative Data Compression and Hard-Error Tolerance in PCM MemoriesabstractLimited write endurance is the main obstacle standing in the way of using phase change memory (PCM) in future computing systems. While several wear-leveling and hard-error tolerant techniques have been proposed for improving PCM lifetime, most of these approaches assume that the underlying memory uses a very simple write traffic reduction scheme (e.g., buffering, differential writes). In particular, most PCM prototypes/chips are equipped with an embedded circuit to support differential writes (DW) - on a write, only the bits that differ between the old and new data are updated. With DW, the bit-pattern of updates in a memory block is usually random, which limits the opportunity to exploit the resulting bit pattern for lifetime enhancement at an architecture level (e.g., using techniques such as wear-leveling and hard-error tolerance). This paper focuses on this inefficiency and proposes a solution based on data compression. Employing compression can improve the lifetime of the PCM memory. Using state-of-the-art compression schemes, the size of the compressed data is usually much smaller than the original data written back to memory from the last-level cache on an eviction. By storing data in a compressed format in the target memory block, first, we limit the number of bit flips to fewer memory cells, enabling more efficient intra-line wear-leveling and error recovery, and second, the unused bits in the memory block can be reused as replacements for faulty bits given the reduced size of the (compressed) data. It can also happen that for a portion of the memory blocks, the resulting compressed data is not very small. This can be due to increased data entropy introduced by compression, where the total number of bit flips will be increased over the baseline system. In this paper, we present an approach that provides collaborative operation of data compression, differential writes, wear-leveling and hard-error tolerant techniques targeting PCM memories. We propose approaches that reap the maximum benefits from compression, while also enjoying the benefits of techniques that reduce the number of high-entropy writes. Using an approach that combines different solutions, our mechanism tolerates 2.9× more cell failures per memory line and achieves a 4.3× increase in PCM memory lifetime, relative to our baseline state-of-the-art PCM DIMM memory. Amin Jadidi, Mohammad Arjomand, Mohammad Khavari Tavana, David R. Kaeli, Mahmut T. Kandemir, Chita R. Das |
DSN | 6 |
| 2017 | Controlled Kernel Launch for Dynamic Parallelism in GPUsabstractDynamic parallelism (DP) is a promising feature for GPUs, which allows on-demand spawning of kernels on the GPU without any CPU intervention. However, this feature has two major drawbacks. First, the launching of GPU kernels can incur significant performance penalties. Second, dynamically-generated kernels are not always able to efficiently utilize the GPU cores due to hardware-limits. To address these two concerns cohesively, we propose SPAWN, a runtime framework that controls the dynamically-generated kernels, thereby directly reducing the associated launch overheads and queuing latency. Moreover, it allows a better mix of dynamically-generated and original (parent) kernels for the scheduler to effectively hide the remaining overheads and improve the utilization of the GPU resources. Our results show that, across 13 benchmarks, SPAWN achieves 69% and 57% speedup over the flat (non-DP) implementation and baseline DP, respectively. Xulong Tang, Ashutosh Pattnaik, Huaipan Jiang, Onur Kayiran, Adwait Jog, Sreepathi Pai, Mohamed Assem Ibrahim, Mahmut T. Kandemir, Chita R. Das |
HPCA | 9 |
| 2017 | Leveraging value locality for efficient design of a hybrid cache in multicore processorsabstractOwing to negligible leakage current, high density and superior scalability, Spin-Transfer Torque RAM (STT-RAM) technology becomes one of the promising candidates for low power and high capacity on-chip caches in multicore systems. While STT-RAM read access latency is comparable to that of SRAM, write operations in STT-RAM are more challenging: writes are slow, consume a large energy, and the lifetime of STT-RAM is limited by the number of write operations to each cell. To overcome these challenges in STT-RAM caches, this paper explores the potential of eliminating redundant writes using the phenomenon of frequent value locality (FVL). According to FLV, few distinct values appear in a large fraction of memory transactions, with emphasis on cache memories in this work. By leveraging frequent value locality, we propose a novel value-based hybrid (STT-RAM +, SRAM) cache that has benefits of both SRAM and STT-RAM technologies — i.e., it is high-performance, power-efficient, and scalable. Our evaluation results for a 8-core chip-multiprocessor with 6MB last-level cache show that our proposed design is able to reduce power consumption of a STT-RAM cache by up to 90% (an average of 82%), enhances its lifetime by up to 52% (29% on average), and improves the system performance by up 30% (11% on average), for a wide range of multi-threaded and multi-program workloads. Mohammad Arjomand, Amin Jadidi, Mahmut T. Kandemir, Chita R. Das |
ICCAD | 4 |
| 2017 | A Scale-Out Enterprise Storage ArchitectureabstractA robust enterprise SSD design should provide scalable throughput and storage capacity by integrating (up to thousands) flash chips in a scale-out fashion. However, the current "channel-based" SSD architecture is not a scalable design choice to allow such a dense integration. Motivated by the inherent architectural scalability of PCIe, we propose UT-SSD, a novel enterprise-scale scale-out SSD architecture, which enables the connection of a large number of (1000s) flash chips using the native PCIe buses instead of the conventional channels. We also propose an architectural enhancement that further improves the performance of our base UT-SSD by maximizing flash utilization. Our experimental analysis of UT-SSD with workloads drawn from various domains shows that the throughput of UT-SSD can reach up to 110 GB/s by successfully aggregating the bandwidth of 4096 flash chips. In addition, our proposed enhancement over this base UT-SSD increases the flash utilization by 50.7%, which in turn results in 116% additional throughput improvement. Wonil Choi, Myoungsoo Jung, Mahmut T. Kandemir, Chita R. Das |
ICCD | 4 |
| 2017 | Phoenix: A Constraint-Aware Scheduler for Heterogeneous DatacentersabstractToday's datacenters are increasingly becoming diverse with respect to both hardware and software architectures in order to support a myriad of applications. These applications are also heterogeneous in terms of job response times and resource requirements (eg., Number of Cores, GPUs, Network Speed) and they are expressed as task constraints. Constraints are used for ensuring task performance guarantees/Quality of Service(QoS) by enabling the application to express its specific resource requirements. While several schedulers have recently been proposed that aim to improve overall application and system performance, few of these schedulers consider resource constraints across tasks while making the scheduling decisions. Furthermore, latencycritical workloads and short-lived jobs that typically constitute about 90% of the total jobs in a datacenter have strict QoS requirements, which can be ensured by minimizing the tail latency through effective scheduling. In this paper, we propose Phoenix, a constraint-aware hybrid scheduler to address both these problems (constraint awareness and ensuring low tail latency) by minimizing the job response times at constrained workers. We use a novel Constraint Resource Vector (CRV) based scheduling, which in turn facilitates reordering of the jobs in a queue to minimize tail latency. We have used the publicly available Google traces to analyze their constraint characteristics and have embedded these constraints in Cloudera and Yahoo cluster traces for studying the impact of traces on system performance. Experiments with Google, Cloudera and Yahoo cluster traces across 15,000 worker node cluster shows that Phoenix improves the 99th percentile job response times on an average by 1.9× across all three traces when compared against a state-of-the-art hybrid scheduler. Further, in comparison to other distributed scheduler like Hawk, it improves the 90thand 99thpercentile job response times by 4.5× and 5× respectively. Prashanth Thinakaran, Jashwant Raj Gunasekaran, Bikash Sharma, Mahmut T. Kandemir, Chita R. Das |
ICDCS | 5 |
| 2017 | Quantifying the Potential Benefits of On-chip Near-Data Computing in Manycore ProcessorsabstractIncreasing data set sizes motivate for a shift of focus from computation-centric systems to data-centric systems, where data movement is treated as a first-class optimization metric. An example of this emerging paradigm is in-situ computing in largescale computing systems. Observing that data movement costs are increasing at an exponential rate even at a node level (as a node itself is fast-becoming a large manycore system), this paper provides a limit study of near-data computing within a manycore chip. Specifically, it makes the following two contributions. First, it quantifies the potential performance benefits of three incarnations of the near-data computing paradigm under the assumption of zero on-chip network latency and an infinite number of extra cores for offloading computations close to data they require. Our detailed experimental evaluation indicates that the most successful of these incarnations can boost the performance of the original execution by as much as 75%. The second contribution of this paper is an investigation of more realistic schemes that can approximate the potential savings achieved by perfect near-data computing. Our results demonstrate performance improvements ranging between 44% and 52%, over the original execution. We also discuss the pros and cons of each of these realistic schemes, and point to further research directions. Jagadish Kotra, Diana R. Guttman, Nachiappan Chidambaram Nachiappan, Mahmut T. Kandemir, Chita R. Das |
MASCOTS | 5 |
| 2017 | DEMM: A Dynamic Energy-Saving Mechanism for Multicore MemoriesabstractSince main memory system contributes to a large and increasing fraction of server/datacenter energy consumption, there have been several efforts to reduce its power and energy consumption. DVFS schemes have been used to reduce the memory power, but they come with a performance penalty. In this work, we propose DEMM, an OS-based, high performance DVFS mechanism that reduces memory power by dynamically scaling individual memory channel frequencies/voltages. Our strategy also involves clustering the running applications based on their sensitivities to memory latency, and assigning memory channels to the application clusters. We introduce a new metric called Discrete Misses per Kilo Cycle (DMPKC) to capture the performance sensitivities of the applications to memory frequency modulation. DEMM allows us to save power in the memory system with negligible impact on performance. We demonstrate around 25% savings in the memory system energy and 10% savings in the total system energy, with only a 4% loss in workload performance. Akbar Sharifi, Wei Ding 0008, Diana R. Guttman, Hui Zhao 0013, Xulong Tang, Mahmut T. Kandemir, Chita R. Das |
MASCOTS | 7 |
| 2017 | Race-to-sleep + content caching + display caching: a recipe for energy-efficient video streaming on handheldsabstractVideo streaming has become the most common application in handhelds and this trend is expected to grow in future to account for about 75% of all mobile data traffic by 2021. Thus, optimizing the performance and energy consumption of video processing in mobile devices is critical for sustaining the handheld market growth. In this paper, we propose three complementary techniques, race-to-sleep, content caching and display caching, to minimize the energy consumption of the video processing flows. Unlike the state-of-the-art frame-by-frame processing of a video decoder, the first scheme, race-to-sleep, uses two approaches, called batching of frames and frequency boosting to prolong its sleep state for saving energy, while avoiding any frame drops. The second scheme, content caching, exploits the content similarity of smaller video blocks, called macroblocks, to design a novel cache organization for reducing the memory pressure. The third scheme, in turn, takes advantage of content similarity at the display controller to facilitate display caching further improving energy efficiency. We integrate these three schemes for developing an end-to-end video processing framework and evaluate our design on a comprehensive mobile system design platform with a variety of video processing workloads. Our evaluations show that the proposed three techniques complement each other in improving performance by avoiding frame drops and reducing the energy consumption of video streaming applications by 21%, on average, compared to the current baseline design. Haibo Zhang 0005, Prasanna Venkatesh Rengasamy, Shulin Zhao 0001, Nachiappan Chidambaram Nachiappan, Anand Sivasubramaniam, Mahmut T. Kandemir, Ravi R. Iyer 0001, Chita R. Das |
MICRO | 8 |
| 2017 | HL-PCM: MLC PCM Main Memory with Accelerated ReadabstractMulti-Level Cell Phase Change Memory (MLC PCM) is a promising candidate technology for DRAM replacement in main memory of modern computers. Despite of its high density and low power advantages, this technology seriously suffers from slow read and write operations. While prior works extensively studied the problem of slow write, this paper targets high read latency problem in MLC PCM and introduces an architecture mechanism to overcome it. To this end, we rely on the fact that reading different bits from an MLC cell takes different latencies, i.e., for a 2-bit MLC, reading its Most-Significant Bit (MSB) is fast, while reading its Least-Significant Bits (LSBs) is slower. We then propose Half-Line PCM (HL-PCM), a novel memory architecture that leverages this non-uniformity in reading MLC PCM's content to send a requested memory block to the processor in different cycles-it sends half of a memory block to the processor ahead of the other half. If the processor requested a word belonging to the first half, it can resume its execution on receiving the first half, while the other half has not sent yet and scheduled to be received by the memory controller later. HL-PCM is easy and simple to implement, i.e., it needs minor modifications at memory controller, the search/evict policies at last level cache, as well as data layout in main memory. Our experimental results show that the proposed design improves the average memory access latency by 33-43 percent and program's execution time by 23 percent, on average, while incurring negligible overhead at memory controller and PCM DIMM, in a 16-core chip multiprocessor (CMP) running memory-intensive benchmarks. Mohammad Arjomand, Amin Jadidi, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2016 | μC-States: Fine-grained GPU Datapath Power ManagementabstractTo improve the performance of Graphics Processing Units (GPUs) beyond simply increasing core count, architects are recently adopting a scale-up approach: the peak throughput and individual capabilities of the GPU cores are increasing rapidly. This big-core trend in GPUs leads to various challenges, including higher static power consumption and lower and imbalanced utilization of the datapath components of a big core. As we show in this paper, two key problems ensue: (1) the lower and imbalanced datapath utilization can waste power as an application does not always utilize all portions of the big core datapath, and (2) the use of big cores can lead to application performance degradation in some cases due to the higher memory system contention caused by the more memory requests generated by each big core. Onur Kayiran, Adwait Jog, Ashutosh Pattnaik, Rachata Ausavarungnirun, Xulong Tang, Mahmut T. Kandemir, Gabriel H. Loh, Onur Mutlu, Chita R. Das |
PACT | 9 |
| 2016 | Scheduling Techniques for GPU Architectures with Processing-In-Memory CapabilitiesabstractProcessing data in or near memory (PIM), as opposed to in conventional computational units in a processor, can greatly alleviate the performance and energy penalties of data transfers from/to main memory. Graphics Processing Unit (GPU) architectures and applications, where main memory bandwidth is a critical bottleneck, can benefit from the use of PIM. To this end, an application should be properly partitioned and scheduled to execute on either the main, powerful GPU cores that are far away from memory or the auxiliary, simple GPU cores that are close to memory (e.g., in the logic layer of 3D-stacked DRAM). Ashutosh Pattnaik, Xulong Tang, Adwait Jog, Onur Kayiran, Asit K. Mishra, Mahmut T. Kandemir, Onur Mutlu, Chita R. Das |
PACT | 8 |
| 2016 | Re-NUCA: A Practical NUCA Architecture for ReRAM Based Last-Level CachesabstractAlthough resistive RAM (ReRAM) technology offers a good combination of high capacity and low-power for cache memories, its long write latency and low endurance are potential showstoppers to its wide commercial adoption. In particular, its low write-endurance can cause fast wear-out of cache lines, bringing reliability issues and leading to capacity reduction over time. This problem is exacerbated when ReRAM cache has dynamic NUCA structure, where each core brings most of its data to the cache banks close to itself and writes become localized. We propose Re-NUCA, a NUCA architecture design for ReRAM cache to address its lifetime problem while keeping its performance high. Re-NUCA relies on performance-wise data criticality: if it realizes a cache line is performance critical, it keeps it in the banks close to the target core, like dynamic NUCA, otherwise, it maps cache lines onto banks using static NUCA to evenly distribute writes over cache banks. This change in mapping of cache lines to banks relaxes the lifetime problem in ReRAM NUCA significantly and wear-levels the lifetime of banks. Re-NUCA needs a logic for detecting performance-wise critical cache lines and a low-overhead changes in TLB for keeping mapping information. Our experimental results of a 16-core chip multiprocessor with 32MB ReRAM L3 cache show that Re-NUCA improves the lifetime of the non-volatile cache by about 42%, on average, with almost no impact on performance. Jagadish Kotra, Mohammad Arjomand, Diana R. Guttman, Mahmut T. Kandemir, Chita R. Das |
IPDPS | 5 |
| 2016 | Boosting Access Parallelism to PCM-Based Main MemoryabstractDespite its promise as a DRAM main memory replacement, Phase Change Memory (PCM) has high write latencies which can be a serious detriment to its widespread adoption. Apart from slowing down a write request, the consequent high latency can also keep other chips of the same rank, that are not involved in this write, idle for long times. There are several practical considerations that make it difficult to allow subsequent reads and/or writes to be served concurrently from the same chips during the long latency write. This paper proposes and evaluates several novel mechanisms - re-constructing data from error correction bits instead of waiting for chips currently busy to serve a read, rotating word mappings across chips of a PCM rank, and rotating the mapping of error detection/correction bits across these chips - to overlap several reads with an ongoing write (RoW) and even a write with an ongoing write (WoW). The paper also presents the necessary micro-architectural enhancements needed to implement these mechanisms, without significantly changing the current interfaces. The resulting PCM access parallelism (PCMap) system incorporating these enhancements, boosts the intra-rank-level parallelism during such writes from a very low baseline value of 2.4 to an average and maximum values of 4.5 and 7.4, respectively (out of a maximum of 8.0), across a wide spectrum of both multiprogrammed and multithreaded workloads. This boost in parallelism results in an average IPC improvement of 15.6% and 16.7% for the multi-programmed and multi-threaded workloads, respectively. Mohammad Arjomand, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das |
ISCA | 4 |
| 2016 | MLC PCM main memory with accelerated readabstractThis paper alleviates the problem of slow reads in the Multi-Level Cell Phase Change Memory (MLC PCM) by exploiting a the fact that the Most-Significant Bit (MSB) of MLCs is read fast, while reading the Least-Significant Bits (LSBs) is slower. We propose Half-Line PCM (HL-PCM), a memory architecture that leverages this property to send half of a cache line to the processor ahead of the other half, so that processor continues its execution if the missed data element is in the first half. Our evaluation shows that HL-PCM improves program execution time by 23%, on average, in a 16-core CMP model for workloads from PARSEC-2 benchmark. Mohammad Arjomand, Amin Jadidi, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das |
ISPASS | 5 |
| 2016 | Storage consolidation: Not always a panacea, but can we ease the pain?abstractThe system administrator, faces an arduous task of figuring out whether to consolidate storage workloads, which workloads to isolate, which workloads to co-locate, and how to reduce the interference on co-located workloads? This paper presents an approach to ease this arduous task. We consider different mixes of enterprise storage applications on a high-end SSD to study their performance, system throughput and fairness in their consolidated execution. The paper also considers a static approach using an integrated combination of resource partitioning and data placement to reduce the interference between the workloads. Narges Shahidi, Mohammad Arjomand, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das |
ISPASS | 5 |
| 2016 | OSCAR: Orchestrating STT-RAM cache traffic for heterogeneous CPU-GPU architecturesabstractAs we integrate data-parallel GPUs with general-purpose CPUs on a single chip, the enormous cache traffic generated by GPUs will not only exhaust the limited cache capacity, but also severely interfere with CPU requests. Such heterogeneous multicores pose significant challenges to the design of shared last-level cache (LLC). This problem can be mitigated by replacing SRAM LLC with emerging non-volatile memories like Spin-Transfer Torque RAM (STT-RAM), which provides larger cache capacity and near-zero leakage power. However, without careful design, the slow write operations of STT-RAM may offset the capacity benefit, and the system may still suffer from contention in the shared LLC and on-chip interconnects. While there are cache optimization techniques to alleviate such problems, we reveal that the true potential of STT-RAM LLC may still be limited because now that the cache hit rate has been improved by the increased capacity, the on-chip network can become a performance bottleneck. CPU and GPU packets contend with each other for the shared network bandwidth. Moreover, the mixed-criticality read/write packets to STT-RAM add another layer of complexity to the network resource allocation. Therefore, being aware of the disparate latency tolerance of CPU/GPU applications and the asymmetric read/write latency of STT-RAM, we propose OSCAR to Orchestrate STT-RAM Caches traffic for heterogeneous ARchitectures. Specifically, an integration of asynchronous batch scheduling and priority based allocation for on-chip interconnect is proposed to maximize the potential of STT-RAM based LLC. Simulation results on a 28-GPU and 14-CPU system demonstrate an average of 17.4% performance improvement for CPUs, 10.8% performance improvement for GPUs, and 28.9% LLC energy saving compared to SRAM based LLC design. Jia Zhan, Onur Kayiran, Gabriel H. Loh, Chita R. Das, Yuan Xie 0001 |
MICRO | 4 |
| 2016 | Exploring the potentials of parallel garbage collection in SSDs for enterprise storage systemsabstractIn the last decade, NAND flash-based SSDs have been widely adopted for high-end enterprise systems in an attempt to provide a high-performance and reliable storage. However, inferior performance is frequently attained mainly due to the need for Garbage Collection (GC). GC in flash memory is the process of identifying and clearing the blocks of unneeded data to create space for the new data to be allocated. GC is a high-latency operation and once it is scheduled for service to a block of a plane in a flash chip (each flash chip consists of multiple planes), it can increase latency for later arriving I/O requests to the same plane. Apart from that, the consequent high latency also keep other planes of the same chip, that are not involved in this GC, idle for a long time. We show that for the baseline SSD with modern FTL, GC considerably reduces the plane-level parallelism, causing significant performance degradation. There are several circuit-level constraints that make it difficult to allow subsequent I/O operations and/or GCs to be served concurrently from the same chip, but different planes, during the long latency GC. This paper proposes a novel GC strategy, called Parallel GC (PaGC), whose goal is to proactively run GC on the remaining planes of a flash chip whenever any of its planes needs to execute on-demand GC. The resulting PaGC system boosts the response time of I/O requests by up to 45% (32% on average) for different GC settings and across a wide spectrum of enterprise I/O workloads. Narges Shahidi, Mohammad Arjomand, Myoungsoo Jung, Mahmut T. Kandemir, Chita R. Das, Anand Sivasubramaniam |
SC | 5 |
| 2016 | Exploiting Core Criticality for Enhanced GPU PerformanceabstractModern memory access schedulers employed in GPUs typically optimize for memory throughput. They implicitly assume that all requests from different cores are equally important. However, we show that during the execution of a subset of CUDA applications, different cores can have different amounts of tolerance to latency. In particular, cores with a larger fraction of warps waiting for data to come back from DRAM are less likely to tolerate the latency of an outstanding memory request. Requests from such cores are more critical than requests from others. Based on this observation, this paper introduces a new memory scheduler, called (C)ritica(L)ity (A)ware (M)emory (S)cheduler (CLAMS), which takes into account the latency-tolerance of the cores that generate memory requests. The key idea is to use the fraction of critical requests in the memory request buffer to switch between scheduling policies optimized for criticality and locality. If this fraction is below a threshold, CLAMS prioritizes critical requests to ensure cores that cannot tolerate latency are serviced faster. Otherwise, CLAMS optimizes for locality, anticipating that there are too many critical requests and prioritizing one over another would not significantly benefit performance. Adwait Jog, Onur Kayiran, Ashutosh Pattnaik, Mahmut T. Kandemir, Onur Mutlu, Ravi R. Iyer 0001, Chita R. Das |
SIGMETRICS | 7 |
| 2015 | Exploiting Inter-Warp Heterogeneity to Improve GPGPU PerformanceabstractIn a GPU, all threads within a warp execute the same instruction in lockstep. For a memory instruction, this can lead to memory divergence: the memory requests for some threads are serviced early, while the remaining requests incur long latencies. This divergence stalls the warp, as it cannot execute the next instruction until all requests from the current instruction complete. In this work, we make three new observations. First, GPGPU warps exhibit heterogeneous memory divergence behavior at the shared cache: some warps have most of their requests hit in the cache (high cache utility), while other warps see most of their request miss (low cache utility). Second, a warp retains the same divergence behavior for long periods of execution. Third, due to high memory level parallelism, requests going to the shared cache can incur queuing delays as large as hundreds of cycles, exacerbating the effects of memory divergence. We propose a set of techniques, collectively called Memory Divergence Correction (MeDiC), that reduce the negative performance impact of memory divergence and cache queuing. MeDiC uses warp divergence characterization to guide three components: (1) a cache bypassing mechanism that exploits the latency tolerance of low cache utility warps to both alleviate queuing delay and increase the hit rate for high cache utility warps, (2) a cache insertion policy that prevents data from highcache utility warps from being prematurely evicted, and (3) a memory controller that prioritizes the few requests received from high cache utility warps to minimize stall time. We compare MeDiC to four cache management techniques, and find that it delivers an average speedup of 21.8%, and 20.1% higher energy efficiency, over a state-of-the-art GPU cache management mechanism across 15 different GPGPU applications. Rachata Ausavarungnirun, Saugata Ghose, Onur Kayiran, Gabriel H. Loh, Chita R. Das, Mahmut T. Kandemir, Onur Mutlu |
PACT | 5 |
| 2015 | Exploiting Staleness for Approximating Loads on CMPsabstractCoherence misses are an important factor in limiting the scalability of multi-threaded shared memory applications on chip multiprocessors (CMPs) that are envisaged to contain dozens of cores in the imminent future. This paper proposes a novel approach to tackling this problem by leveraging the growingly important paradigm of approximate computing. Many applications are either tolerant to slight errors in the output or if stringent, have in-built resiliency to tolerate some errors in the execution. The approximate computing paradigm suggests breaking conventional barriers of mandating stringent correctness on the hardware, allowing more flexibility in the performance-power-reliability design space. Taking the multi-threaded applications in the SPLASH-2 benchmark suite, we note that nearly all these applications have such inherent resiliency and/or tolerance to slight errors in the output. Based on this observation, we propose to approximate coherence-related load misses by returning stale values, i.e., the version at the time of the invalidation. We show that returning such values from the invalidated lines already present in d-L1 offers only limited scope for improvement since those lines get evicted fairly soon due to the high pressure on d-L1. Instead, we propose a very small (8 lines) Stale Victim Cache (SVC), to hold such lines upon d-L1 eviction. While this does offer significant improvement, there is the possibility of data getting very stale in such a structure, making it highly sensitive to the choice of what data to keep, and for how long. To address these concerns, we propose to time-out these lines from the SVC to limit their staleness in a mechanism called SVC+TB. We show that SVC+TB provides as much as 28.6% speedup in some SPLASH-2 applications, with an average speedup between 10-15% across the entire suite, becoming comparable to an ideal execution that does not incur coherence misses. Further, the consequent approximations have little impact on the correctness, allowing all of them to complete. There were no errors, because of inherent application resilience, in eleven applications, and the maximum error was at most 0.08% across the entire suite. Prasanna Venkatesh Rengasamy, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das |
PACT | 4 |
| 2015 | Storage Consolidation on SSDs: Not Always a Panacea, but Can We Ease the Pain?abstractStorage Consolidation is increasing being adopted to reduce system costs, simplify the storage infrastructure, and enhance availability and resource management. However, consolidation leads to interference in shared resources. This poster shows the effect of consolidation on performance of co-located applications and proposes a static approach to reduce interference. Narges Shahidi, Anand Sivasubramanian, Mahmut T. Kandemir, Chita R. Das |
PACT | 4 |
| 2015 | Domain knowledge based energy management in handheldsabstractEnergy management in handheld devices is becoming a daunting task with the growing number of accelerators, increasing memory demands and high computing capacities required to support applications with stringent QoS needs. Current DVFS techniques that modulate power states of a single hardware component, or even recent proposals that manage multiple components, can lose out opportunities for attaining high energy efficiencies that may be possible by leveraging application domain knowledge. Thus, this paper proposes a coordinated multi-component energy optimization mechanism for handheld devices, where the energy profile of different components such as CPU, memory, GPU and IP cores are considered in unison to trigger the appropriate DVFS state by exploiting the application domain knowledge. Specifically, we show that for the important class of frame-based applications, the domain knowledge - frame processing rates, component utilization and available slack - can be used to decide effective DVFS states for each component from among the numerous choices. With such knowledge, rather than a brute force search of all speed setting choices, we propose two simpler heuristics, called Greedy policy and Kaldor-Hicks compensation policy, to make the decisions at frame boundaries. Our evaluations with 7 commonly-used Android apps show that our domain-aware coordinated DVFS policies have 23% better energy efficiency than the conventionally used Android governors, and are within ~9% of an optimal policy that does not drop any frames. Nachiappan Chidambaram Nachiappan, Praveen Yedlapalli, Niranjan Soundararajan, Anand Sivasubramaniam, Mahmut T. Kandemir, Ravi R. Iyer 0001, Chita R. Das |
HPCA | 7 |
| 2015 | VIP: virtualizing IP chains on handheld platformsabstractEnergy-efficient user-interactive and display-oriented applications on handhelds rely heavily on multiple accelerators (termed IP cores) to meet their periodic frame processing needs. Further, these platforms are starting to host multiple applications concurrently on the multiple CPU cores. Unfortunately, today's hardware exposes an interface that forces the host software (Android drivers) to treat each IP core as an isolated device. Consequently, the host CPU has to get involved in the (i) processing of each frame, (ii) scheduling them to ensure timely progress through the IP cores to meet their QoS needs, and (iii) explicitly having to move data from one IP core to the next, with main memory serving as the common staging area. Nachiappan Chidambaram Nachiappan, Haibo Zhang 0005, Jihyun Ryoo, Niranjan Soundararajan, Anand Sivasubramaniam, Mahmut T. Kandemir, Ravi R. Iyer 0001, Chita R. Das |
ISCA | 8 |
| 2015 | A case for core-assisted bottleneck acceleration in GPUs: enabling flexible data compression with assist warpsabstractModern Graphics Processing Units (GPUs) are well provisioned to support the concurrent execution of thousands of threads. Unfortunately, different bottlenecks during execution and heterogeneous application requirements create imbalances in utilization of resources in the cores. For example, when a GPU is bottlenecked by the available off-chip memory bandwidth, its computational resources are often overwhelmingly idle, waiting for data from memory to arrive. Nandita Vijaykumar, Gennady Pekhimenko, Adwait Jog, Abhishek Bhowmick 0002, Rachata Ausavarungnirun, Chita R. Das, Mahmut T. Kandemir, Todd C. Mowry, Onur Mutlu |
ISCA | 6 |
| 2014 | Trading cache hit rate for memory performanceabstractMost of the prior compiler based data locality optimization works target exclusively cache locality optimization, and row-buffer locality in DRAM banks received much less attention. In particular, to the best of our knowledge, there is no single compiler based approach that can improve row-buffer locality in executing irregular applications. This presents a critical problem considering the fact that executing irregular applications in a power and performance efficient manner will be a key requirement to extract maximum benefits from emerging multicore machines and exascale systems. Motivated by these observations, this paper makes the following contributions. First, it presents a compiler-runtime cooperative data layout optimization approach that takes as input an irregular program that has already been optimized for cache locality and generates an output code with the same cache performance but better row-buffer locality (lower number of row-buffer misses). Second, it discusses a more aggressive strategy that sacrifices some cache performance in order to further improve row-buffer performance (i.e., it trades cache performance for memory system performance). The ultimate goal of this strategy is to find the right tradeoff point between cache performance and row-buffer performance so that the overall application performance is improved. Third, the paper performs a detailed evaluation of these two approaches using both an AMD Opteron based multicore system and a multicore simulator. The experimental results, collected using five real-world irregular applications, show that (i) conventional cache optimizations do not improve row-buffer locality significantly; (ii) our first approach achieves about 9.8% execution time improvement by keeping the number of cache misses the same as a cache-optimized code but reducing the number of row-buffer misses; and (iii) our second approach achieves even higher execution time improvements (13.8% on average) by sacrificing cache performance for additional memory performance. Wei Ding 0008, Mahmut T. Kandemir, Diana R. Guttman, Adwait Jog, Chita R. Das, Praveen Yedlapalli |
PACT | 5 |
| 2014 | Managing GPU Concurrency in Heterogeneous ArchitecturesabstractHeterogeneous architectures consisting of general-purpose CPUs and throughput-optimized GPUs are projected to be the dominant computing platforms for many classes of applications. The design of such systems is more complex than that of homogeneous architectures because maximizing resource utilization while minimizing shared resource interference between CPU and GPU applications is difficult. We show that GPU applications tend to monopolize the shared hardware resources, such as memory and network, because of their high thread-level parallelism (TLP), and discuss the limitations of existing GPU-based concurrency management techniques when employed in heterogeneous systems. To solve this problem, we propose an integrated concurrency management strategy that modulates the TLP in GPUs to control the performance of both CPU and GPU applications. This mechanism considers both GPU core state and system-wide memory and network congestion information to dynamically decide on the level of GPU concurrency to maximize system performance. We propose and evaluate two schemes: one (CM-CPU) for boosting CPU performance in the presence of GPU interference, the other (CM-BAL) for improving both CPU and GPU performance in a balanced manner and thus overall system performance. Our evaluations show that the first scheme improves average CPU performance by 24%, while reducing average GPU performance by 11%. The second scheme provides 7% average performance improvement for both CPU and GPU applications. We also show that our solution allows the user to control performance trade-offs between CPUs and GPUs. Onur Kayiran, Nachiappan Chidambaram Nachiappan, Adwait Jog, Rachata Ausavarungnirun, Mahmut T. Kandemir, Gabriel H. Loh, Onur Mutlu, Chita R. Das |
MICRO | 8 |
| 2014 | Short-Circuiting Memory Traffic in Handheld PlatformsabstractHandheld devices are ubiquitous in today's world. With their advent, we also see a tremendous increase in device-user interactivity and real-time data processing needs. Media (audio/video/camera) and gaming use-cases are gaining substantial user attention and are defining product successes. The combination of increasing demand from these use-cases and having to run them at low power (from a battery) means that architects have to carefully study the applications and optimize the hardware and software stack together to gain significant optimizations. In this work, we study workloads from these domains and identify the memory subsystem (system agent) to be a critical bottleneck to performance scaling. We characterize the lifetime of the "frame-based" data used in these workloads through the system and show that, by communicating at frame granularity, we miss significant performance optimization opportunities, caused by large IP-to-IP data reuse distances. By carefully breaking these frames into sub-frames, while maintaining correctness, we demonstrate substantial gains with limited hardware requirements. Specifically, we evaluate two techniques, flow-buffering and IP-IP short-circuiting, and show that these techniques bring both power-performance benefits and enhanced user experience. Praveen Yedlapalli, Nachiappan Chidambaram Nachiappan, Niranjan Soundararajan, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das |
MICRO | 6 |
| 2014 | GemDroid: a framework to evaluate mobile platformsabstractAs the demand for feature-rich mobile systems such as smartphones and tablets has outpaced other computing systems and is expected to continue at a faster rate, it is projected that SoCs with tens of cores and hundreds of IPs (or accelerator) will be designed to provide unprecedented level of features and functionality in future. Design of such mobile systems with required QoS and power budgets along with other design constraints will be a daunting task for computer architects since any ad hoc, piece-meal solution is unlikely to result in an optimal design. This requires early exploration of the complete design space to understand the system-level design trade-offs. To the best of our knowledge, there is no such publicly available tool to conduct a holistic evaluation of mobile platforms consisting of cores, IPs and system software. Nachiappan Chidambaram Nachiappan, Praveen Yedlapalli, Niranjan Soundararajan, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das |
SIGMETRICS | 6 |
| 2013 | Neither more nor less: Optimizing thread-level parallelism for GPGPUsabstractGeneral-purpose graphics processing units (GPG-PUs) are at their best in accelerating computation by exploiting abundant thread-level parallelism (TLP) offered by many classes of HPC applications. To facilitate such high TLP, emerging programming models like CUDA and OpenCL allow programmers to create work abstractions in terms of smaller work units, called cooperative thread arrays (CTAs). CTAs are groups of threads and can be executed in any order, thereby providing ample opportunities for TLP. The state-of-the-art GPGPU schedulers allocate maximum possible CTAs per-core (limited by available on-chip resources) to enhance performance by exploiting TLP. However, we demonstrate in this paper that executing the maximum possible number of CTAs on a core is not always the optimal choice from the performance perspective. High number of concurrently executing threads might cause more memory requests to be issued, and create contention in the caches, network and memory, leading to long stalls at the cores. To reduce resource contention, we propose a dynamic CTA scheduling mechanism, called DYNCTA, which modulates the TLP by allocating optimal number of CTAs, based on application characteristics. To minimize resource contention, DYNCTA allocates fewer CTAs for applications suffering from high contention in the memory subsystem, compared to applications demonstrating high throughput. Simulation results on a 30-core GPGPU platform with 31 applications show that the proposed CTA scheduler provides 28% average improvement in performance compared to the existing CTA scheduler. Onur Kayiran, Adwait Jog, Mahmut T. Kandemir, Chita R. Das |
PACT | 4 |
| 2013 | Meeting midway: Improving CMP performance with memory-side prefetchingabstractBoth on-chip resource contention and off-chip latencies have a significant impact on memory requests in large-scale chip multiprocessors. We propose a memory-side prefetcher, which brings data on-chip from DRAM, but does not proactively further push this data to the cores/caches. Sitting close to memory, it avails close knowledge of DRAM state and memory channels to leverage DRAM row buffer locality and channel state to bring data (from the current row buffer) on-chip ahead of need. This not only reduces the number of off-chip accesses for demand requests, but also reduces row buffer conflicts, effectively improving DRAM access times. At the same time, our prefetcher maintains this data in a small buffer at each memory controller instead of pushing it into the caches to avoid on-chip resource contention. We show that the proposed memory-side prefetcher outperforms a state-of-the-art core-side prefetcher and an existing memory-side prefetcher. More importantly, our prefetcher can also work in tandem with the core-side prefetcher to amplify the benefits. Using a wide range of multiprogrammed and multi-threaded workloads, we show that this memory-side prefetcher provides IPC improvements of 6.2% (maximum of 33.6%), and 10% (maximum of 49.6%), on an average when running alone and when combined with a core-side prefetcher, respectively. By meeting requests midway, our solution reduces the off-chip latencies while avoiding the on-chip resource contention caused by inaccurate and ill-timed prefetches. Praveen Yedlapalli, Jagadish Kotra, Emre Kultursay, Mahmut T. Kandemir, Chita R. Das, Anand Sivasubramaniam |
PACT | 5 |
| 2013 | OWL: cooperative thread array aware scheduling techniques for improving GPGPU performanceabstractEmerging GPGPU architectures, along with programming models like CUDA and OpenCL, offer a cost-effective platform for many applications by providing high thread level parallelism at lower energy budgets. Unfortunately, for many general-purpose applications, available hardware resources of a GPGPU are not efficiently utilized, leading to lost opportunity in improving performance. A major cause of this is the inefficiency of current warp scheduling policies in tolerating long memory latencies. Adwait Jog, Onur Kayiran, Nachiappan Chidambaram Nachiappan, Asit K. Mishra, Mahmut T. Kandemir, Onur Mutlu, Ravi R. Iyer 0001, Chita R. Das |
ASPLOS | 8 |
| 2013 | A heterogeneous multiple network-on-chip design: an application-aware approachabstractCurrent network-on-chip designs in chip-multiprocessors are agnostic to application requirements and hence are provisioned for the general case, leading to wasted energy and performance. We observe that applications can generally be classified as either network bandwidth-sensitive or latency-sensitive. We propose the use of two separate networks on chip, where one network is optimized for bandwidth and the other for latency, and the steering of applications to the appropriate network. We further observe that not all bandwidth (latency) sensitive applications are equally sensitive to network bandwidth (latency). Hence, within each network, we prioritize packets based on the relative sensitivity of the applications they belong to. We introduce two metrics, network episode height and length, as proxies to estimate bandwidth and latency sensitivity, to classify and rank applications. Our evaluations show that the resulting heterogeneous two-network design can provide significant energy savings and performance improvements across a variety of workloads compared to a single one-size-fits-all single network and homogeneous multiple networks. Asit K. Mishra, Onur Mutlu, Chita R. Das |
DAC | 3 |
| 2013 | CloudPD: Problem determination and diagnosis in shared dynamic cloudsabstractIn this work, we address problem determination in virtualized clouds. We show that high dynamism, resource sharing, frequent reconfiguration, high propensity to faults and automated management introduce significant new challenges towards fault diagnosis in clouds. Towards this, we propose CloudPD, a fault management framework for clouds. CloudPD leverages (i) a canonical representation of the operating environment to quantify the impact of sharing; (ii) an online learning process to tackle dynamism; (iii) a correlation-based performance models for higher detection accuracy; and (iv) an integrated end-to-end feedback loop to synergize with a cloud management ecosystem. Using a prototype implementation with cloud representative batch and transactional workloads like Hadoop, Olio and RUBiS, it is shown that CloudPD detects and diagnoses faults with low false positives (<; 16%) and high accuracy of 88%, 83% and 83%, respectively. In an enterprise trace-based case study, CloudPD diagnosed anomalies within 30 seconds and with an accuracy of 77%, demonstrating its effectiveness in real-life operations. Bikash Sharma, Praveen Jayachandran, Akshat Verma, Chita R. Das |
DSN | 4 |
| 2013 | HybridMR: A Hierarchical MapReduce Scheduler for Hybrid Data CentersabstractVirtualized environments are attractive because they simplify cluster management, while facilitating cost-effective workload consolidation. As a result, virtual machines in public clouds or private data centers, have become the norm for running transactional applications like web services and virtual desktops. On the other hand, batch workloads like MapReduce, are typically deployed in a native cluster to avoid the performance overheads of virtualization. While both these virtual and native environments have their own strengths and weaknesses, we demonstrate in this work that it is feasible to provide the best of these two computing paradigms in a hybrid platform. In this paper, we make a case for a hybrid data center consisting of native and virtual environments, and propose a 2-phase hierarchical scheduler, called HybridMR, for the effective resource management of interactive and batch workloads. In the first phase, HybridMR classifies incoming MapReduce jobs based on the expected virtualization overheads, and uses this information to automatically guide placement between physical and virtual machines. In the second phase, HybridMR manages the run-time performance of MapReduce jobs collocated with interactive applications in order to provide best effort delivery to batch jobs, while complying with the Service Level Agreements (SLAs) of interactive applications. By consolidating batch jobs with over-provisioned foreground applications, the available unused resources are better utilized, resulting in improved application performance and energy efficiency. Evaluations on a hybrid cluster consisting of 24 physical servers and 48 virtual machines, with diverse workload mix of interactive and batch MapReduce applications, demonstrate that HybridMR can achieve up to 40% improvement in the completion times of MapReduce jobs, over the virtual-only case, while complying with the SLAs of interactive applications. Compared to the native-only cluster, at the cost of minimal performance penalty, HybridMR boosts resource utilization by 45%, and achieves up to 43% energy savings. These results indicate that a hybrid data center with an efficient scheduling mechanism can provide a cost-effective solution for hosting both batch and interactive workloads. Bikash Sharma, Timothy Wood 0001, Chita R. Das |
ICDCS | 3 |
| 2013 | Orchestrated scheduling and prefetching for GPGPUsabstractIn this paper, we present techniques that coordinate the thread scheduling and prefetching decisions in a General Purpose Graphics Processing Unit (GPGPU) architecture to better tolerate long memory latencies. We demonstrate that existing warp scheduling policies in GPGPU architectures are unable to effectively incorporate data prefetching. The main reason is that they schedule consecutive warps, which are likely to access nearby cache blocks and thus prefetch accurately for one another, back-to-back in consecutive cycles. This either 1) causes prefetches to be generated by a warp too close to the time their corresponding addresses are actually demanded by another warp, or 2) requires sophisticated prefetcher designs to correctly predict the addresses required by a future "far-ahead" warp while executing the current warp. Adwait Jog, Onur Kayiran, Asit K. Mishra, Mahmut T. Kandemir, Onur Mutlu, Ravi R. Iyer 0001, Chita R. Das |
ISCA | 7 |
| 2013 | Editorial to special section on networks on chip: Architecture, tools, and methodologiesabstractNo abstract available. Diana Marculescu, Chita R. Das |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2012 | MROrchestrator: A Fine-Grained Resource Orchestration Framework for MapReduce ClustersabstractEfficient resource management in data centers and clouds running large distributed data processing frameworks like MapReduce is crucial for enhancing the performance of hosted applications and increasing resource utilization. However, existing resource scheduling schemes in Hadoop MapReduce allocate resources at the granularity of fixed-size, static portions of nodes, called slots. In this work, we show that MapReduce jobs have widely varying demands for multiple resources, making the static and fixed-size slot-level resource allocation a poor choice both from the performance and resource utilization standpoints. Furthermore, lack of coordination in the management of multiple resources across nodes prevents dynamic slot reconfiguration, and leads to resource contention. Motivated by this, we propose MROrchestrator, a MapReduce resource Orchestrator framework, which can dynamically identify resource bottlenecks, and resolve them through fine-grained, coordinated, and on-demand resource allocations. We have implemented MROrchestrator on two 24-node native and virtualized Hadoop clusters. Experimental results with a suite of representative MapReduce benchmarks demonstrate up to 38% reduction in job completion times, and up to 25% increase in resource utilization. We further demonstrate the performance boost in existing resource managers like NGM and Mesos, when augmented with MROrchestrator. Bikash Sharma, Ramya Prabhakar, Seung-Hwan Lim, Mahmut T. Kandemir, Chita R. Das |
IEEE CLOUD | 5 |
| 2012 | Application-aware prefetch prioritization in on-chip networksabstractData prefetching is an effective technique for hiding memory latency. When issued prefetches are inaccurate, performance can degrade. Prior research provided solutions to deal with inaccurate prefetches at the cache and memory levels, but not in the interconnect of a large-scale multiprocessor system. This work introduces application-aware prefetch prioritization techniques to mitigate the negative effects of prefetching in a network-on-chip (NoC) based multicore system. The idea is to rank prefetches from different applications based on their potential utility for the application and propensity to cause interference to other applications. Our evaluation shows that this approach provides significant performance improvements over a baseline that does not distinguish between prefetches from different applications. Nachiappan Chidambaram Nachiappan, Asit K. Mishra, Mahmut T. Kandemir, Anand Sivasubramaniam, Onur Mutlu, Chita R. Das |
PACT | 6 |
| 2012 | PEPON: performance-aware hierarchical power budgeting for NoC based multicoresabstractTargeting NoC based multicores, we propose a two-level power budget distribution mechanism, called PEPON, where the first level distributes the overall power budget of the multicore system among various types of on-chip resources like the cores, caches, and NoC, and the second level determines the allocation of power to individual instances of each type of resource. Both these distributions are oriented towards maximizing workload performance without exceeding the specified power budget. Extensive experimental evaluations of the proposed power distribution scheme using a full system simulation and detailed power models emphasize the importance of power budget partitioning at both levels. Specifically, our results show that the proposed scheme can provide up to 29% performance improvement as compared to no power budgeting, and performs 13% better than a competing scheme, under the same chip-wide power cap. Akbar Sharifi, Asit K. Mishra, Shekhar Srikantaiah, Mahmut T. Kandemir, Chita R. Das |
PACT | 5 |
| 2012 | Cache revive: architecting volatile STT-RAM caches for enhanced performance in CMPsabstractHigh density, low leakage and non-volatility are the attractive features of Spin-Transfer-Torque-RAM (STT-RAM), which has made it a strong competitor against SRAM as a universal memory replacement in multi-core systems. However, STT-RAM suffers from high write latency and energy which has impeded its widespread adoption. To this end, we look at trading-off STT-RAM's non-volatility property (data-retention-time) to overcome these problems. We formulate the relationship between retention-time and write-latency, and find optimal retention-time for architecting an efficient cache hierarchy using STT-RAM. Our results show that, compared to SRAM-based design, our proposal can improve performance and energy consumption by 18% and 60%, respectively. Adwait Jog, Asit K. Mishra, Cong Xu 0002, Yuan Xie 0001, Narayanan Vijaykrishnan, Ravi R. Iyer 0001, Chita R. Das |
DAC | 7 |
| 2012 | Addressing End-to-End Memory Access Latency in NoC-Based MulticoresabstractTo achieve high performance in emerging multicores, it is crucial to reduce the number of memory accesses that suffer from very high latencies. However, this should be done with care as improving latency of an access can worsen the latency of another as a result of resource sharing. Therefore, the goal should be to balance latencies of memory accesses issued by an application in an execution phase, while ensuring a low average latency value. Targeting Network-on-Chip (NoC) based multicores, we propose two network prioritization schemes that can cooperatively improve performance by reducing end-to-end memory access latencies. Our first scheme prioritizes memory response messages such that, in a given period of time, messages of an application that experience higher latencies than the average message latency for that application are expedited and a more uniform memory latency pattern is achieved. Our second scheme prioritizes the request messages that are destined for idle memory banks over others, with the goal of improving bank utilization and preventing long queues from being built in front of the memory banks. These two network prioritization-based optimizations together lead to uniform memory access latencies with a low average value. Our experiments with a 4×8 mesh network-based multicore show that, when applied together, our schemes can achieve 15%, 10% and 13% performance improvement on memory intensive, memory non-intensive, and mixed multiprogrammed workloads, respectively. Akbar Sharifi, Emre Kultursay, Mahmut T. Kandemir, Chita R. Das |
MICRO | 4 |
| 2012 | D-factor: a quantitative model of application slow-down in multi-resource shared systemsabstractScheduling multiple jobs onto a platform enhances system utilization by sharing resources. The benefits from higher resource utilization include reduced cost to construct, operate, and maintain a system, which often include energy consumption. Maximizing these benefits, while satisfying performance limits, comes at a price -- resource contention among jobs increases job completion time. In this paper, we analyze slow-downs of jobs due to contention for multiple resources in a system; referred to as dilation factor. We observe that multiple-resource contention creates non-linear dilation factors of jobs. From this observation, we establish a general quantitative model for dilation factors of jobs in multi-resource systems. A job is characterized by a vector-valued loading statistics and dilation factors of a job set are given by a quadratic function of their loading vectors. We demonstrate how to systematically characterize a job, maintain the data structure to calculate the dilation factor (loading matrix), and calculate the dilation factor of each job. We validated the accuracy of the model with multiple processes running on a native Linux server, virtualized servers, and with multiple MapReduce workloads co-scheduled in a cluster. Evaluation with measured data shows that the D-factor model has an error margin of less than 16%. We also show that the model can be integrated with an existing on-line scheduler to minimize the makespan of workloads. Seung-Hwan Lim, Jae-Seok Huh, Youngjae Kim 0001, Galen M. Shipman, Chita R. Das |
SIGMETRICS | 5 |
| 2012 | Cache invalidation strategies for Internet-based vehicular ad hoc networks
Sunho Lim, Chansu Yu, Chita R. Das |
Comput. Commun. | 3 |
| 2011 | Modeling and synthesizing task placement constraints in Google compute clustersabstractEvaluating the performance of large compute clusters requires benchmarks with representative workloads. At Google, performance benchmarks are used to obtain performance metrics such as task scheduling delays and machine resource utilizations to assess changes in application codes, machine configurations, and scheduling algorithms. Existing approaches to workload characterization for high performance computing and grids focus on task resource requirements for CPU, memory, disk, I/O, network, etc. Such resource requirements address how much resource is consumed by a task. However, in addition to resource requirements, Google workloads commonly include task placement constraints that determine which machine resources are consumed by tasks. Task placement constraints arise because of task dependencies such as those related to hardware architecture and kernel version. Bikash Sharma, Victor Chudnovsky, Joseph L. Hellerstein, Rasekh Rifaat, Chita R. Das |
SoCC | 5 |
| 2011 | ACCESS: Smart scheduling for asymmetric cache CMPsabstractIn current Chip-multiprocessors (CMPs), a significant portion of the die is consumed by the last-level cache. Until recently, the balance of cache and core space has been primarily guided by the needs of single applications. However, as multiple applications or virtual machines (VMs) are consolidated on such a platform, researchers have observed that not all VMs or applications require significant amount of cache space. In order to take advantage of this phenomenon, we explore the use of asymmetric last-level caches in a CMP platform. While asymmetric cache CMPs provide the benefit of reduced power and area, it is important to build in hardware/software support to appropriately schedule applications on to cores with suitable cache capacity. In this paper, we address this problem with our ACCESS architecture comprising of: (a) asymmetric caches across a group of cores, (b) hardware support that enables prediction of cache performance on the different sized caches and (c) OS scheduler support to make use of the prediction capability and appropriately schedule applications on to core with suitable cache capacity. Measurements on a working prototype using SPEC2006 benchmarks show that our ACCESS architecture can effectively schedule jobs in an asymmetric cache CMP and provide 23% performance improvement compared to a naive scheduler, and is 97% close to an oracle scheduler in making schedules. Xiaowei Jiang, Asit K. Mishra, Li Zhao 0002, Ravi R. Iyer 0001, Zhen Fang 0002, Sadagopan Srinivasan, Srihari Makineni, Paul Brett, Chita R. Das |
HPCA | 9 |
| 2011 | Architecting on-chip interconnects for stacked 3D STT-RAM caches in CMPsabstractEmerging memory technologies such as STT-RAM, PCRAM, and resistive RAM are being explored as potential replacements to existing on-chip caches or main memories for future multi-core architectures. This is due to the many attractive features these memory technologies posses: high density, low leakage, and non-volatility. However, the latency and energy overhead associated with the write operations of these emerging memories has become a major obstacle in their adoption. Previous works have proposed various circuit and architectural level solutions to mitigate the write overhead. In this paper, we study the integration of STT-RAM in a 3D multi-core environment and propose solutions at the on-chip network level to circumvent the write overhead problem in the cache architecture with STT-RAM technology. Our scheme is based on the observation that instead of staggering requests to a write-busy STT-RAM bank, the network should schedule requests to other idle cache banks for effectively hiding the latency. Thus, we prioritize cache accesses to the idle banks by delaying accesses to the STT-RAM cache banks that are currently serving long latency write requests. Through a detailed characterization of the cache access patterns of 42 applications, we propose an efficient mechanism to facilitate such delayed writes to cache banks by (a) accurately estimating the busy time of each cache bank through logical partitioning of the cache layer and (b) prioritizing packets in a router requesting accesses to idle banks. Evaluations on a 3D architecture, consisting of 64 cores and 64 STT-RAM cache banks, show that our proposed approach provides 14% average IPC improvement for multi-threaded benchmarks, 19% instruction throughput benefits for multi-programmed workloads, and 6% latency reduction compared to a recently proposed write buffering mechanism. Asit K. Mishra, Xiangyu Dong 0001, Guangyu Sun 0003, Yuan Xie 0001, Narayanan Vijaykrishnan, Chita R. Das |
ISCA | 6 |
| 2011 | A case for heterogeneous on-chip interconnects for CMPsabstractNetwork-on-chip (NoC) has become a critical shared resource in the emerging Chip Multiprocessor (CMP) era. Most prior NoC designs have used the same type of router across the entire network. While this homogeneous network design eases the burden on a network designer, partitioning the resources equally among all routers across the network does not lead to optimal resource usage, and hence, affects the performance-power envelope. In this work, we propose to apportion the resources in an NoC to leverage the non-uniformity in network resource demand. Our proposal includes partitioning the network resources, specifically buffers and links, in an optimal manner. This approach results in redistributing resources such that routers that require more resources are allocated more buffers and wider links compared to routers demanding fewer resources. This results in a novel heterogeneous network, called HeteroNoC, which is composed of two types of routers -- small power efficient routers, and big high performance routers. We evaluate a number of heterogeneous network configurations, composed of big and small routers, and show that giving more resources to routers along the diagonals in a mesh network provides maximum benefits in terms of performance and power. We also show the potential benefits of the HeteroNoC design by co-evaluating it with memory-controllers and configuring it with an asymmetric CMP consisting of heterogeneous cores. Asit K. Mishra, Narayanan Vijaykrishnan, Chita R. Das |
ISCA | 3 |
| 2011 | A dynamic energy management scheme for multi-tier data centersabstractMulti-tier data centers have become a norm for hosting modern Internet applications because they provide a flexible, modular, scalable and high performance environment. However, these benefits come at a price of the economic dent incurred in powering and cooling these large hosting centers. Thus, energy efficiency has become a critical consideration in designing Internet data centers. In this paper, we propose a multifaceted approach, Hybrid, consisting of dynamic provisioning, frequency scaling and dynamic power management (DPM) schemes to reduce the energy consumption of multi-tier data centers, while meeting the Service Level Agreements (SLAs). We formulate a mathematical model of the energy and performance/SLA optimization problem followed by a queueing theory based approach to develop two heuristics for solving the optimization problem. The first heuristic dynamically provisions the optimal number of servers required in each tier. The second heuristic proactively decides the CPU speed and the duration of sleep states of a server to achieve further energy savings. We evaluate our heuristics using a simulator that was validated with real measurements on a prototype three-tier data center consisting of 25 servers with two multi-tier application benchmarks. Our experimental results indicate that the proposed scheme, Hybrid, can reduce the energy consumption by 50% relative to static provisioning without CPU frequency scaling and DPM. We demonstrate that Hybrid satisfies the SLAs for dynamically varying workloads. In addition, the proposed multifaceted approach is more energy efficient than the other methods such as dynamic provisioning with exploiting deep sleep states. Seung-Hwan Lim, Bikash Sharma, Byung-Chul Tak, Chita R. Das |
ISPASS | 4 |
| 2011 | METE: meeting end-to-end QoS in multicores through system-wide resource managementabstractManagement of shared resources in emerging multicores for achieving predictable performance has received considerable attention in recent times. In general, almost all these approaches attempt to guarantee a certain level of performance QoS (weighted IPC, harmonic speedup, etc) by managing a single shared resource or at most a couple of interacting resources. A fundamental shortcoming of these approaches is the lack of coordination between these shared resources to satisfy a system level QoS. This is undesirable because providing end-to-end QoS in future multicores is essential for supporting wide-spread adoption of these architectures in virtualized servers and cloud computing systems. An initial step towards such an end-to-end QoS support in multicores is to ensure that at least the major computational and memory resources on-chip are managed efficiently in a coordinated fashion. Akbar Sharifi, Shekhar Srikantaiah, Asit K. Mishra, Mahmut T. Kandemir, Chita R. Das |
SIGMETRICS | 5 |
| 2011 | RAFT: A router architecture with frequency tuning for on-chip networks
Asit K. Mishra, Aditya Yanamandra, Reetuparna Das, Soumya Eachempati, Ravi R. Iyer 0001, Narayanan Vijaykrishnan, Chita R. Das |
J. Parallel Distributed Comput. | 7 |
| 2010 | Cost-driven 3D integration with interconnect layersabstractThe ever increasing die area of Chip Multiprocessors (CMPs) affects manufacturing yield, resulting in higher manufacture cost. Meanwhile, network-on-chip (NoC) has emerged as a promising and scalable solution for interconnecting the cores in CMPs, however it consumes significant portion of the total die area. In this paper, we propose to decouple the interconnect fabric from computing and storage layers, forming a separate layer called Interconnect Service Layer (ISL), in the context of three-dimensional (3D) chip integration. Such decoupling helps reduce the die area for each layer in 3D stacking. ISL itself can integrate multiple super-imposed interconnect topologies. More importantly, ISL can be designed, manufactured, and tested as a separate Intellectual Property (IP) component, which supports multiple designs in the computing and storage layers. The resulting methodology also helps support different manufacturing volume in each die of 3D to reduce the overall manufacturing cost. We demonstrate the proposed methodology with an ISL design example and compare to its 2D and 3D counterparts without ISL support. The results show that 3D design with ISL not only provides significant cost reduction, but also achieves power-performance improvement thanks to the efficient usage of ISL. Xiaoxia Wu, Guangyu Sun 0003, Xiangyu Dong 0001, Reetuparna Das, Yuan Xie 0001, Chita R. Das, Jian Li 0059 |
DAC | 6 |
| 2010 | Performance Analysis of Communications & Radar Coexistence in a Covert UWB OSA SystemabstractFar-field target detection, multi-sensor communications, and security are essential requirements of first-emergency networks. A radar-communications system is a potential opportunistic spectrum access (OSA) solution, harnessing the coexisting advantages of radio detection and ranging (RADAR), and wireless communications. A multi-functional waveform has been designed, by embedding an Orthogonal Frequency Division Multiplexing (OFDM) signal within a spectrally notched ultra-wideband (UWB) random noise waveform. Extending that development, this paper analyzes the waveform's Bit-Error-Rate (BER) and Ambiguity Function (AF) formulations to demonstrate its OSA ability, that offers reliable multi-user communications, and high range and Doppler resolution in target detection. We further conclude that up to 30% of the available UWB bandwidth can be simultaneously utilized for concealed data communications without adversely affecting radar performance or its physical layer covertness. Shrawan Chittoor Surender, Ram M. Narayanan, Chita R. Das |
GLOBECOM | 3 |
| 2010 | Aérgia: exploiting packet latency slack in on-chip networksabstractTraditional Network-on-Chips (NoCs) employ simple arbitration strategies, such as round-robin or oldest-first, to decide which packets should be prioritized in the network. This is counter-intuitive since different packets can have very different effects on system performance due to, e.g., different level of memory-level parallelism (MLP) of applications. Certain packets may be performance-critical because they cause the processor to stall, whereas others may be delayed for a number of cycles with no effect on application-level performance as their latencies are hidden by other outstanding packets'latencies. In this paper, we define slack as a key measure that characterizes the relative importance of a packet. Specifically, the slack of a packet is the number of cycles the packet can be delayed in the network with no effect on execution time. This paper proposes new router prioritization policies that exploit the available slack of interfering packets in order to accelerate performance-critical packets and thus improve overall system performance. When two packets interfere with each other in a router, the packet with the lower slack value is prioritized. We describe mechanisms to estimate slack, prevent starvation, and combine slack-based prioritization with other recently proposed application-aware prioritization mechanisms. Reetuparna Das, Onur Mutlu, Thomas Moscibroda, Chita R. Das |
ISCA | 4 |
| 2010 | CPM in CMPs: Coordinated Power Management in Chip-MultiprocessorsabstractMultiple clock domain architectures have recently been proposed to alleviate the power problem in CMPs by having different frequency/voltage values assigned to each domain based on workload requirements. However, accurate allocation of power to these voltage/frequency islands based on time varying workload characteristics as well as controlling the power consumption at the provisioned power level is quite non-trivial. Toward this end, we propose a two-tier feedback-based control theoretic solution. Our first-tier consists of a global power manager that allocates power targets to individual islands based on the workload dynamics. The power consumptions of these islands are in turn controlled by a second-tier, consisting of local controllers that regulate island power using dynamic voltage and frequency scaling in response to workload requirements. Asit K. Mishra, Shekhar Srikantaiah, Mahmut T. Kandemir, Chita R. Das |
SC | 4 |
| 2010 | Coordinated power management of voltage islands in CMPsabstractMultiple clock domain architectures have recently been proposed to alleviate the power problem in CMPs by having different frequency/voltage values assigned to each domain based on workload requirements. However, accurate allocation of power to these voltage/frequency islands based on time varying workload characteristics as well as controlling the power consumption at the provisioned power level is non-trivial. Toward this end, we propose a two-tier feedback-based control theoretic solution. Our first-tier consists of a global power manager that allocates power targets to individual islands based on the workload dynamics. The power consumptions of these islands are in turn controlled by a second-tier, consisting of local controllers that regulate island power using dynamic voltage and frequency scaling in response to workload requirements. Asit K. Mishra, Shekhar Srikantaiah, Mahmut T. Kandemir, Chita R. Das |
SIGMETRICS | 4 |
| 2010 | Integration of admission, congestion, and peak power control in QoS-aware clusters
Ki Hwan Yum, Yuho Jin, Eun Jung Kim 0001, Chita R. Das |
J. Parallel Distributed Comput. | 4 |
| 2010 | A Superscalar software architecture model for Multi-Core Processors (MCPs)
Gyu Sang Choi, Chita R. Das |
J. Syst. Softw. | 2 |
| 2010 | On the Effects of Process Variation in Network-on-Chip ArchitecturesabstractThe advent of diminutive technology feature sizes has led to escalating transistor densities. Burgeoning transistor counts are casting a dark shadow on modern chip design: global interconnect delays are dominating gate delays and affecting overall system performance. Networks-on-Chip (NoC) are viewed as a viable solution to this problem because of their scalability and optimized electrical properties. However, on-chip routers are susceptible to another artifact of deep submicron technology, Process Variation (PV). PV is a consequence of manufacturing imperfections, which may lead to degraded performance and even erroneous behavior. In this work, we present the first comprehensive evaluation of NoC susceptibility to PV effects, and we propose an array of architectural improvements in the form of a new router design-called SturdiSwitch-to increase resiliency to these effects. Through extensive reengineering of critical components, SturdiSwitch provides increased immunity to PV while improving performance and increasing area and power efficiency. Chrysostomos Nicopoulos, Suresh Srinivasan, Aditya Yanamandra, Dongkook Park, Narayanan Vijaykrishnan, Chita R. Das, Mary Jane Irwin |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2010 | Cooperative Caching in Wireless P2P Networks: Design, Implementation, and EvaluationabstractSome recent studies have shown that cooperative cache can improve the system performance in wireless P2P networks such as ad hoc networks and mesh networks. However, all these studies are at a very high level, leaving many design and implementation issues unanswered. In this paper, we present our design and implementation of cooperative cache in wireless P2P networks, and propose solutions to find the best place to cache the data. We propose a novel asymmetric cooperative cache approach, where the data requests are transmitted to the cache layer on every node, but the data replies are only transmitted to the cache layer at the intermediate nodes that need to cache the data. This solution not only reduces the overhead of copying data between the user space and the kernel space, it also allows data pipelines to reduce the end-to-end delay. We also study the effects of different MAC layers, such as 802.11-based ad hoc networks and multi-interface-multichannel-based mesh networks, on the performance of cooperative cache. Our results show that the asymmetric approach outperforms the symmetric approach in traditional 802.11-based ad hoc networks by removing most of the processing overhead. In mesh networks, the asymmetric approach can significantly reduce the data access delay compared to the symmetric approach due to data pipelines. Jing Zhao 0001, Guohong Cao, Chita R. Das |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2009 | MDCSim: A multi-tier data center simulation, platformabstractPerformance and power issues are becoming increasingly important in the design of large, cluster-based multitier data centers for supporting a multitude of services. The design and analysis of such large/complex distributed systems often suffer from the lack of availability of an adequate physical infrastructure. This paper presents a comprehensive, flexible, and scalable simulation platform for in-depth analysis of multi-tier data centers. Designed as a pluggable three-level architecture, our simulator captures all the important design specifics of the underlying communication paradigm, kernel level scheduling artifacts, and the application level interactions among the tiers of a three-tier data center. The flexibility of the simulator is attributed to its ability in experimenting with different design alternatives in the three layers, and in analyzing both the performance and power consumption with realistic workloads. The scalability of the simulator is demonstrated with analyses of different data center configurations. In addition, we have designed a prototype three-tier data center on an Infiniband Architecture (IBA) connected Linux cluster to validate the simulator. Using RUBiS benchmark workload, it is shown that the simulator is quite accurate in estimating the throughput, response time, and power consumption. We then demonstrate the applicability of the simulator in conducting three different types of studies. First, we conduct a comparative analysis of the IBA and 10 Gigabit Ethernet (10GigE) under different traffic conditions and with varying size clusters for understanding their relative merits in designing cluster-based servers. Second, measurement and characterization of power consumption across the servers of a three-tier data center is done. Third, we perform a configuration analysis of the Web server (WS), Application Server (AS), and Database Server (DB) for performance optimization. We believe that such a comprehensive simulation infrastructure is critical for providing guidelines in designing efficient and cost-effective multi-tier data centers. Seung-Hwan Lim, Bikash Sharma, Gunwoo Nam, Eun-Kyoung Kim, Chita R. Das |
CLUSTER | 5 |
| 2009 | Design and evaluation of a hierarchical on-chip interconnect for next-generation CMPsabstractPerformance and power consumption of an on-chip interconnect that forms the backbone of chip multiprocessors (CMPs), are directly influenced by the underlying network topology. Both these parameters can also be optimized by application induced communication locality since applications mapped on a large CMP system will benefit from clustered communication, where data is placed in cache banks closer to the cores accessing it. Thus, in this paper, we design a hierarchical network topology that takes advantage of such communication locality. The two-tier hierarchical topology consists of local networks that are connected via a global network. The local network is a simple, high-bandwidth, low-power shared bus fabric, and the global network is a low-radix mesh. The key insight that enables the hybrid topology is that most communication in CMP applications can be limited to the local network, and thus, using a fast, low-power bus to handle local communication will improve both packet latency and power-efficiency. The proposed hierarchical topology provides up to 63% reduction in energy-delay-product over mesh, 47% over flattened butterfly, and 33% with respect to concentrated mesh across network sizes with uniform and non-uniform synthetic traffic. For real parallel workloads, the hybrid topology provides up to 14% improvement in system performance (IPC) and in terms of energy-delay-product, improvements of 70%, 22%, 30% over the mesh, flattened butterfly, and concentrated mesh, respectively, for a 32-way CMP. Although the hybrid topology scales in a power- and bandwidth-efficient manner with network size, while keeping the average packet latency low in comparison to high radix topologies, it has lower throughput due to high concentration. To improve the throughput of the hybrid topology, we propose a novel router micro-architecture, called XShare, which exploits data value locality and bimodal traffic characteristics of CMP applications to transfer multiple small flits over a single channel. This helps in enhancing the network throughput by 35%, providing a latency reduction of 14% with synthetic traffic, and improving IPC on an average 4% with application workloads. Reetuparna Das, Soumya Eachempati, Asit K. Mishra, Narayanan Vijaykrishnan, Chita R. Das |
HPCA | 5 |
| 2009 | Path-Centric On-Demand Rate Adaptation for Mobile Ad Hoc NetworksabstractExploiting the multirate capability in mobile ad hoc networks (MANETs) is more complex than in single-hop WLANs because of the rate-distance and rate-hop count tradeoffs. This paper proposes path-centric on-demand rate adaptation for MANETs (PRAM) protocol. A unique feature that sets PRAM apart from most of previous studies is its path-centric approach. While others focus on finding the best data rate for each link and offering a routing path as a collection of links at their best rates, PRAM finds the best data rate for a source-destination pair and then, dynamically adapts it based on path lifetime and other factors. Another distinctive feature of PRAM is that it can be seamlessly incorporated with an on-demand routing protocol. Extensive performance study based on NS-2 has demonstrated that PRAM achieves as much as 71.7% higher packet delivery ratio than fixed-rate cases (6~54 Mbps) and as much as 43.2% higher than the multihop version of the well- known ARF mechanism in a wide range of network scenarios. It is also shown that PRAM is capable of using a mixture of data rates in an adaptive manner. Saehoon Kang, Chansu Yu, Chita R. Das, Guohong Cao |
ICCCN | 3 |
| 2009 | Cooperative Cache Invalidation Strategies for Internet-Based Vehicular Ad Hoc NetworksabstractInternet-based vehicular ad hoc network (IVANET) is an emerging technique that combines a wired Internet and a vehicular ad hoc network (VANET) for providing universal information and service accessibility. A key design optimization technique in IVANETs is to cache the frequently accessed data items in a local storage of vehicles. Since vehicles are not critically limited by the storage/memory space and power consumption, cache replacement scheme for accommodating new data items is not an issue. Rather, a more critical design question is how to keep the cached copies valid or to invalidate them when the original data items are updated. This is particularly important in IVANETs, where vehicles move very fast. This paper proposes state-aware cooperative cache invalidation (CCI) scheme and its enhancement (ECCI) that take advantage of the underlying location management mechanism. Extensive performance study shows that the proposed schemes can reduce the query delay as much as 69% and increase the cache hit rate up to 57% in comparison to two existing cache invalidation techniques, called poll-each-read (PER) and extended asynchronous (EAS). Note that PER and EAS have been modified to work in IVANETs. Sunho Lim, Chansu Yu, Chita R. Das |
ICCCN | 3 |
| 2009 | Mass Purging of Stale TCP Flows in Per-Flow Monitoring SystemsabstractTimely deletion of a large number of stale sessions monitored by Internet routers, particularly in the presence of SYN floods, is critical to prevent flow table explosion. We investigate two frameworks for purging of stale sessions: "opportunistic" purging that employs a free-list of pointers to memory and "deterministic purging" involving logical swapping of a 1-bit flow enable and touch-bit vectors without requiring a free list. We compare the performance of our algorithms with a state-of-the-art algorithm, namely finger-compressed filter (FCF). Our analysis using Internet traces shows that the deterministic purging, with no purging overhead, is ideal in that it reduces false positive and negative rates as compared to FCF by 52.5% and 59.2%, when the table size is twice the average number of active flows. Gunwoo Nam, Pushkar Patankar, George Kesidis, Chita R. Das, Cetin Seren |
ICCCN | 4 |
| 2009 | Clock-like Flow Replacement Schemes for Resilient Flow MonitoringabstractIn the context of a collaborating surveillance system for active TCP sessions handled by a networking device, we consider two problems. The first is the problem of protecting a flow table from overflow and the second is developing an efficient algorithm for estimating the number of active flows coupled with the identification of "heavy-hitter" TCP sessions. Our proposed techniques are sensitive to limited hardware and software resources allocated for this purpose in the linecards in addition to the very high data rates that modern line cards handle; specifically we are interested in cooperatively maintaining a per-flow state with a low cost, which has resiliency on dynamic traffic mix. We investigate a traditional timeout processing mechanism to manage the flow table for per-flow monitoring, called Timeout-Based Purging (TBP), our proposed Clock-like Flow Replacement (CFR) algorithms using a replacement policy, called "clock", and a hybrid approach combining these two. Experiments with Internet traces show that our CFR schemes can significantly reduce both false positive and false negative rates regardless of whether the flow table is fully occupied or sufficiently empty, even under SYN flooding. Our hybrid scheme estimates the number of active flows accurately, and confines the heavy-hitters without storing packet counters. Gunwoo Nam, Pushkar Patankar, Seung-Hwan Lim, Bikash Sharma, George Kesidis, Chita R. Das |
ICDCS | 6 |
| 2009 | On Interest Locality in Content-Based Routing for Large-scale MANETsabstractTo disseminate content with content-based routing (CBR), the routing paths of subscription and publication cannot be determined a priori and have to be computed hop-by-hop, which brings in scalability and robustness challenges in large scale mobile ad hoc networks (MANETs). In this paper, we propose a novel two-tier content-based routing protocol called CLONE (Community and Location aware content based routing). In CLONE, we map the human community structure of social networks to MANETs. The whole network can be self-organized into communities based on the interest locality, so that most subscriptions inside a community can be served in an intra-community fashion, reducing the communication overhead and the response delay. Community construction is self-organized and completely distributed. Analytical and simulation results demonstrate the effectiveness of CLONE in large-scale MANETs. Yang Zhang 0017, Jing Zhao 0001, Guohong Cao, Chita R. Das |
MASS | 4 |
| 2009 | Application-aware prioritization mechanisms for on-chip networksabstractNetwork-on-Chips (NoCs) are likely to become a critical shared resource in future many-core processors. The challenge is to develop policies and mechanisms that enable multiple applications to efficiently and fairly share the network, to improve system performance. Existing local packet scheduling policies in the routers fail to fully achieve this goal, because they treat every packet equally, regardless of which application issued the packet. Reetuparna Das, Onur Mutlu, Thomas Moscibroda, Chita R. Das |
MICRO | 4 |
| 2009 | A case for dynamic frequency tuning in on-chip networksabstractPerformance and power are the first order design metrics for Network-on-Chips (NoCs) that have become the de-facto standard in providing scalable communication backbones for multicores/CMPs. However, NoCs can be plagued by higher power consumption and degraded throughput if the network and router are not designed properly. Towards this end, this paper proposes a novel router architecture, where we tune the frequency of a router in response to network load to manage both performance and power. We propose three dynamic frequency tuning techniques, FreqBoost, FreqThrtl and FreqTune, targeted at congestion and power management in NoCs. As enablers for these techniques, we exploit Dynamic Voltage and Frequency Scaling (DVFS) and the imbalance in a generic router pipeline through time stealing. Experiments using synthetic workloads on a 8x8 wormhole-switched mesh interconnect show that FreqBoost is a better choice for reducing average latency (maximum 40%) while, FreqThrtl provides the maximum benefits in terms of power saving and energy delay product (EDP). The FreqTune scheme is a better candidate for optimizing both performance and power, achieving on an average 36% reduction in latency, 13% savings in power (up to 24% at high load), and 40% savings (up to 70% at high load) in EDP. With application benchmarks, we observe IPC improvement up to 23% using our design. The performance and power benefits also scale for larger NoCs. Asit K. Mishra, Reetuparna Das, Soumya Eachempati, Ravi R. Iyer 0001, Narayanan Vijaykrishnan, Chita R. Das |
MICRO | 6 |
| 2009 | A case for integrated processor-cache partitioning in chip multiprocessorsabstractExisting cache partitioning schemes are designed in a manner oblivious to the implicit processor partitioning enforced by the operating system. This paper examines an operating system directed integrated processor-cache partitioning scheme that partitions both the available processors and the shared cache in a chip multiprocessor among different multi-threaded applications. Extensive simulations using a set of multiprogrammed workloads show that our integrated processor-cache partitioning scheme facilitates achieving better performance isolation as compared to state of the art hardware/software based solutions. Specifically, our integrated processor-cache partitioning approach performs, on an average, 20.83% and 14.14% better than equal partitioning and the implicit partitioning enforced by the underlying operating system, respectively, on the fair speedup metric on an 8 core system. We also compare our approach to processor partitioning alone and a state-of-the-art cache partitioning scheme and our scheme fares 8.21% and 9.19% better than these schemes on a 16 core system. Shekhar Srikantaiah, Reetuparna Das, Asit K. Mishra, Chita R. Das, Mahmut T. Kandemir |
SC | 4 |
| 2009 | RandomCast: An Energy-Efficient Communication Scheme for Mobile Ad Hoc NetworksabstractIn mobile ad hoc networks (MANETs), every node overhears every data transmission occurring in its vicinity and thus, consumes energy unnecessarily. However, since some MANET routing protocols such as dynamic source routing (DSR) collect route information via overhearing, they would suffer if they are used in combination with 802.11 PSM. Allowing no overhearing may critically deteriorate the performance of the underlying routing protocol, while unconditional overhearing may offset the advantage of using PSM. This paper proposes a new communication mechanism, called RandomCast, via which a sender can specify the desired level of overhearing, making a prudent balance between energy and routing performance. In addition, it reduces redundant rebroadcasts for a broadcast packet, and thus, saves more energy. Extensive simulation using NS-2 shows that RandomCast is highly energy-efficient compared to conventional 802.11 as well as 802.11 PSM-based schemes, in terms of total energy consumption, energy goodput, and energy balance. Sunho Lim, Chansu Yu, Chita R. Das |
IEEE Trans. Mob. Comput. | 3 |
| 2008 | Performance and power optimization through data compression in Network-on-Chip architecturesabstractThe trend towards integrating multiple cores on the same die has accentuated the need for larger on-chip caches. Such large caches are constructed as a multitude of smaller cache banks interconnected through a packet-based network-on-chip (NoC) communication fabric. Thus, the NoC plays a critical role in optimizing the performance and power consumption of such non-uniform cache-based multicore architectures. While almost all prior NoC studies have focused on the design of router microarchitectures for achieving this goal, in this paper, we explore the role of data compression on NoC performance and energy behavior. In this context, we examine two different configurations that explore combinations of storage and communication compression: (1) Cache compression (CC) and (2) Compression in the NIC (NC). We also address techniques to hide the decompression latency by overlapping with NoC communication latency. Our simulation results with a diverse set of scientific and commercial benchmark traces reveal that CC can provide up to 33% reduction in network latency and up to 23% power savings. Even in the case of NC - where the data is compressed only when passing through the NoC fabric of the NUCA architecture and stored uncompressed - performance and power savings of up to 32% and 21%, respectively, can be obtained. These performance benefits in the interconnect translate up to 17% reduction in CPI. These benefits are orthogonal to any router architecture and make a strong case for utilizing compression for optimizing the performance and power envelope of NoC architectures. In addition, the study demonstrates the criticality of designing faster routers in shaping the performance behavior. Reetuparna Das, Asit K. Mishra, Chrysostomos Nicopoulos, Dongkook Park, Narayanan Vijaykrishnan, Ravi R. Iyer 0001, Mazin S. Yousif, Chita R. Das |
HPCA | 8 |
| 2008 | Exploring Anti-Spam Models in Large Scale VoIP SystemsabstractAlthough the problem of spam detection in email is well understood and has been extensively researched, a significant portion of emails today are spam. A most widely used method to detect spam involves content filtering, where the spam detector scans the received email for keywords. However, the same approach cannot be applied to detect Voice over IP (VoIP) spam, since a call has to be categorized as a legitimate or a spam (each to a degree with a certain reliability) before the connection is established. Also, spammers over IP can potentially generate orders of magnitude more spam volume, at far less cost, and with greater anonymity than telemarketers using the Public Switch Telephone Network (PSTN). The spam problem in VoIP is further compounded by the absence of a do-not-call-list, which has been the main reason for the reduction of spam calls in PSTN. Thus, the spam issue for VoIP is as important as those pertaining to quality-of-service (QoS) of the voice traffic itself. To this end, we propose two different anti-spam frameworks for large scale VoIP systems. The first one is a centralized SIP-based spam detection framework that relies on SIP messages during the call establishment phase to identify spam calls, and the second one is a distributed referral social network model, where a user is assigned a reputation score by its neighbors. Based on the reputation, a callee can decide either to accept or decline a call. Our simulation results indicate that the referral model can provide better anti-spam capabilities by isolating a spammer faster than the SIP based approach, and can also correctly identify spam calls over 98% of time. Pushkar Patankar, Gunwoo Nam, George Kesidis, Chita R. Das |
ICDCS | 4 |
| 2008 | MIRA: A Multi-layered On-Chip Interconnect Router ArchitectureabstractRecently, Network-on-Chip (NoC) architectures have gained popularity to address the interconnect delay problem for designing CMP / multi-core/SoC systems in deep sub-micron technology. However, almost all prior studies have focused on 2D NoC designs. Since three dimensional (3D) integration has emerged to mitigate the interconnect delay problem, exploring the NoC design space in 3D can provide ample opportunities to design high performance and energy-efficient NoC architectures. In this paper, we propose a 3D stacked NoC router architecture, called MIRA, which unlike the 3D routers in previous works, is stacked into multiple layers and optimized to reduce the overall area requirements and power consumption. We discuss the design details of a four-layer 3D NoC and its enhanced version with additional express channels, and compare them against a (6times6) 2D design and a baseline 3D design. All the designs are evaluated using a cycle-accurate 3D NoC simulator, and integrated with the Orion power model for performance and power analysis. The simulation results with synthetic and application traces demonstrate that the proposed multi-layered NoC routers can outperform the 2D and naive 3D designs in terms of performance and power. It can achieve up to 42% reduction in power consumption and up to 51% improvement in average latency with synthetic workloads. With real workloads, these benefits are around 67% and 38%, respectively. Dongkook Park, Soumya Eachempati, Reetuparna Das, Asit K. Mishra, Yuan Xie 0001, Narayanan Vijaykrishnan, Chita R. Das |
ISCA | 7 |
| 2008 | On cache invalidation for internet-based vehicular ad hoc networksabstractInternet-based vehicular ad hoc network (IVANET) is an emerging technique that combines a wired Internet and a vehicular ad hoc network (VANET) for developing a next generation of ubiquitous communication infrastructure and improving universal information and service accessibility. A key optimization technique in IVANETs is to cache frequently accessed data items in a local storage of vehicles. Since vehicles are not critically limited by the storage space, it is a less of a problem which data items to cache. Rather, a critical design issue is how to invalidate them when data items are updated. This is particularly a concern due to vehiclespsila high-speed mobility. In this paper, we propose a novel cache invalidation algorithm that takes advantage of the underlying location management scheme to reduce the number of broadcast operations and the corresponding query delay. Numerical results indicate that the proposed scheme significantly reduces the communication cost, and thus is proven to be a viable solution for IVANETs. Sunho Lim, Soo Hoan Chae, Chansu Yu, Chita R. Das |
MASS | 4 |
| 2008 | Coscheduled distributed-Web servers on system area network
Jin-Ha Kim, Gyu Sang Choi, Chita R. Das |
J. Parallel Distributed Comput. | 3 |
| 2008 | Proxy-RED: an AQM scheme for wireless local area networksabstractAbstract Wireless access points (APs) act as bridges between wired and wireless networks. Since the actually available bandwidth in wireless networks is much smaller than the bandwidth in wired networks, there is a disparity in channel capacity which makes the access point a significant network congestion point in the downstream direction. A current architectural trend in wireless local area networks (WLAN) is to move functionality from APs to a centralized gateway in order to reduce cost and improve features. In this paper, we study the use of RED, a well known active queue management (AQM) scheme, and explicit congestion notification (ECN) to handle bandwidth disparity between the wired and the wireless interface of an access point. Then, we propose the Proxy‐RED scheme, as a solution for reducing the AQM overhead from the access point. Simulations‐based performance analysis indicates that the proposed Proxy‐RED scheme improves the overall performance of a network. In particular, the Proxy‐RED scheme significantly reduces packet loss rate and improves goodput for a small buffer, and minimizes delay for a large buffer size. Copyright © 2006 John Wiley & Sons, Ltd. Sungwon Yi, Martin Kappes, Sachin Garg, Xidong Deng, George Kesidis, Chita R. Das |
Wirel. Commun. Mob. Comput. | 6 |
| 2007 | Characterizing Network Traffic in a Cluster-based, Multi-tier Data CenterabstractWith the increasing use of various Web-based services, design of high performance, scalable and dependable data centers has become a critical issue. Recent studies show that a clustered, multi-tier architecture is a cost-effective approach to design such servers. Since these servers are highly distributed and complex, understanding the workloads driving them is crucial for the success of the ongoing research to improve them. In view of this, there has been a significant amount of work to characterize the workloads of Web-based services. However, all of the previous studies focus on a high level view of these servers, and analyze request-based or session-based characteristics of the workloads. In this paper, we focus on the characteristics of the network behavior within a clustered, multi-tiered data center. Using a real implementation of a clustered three-tier data center, we analyze the arrival rate and inter-arrival time distribution of the requests to individual server nodes, the network traffic between tiers, and the average size of messages exchanged between tiers. The main results of this study are; (1) in most cases, the request inter-arrival rates follow log-normal distribution, and self-similarity exists when the data center is heavily loaded, (2) message sizes can be modeled by the log-normal distribution, and (3) service times fit reasonably well with the Pareto distribution and show heavy tailed behavior at heavy loads. Deniz Ersoz, Mazin S. Yousif, Chita R. Das |
ICDCS | 3 |
| 2007 | A novel dimensionally-decomposed router for on-chip communication in 3D architecturesabstractMuch like multi-storey buildings in densely packed metropolises, three-dimensional (3D) chip structures are envisioned as a viable solution to skyrocketing transistor densities and burgeoning die sizes in multi-core architectures. Partitioning a larger die into smaller segments and then stacking them in a 3D fashion can significantly reduce latency and energy consumption. Such benefits emanate from the notion that inter-wafer distances are negligible compared to intra-wafer distances. This attribute substantially reduces global wiring length in 3D chips. The work in this paper integrates the increasingly popular idea of packet-based Networks-on-Chip (NoC) into a 3D setting. While NoCs have been studied extensively in the 2D realm, the microarchitectural ramifications of moving into the third dimension have yet to be fully explored. This paper presents a detailed exploration of inter-strata communication architectures in 3D NoCs. Three design options are investigated; a simple bus-based inter-wafer connection, a hop-by-hop standard 3D design, and a full 3D crossbar implementation. In this context, we propose a novel partially-connected 3D crossbar structure, called the 3D Dimensionally-Decomposed (DimDe) Router, which provides a good tradeoff between circuit complexity and performance benefits. Simulation results using (a) a stand-alone cycle-accurate 3D NoC simulator running synthetic workloads, and (b) a hybrid 3D NoC/cache simulation environment running real commercial and scientific benchmarks, indicate that the proposed DimDe design provides latency and throughput improvements of over 20% on average over the other 3D architectures, while remaining within 5% of the full 3D crossbar performance. Furthermore, based on synthesized hardware implementations in 90 nm technology, the DimDe architecture outperforms all other designs -- including the full 3D crossbar -- by an average of 26% in terms of the Energy-Delay Product (EDP). Jongman Kim, Chrysostomos Nicopoulos, Dongkook Park, Reetuparna Das, Yuan Xie 0001, Narayanan Vijaykrishnan, Mazin S. Yousif, Chita R. Das |
ISCA | 8 |
| 2007 | Cache invalidation strategies for internet-based mobile ad hoc networks
Sunho Lim, Wang-Chien Lee, Guohong Cao, Chita R. Das |
Comput. Commun. | 4 |
| 2007 | An analytical model for interval caching in interactive video servers
Suneuy Kim, Chita R. Das |
J. Netw. Comput. Appl. | 2 |
| 2007 | A comprehensive performance and energy consumption analysis of scheduling alternatives in clusters
Gyu Sang Choi, Jin-Ha Kim, Deniz Ersoz, Andy B. Yoo, Chita R. Das |
J. Supercomput. | 5 |
| 2007 | An SSL Back-End Forwarding Scheme in Cluster-Based Web ServersabstractState-of-the-art cluster-based data centers consisting of three tiers (Web server, application server, and database server) are being used to host complex Web services such as e-commerce applications. The application server handles dynamic and sensitive Web contents that need protection from eavesdropping, tampering, and forgery. Although the secure sockets layer (SSL) is the most popular protocol to provide a secure channel between a client and a cluster-based network server, its high overhead degrades the server performance considerably and, thus, affects the server scalability. Therefore, improving the performance of SSL-enabled network servers is critical for designing scalable and high-performance data centers. In this paper, we examine the impact of SSL offering and SSL-session-aware distribution in cluster-based network servers. We propose a back-end forwarding scheme, called ssl_with_bf, that employs a low-overhead user-level communication mechanism like virtual interface architecture (VIA) to achieve a good load balance among server nodes. We compare three distribution models for network servers, round robin (RR), ssl_with_session, and ssl_with_bf, through simulation. The experimental results with 16-node and 32-node cluster configurations show that, although the session reuse of ssl_with_session is critical to improve the performance of application servers, the proposed back-end forwarding scheme can further enhance the performance due to better load balancing. The ssl_with_bf scheme can minimize the average latency by about 40 percent and improve throughput across a variety of workloads. Jin-Ha Kim, Gyu Sang Choi, Chita R. Das |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2007 | Exploring IBA Design Space for Improved PerformanceabstractInfiniBand architecture (IBA) is envisioned to be the default communication fabric for future system area networks (SANs) or clusters. However, IBA design is currently in its infancy since the released specification outlines only higher level functionalities, leaving it open for exploring various design alternatives. In this paper, we investigate four corelated techniques for providing high and predictable performance in IBA. These are: 1) using the shortest path first (SPF) algorithm for deterministic packet routing, 2) developing a multipath routing mechanism for minimizing congestion, 3) developing a selective packet dropping scheme to handle deadlock and congestion, and 4) providing multicasting support for customized applications. These designs are implemented in a pipelined, IBA-style switch architecture, and are evaluated using an integrated workload consisting of MPEG-2 video streams, best- effort traffic, and control traffic on a versatile IBA simulation testbed. Simulation results with 15-node and 30-node irregular networks indicate that the SPF routing, multipath routing, packet dropping, and multicasting schemes are quite effective in delivering high and assured performance in clusters Eun Jung Kim 0001, Ki Hwan Yum, Chita R. Das, Mazin S. Yousif, José Duato |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2006 | Exploring Fault-Tolerant Network-on-Chip ArchitecturesabstractThe advent of deep sub-micron technology has exacerbated reliability issues in on-chip interconnects. In particular, single event upsets, such as soft errors, and hard faults are rapidly becoming a force to be reckoned with. This spiraling trend highlights the importance of detailed analysis of these reliability hazards and the incorporation of comprehensive protection measures into all Network-on-Chip (NoC) designs. In this paper, we examine the impact of transient failures on the reliability of on-chip interconnects and develop comprehensive counter-measures to either prevent or recover from them. In this regard, we propose several novel schemes to remedy various kinds of soft error symptoms, while keeping area and power overhead at a minimum. Our proposed solutions are architected to fully exploit the available infrastructures in an NoC and enable versatile reuse of valuable resources. The effectiveness of the proposed techniques has been validated using a cycle-accurate simulator. Dongkook Park, Chrysostomos Nicopoulos, Jongman Kim, Narayanan Vijaykrishnan, Chita R. Das |
DSN | 5 |
| 2006 | A Gracefully Degrading and Energy-Efficient Modular Router Architecture for On-Chip NetworksabstractPacket-based on-chip networks are increasingly being adopted in complex system-on-chip (SoC) designs supporting numerous homogeneous and heterogeneous functional blocks. These network-on-chip (NoC) architectures are required to not only provide ultra-low latency, but also occupy a small footprint and consume as little energy as possible. Further, reliability is rapidly becoming a major challenge in deep sub-micron technologies due to the increased prominence of permanent faults resulting from accelerated aging effects and manufacturing/testing challenges. Towards the goal of designing low-latency, energy-efficient and reliable on-chip communication networks, we propose a novel fine-grained modular router architecture. The proposed architecture employs decoupled parallel arbiters and uses smaller crossbars for row and column connections to reduce output port contention probabilities as compared to existing designs. Furthermore, the router employs a new switch allocation technique known as "mirroring effect" to reduce arbitration depth and increase concurrency. In addition, the modular design permits graceful degradation of the network in the event of permanent faults and also helps to reduce the dynamic power consumption. Our simulation results indicate that in an 8 times 8 mesh network, the proposed architecture reduces packet latency by 4-40% and power consumption by 6-20% as compared to two existing router architectures. Evaluation using a combined performance, energy and fault-tolerance metric indicates that the proposed architecture provides 35-50% overall improvement compared to the two earlier routers Jongman Kim, Chrysostomos Nicopoulos, Dongkook Park, Narayanan Vijaykrishnan, Mazin S. Yousif, Chita R. Das |
ISCA | 6 |
| 2006 | Clustered Mobility Model for Scale-Free Wireless NetworksabstractRecently, researchers have discovered that many of social, natural and biological networks are characterized by scale-free power-law connectivity distribution and a few densely populated nodes, known as hubs. We envision that wireless communication or sensor networks are directly deployed over such real-world networks to facilitate communication among participating entities. Here nodes move in such a way that they exhibit scale-free connectivity distribution at any instance, which cannot be modeled by most of the prior mobility models such as random waypoint (RWP) mobility model. This paper proposes clustered mobility model (CMM), which facilitates in forming hubs in a network satisfying the scale-free property. We call this a scale-free wireless network (SFWN). In CMM, it is possible to control the degree of node concentration or non-homogeneity to easily assess the strengths and weaknesses of the scale-free phenomena. To the best of the authors' knowledge, there has been no such mobility model reported in the literature and we believe the proposed CMM can be usefully used to investigate the properties of the SFWNs that are likely to occur in a real deployment of wireless multihop and sensor networks. Another important feature of CMM is that it does not possess any unintended spatial and temporal characteristics found in other mobility models such as RWP. Finally, to highlight the difference between a SFWN and a conventional wireless network, extensive simulation study has been conducted to measure network capacities at the physical, link and network layers Sunho Lim, Chansu Yu, Chita R. Das |
LCN | 3 |
| 2006 | ViChaR: A Dynamic Virtual Channel Regulator for Network-on-Chip RoutersabstractThe advent of deep sub-micron technology has recently highlighted the criticality of the on-chip interconnects. As diminishing feature sizes have led to increases in global wiring delays, network-on-chip (NoC) architectures are viewed as a possible solution to the wiring challenge and have recently crystallized into a significant research thrust. Both NoC performance and energy budget depend heavily on the routers' buffer resources. This paper introduces a novel unified buffer structure, called the dynamic virtual channel regulator (ViChaR), which dynamically allocates virtual channels (VC) and buffer resources according to network traffic conditions. ViChaR maximizes throughput by dispensing a variable number of VCs on demand. Simulation results using a cycle-accurate simulator show a performance increase of 25% on average over an equal-size generic router buffer, or similar performance using a 50% smaller buffer. ViChaR's ability to provide similar performance with half the buffer size of a generic router is of paramount importance, since this can yield total area and power savings of 30% and 34%, respectively, based on synthesized designs in 90 nm technology Chrysostomos Nicopoulos, Dongkook Park, Jongman Kim, Narayanan Vijaykrishnan, Mazin S. Yousif, Chita R. Das |
MICRO | 6 |
| 2006 | A novel caching scheme for improving Internet-based mobile ad hoc networks performance
Sunho Lim, Wang-Chien Lee, Guohong Cao, Chita R. Das |
Ad Hoc Networks | 4 |
| 2005 | Design and analysis of an NoC architecture from performance, reliability and energy perspectiveabstractNetwork-on-Chip (NoC) architectures employing packet-based communication are being increasingly adopted in System-on-Chip (SoC) designs. In addition to providing high performance, the fault tolerance and reliability of these networks is becoming a critical issue due to several artifacts of deep sub-micron technologies. Consequently, it is important for a designer to have access to fast methods for evaluating the performance, reliability, and energy-efficiency of an on-chip network. Towards this end, first, we propose a novel path-sensitive router architecture for low-latency applications. Next, we present a queuing-theory-based model for evaluating the performance and energy behavior of on-chip networks. Then the model is used to demonstrate the effectiveness of our proposed router. The performance (average latency) and energy consumption results from the analytical model are validated with those obtained from a cycle-accurate simulator. Finally, we explore error detection and correction mechanisms that provide different energy-reliability- performance tradeoffs and extend our model to evaluate the on-chip network in the presence of these error protection schemes. Our reliability exploration culminates with the introduction of an array of transient fault protection techniques, both architectural and algorithmic, to tackle reliability issues within the router's individual hardware components. We propose a complete solution safeguarding against both the traditional link faults and internal router upsets, without incurring any significant latency, area and power overhead. Jongman Kim, Dongkook Park, Chrysostomos Nicopoulos, Narayanan Vijaykrishnan, Chita R. Das |
ANCS | 5 |
| 2005 | Exploiting NIC Memory for Improving Cluster-Based Webserver PerformanceabstractImproving the performance of Web servers has become a critical issue to handle the increasing demand on various network-based services. In this context, we exploit the local memory of programmable network interface cards (NICs) to improve the performance of cluster-based Web servers, which are increasingly used in designing Web server platforms. We use the NIC memory for caching recently accessed data blocks to improve server performance. We have implemented a prototype of the proposed NIC caching mechanism for a distributed Web server, based on PRESS (Carrera et al., 2002), on an 8-node, Myrinet-connected Linux cluster. Measurements with several server workloads show that NIC caching can enhance throughput by up to 27% compared to the original PRESS Web server without NIC caching, by minimizing the DMA and PCI bus overhead Gyu Sang Choi, Jin-Ha Kim, Deniz Ersoz, Mazin S. Yousif, Chita R. Das |
CLUSTER | 5 |
| 2005 | A Load Balancing Scheme for Cluster-based Secure Network ServersabstractAlthough the secure sockets layer (SSL) is the most popular protocol to provide a secure channel between a client and a cluster-based network server, its high overhead degrades the server performance considerably, and thus, affects the server scalability. Therefore, improving the performance of SSL-enabled network servers is critical for designing scalable and high performance data centers. In this paper, we examine the impact of SSL offering and SSL-session aware distribution in cluster-based network servers. We propose a backend forwarding scheme, called ssl_with_bf that employs a low-overhead user-level communication mechanism like VIA to achieve good load balance among server nodes. We compare three distribution models for network servers: round robin (RR), ssl_with_session and ssl_with_bf through simulation. The experimental results with 16-node and 32-node cluster configurations show that while session reuse of ss_with_session is critical to improve the performance of application servers, the proposed backend forwarding scheme can further enhance the performance due to better load balancing. The ssl_with_bf scheme can minimize average latency by about 40% and improve throughput across a variety of workloads Jin-Ha Kim, Gyu Sang Choi, Chita R. Das |
CLUSTER | 3 |
| 2005 | A low latency router supporting adaptivity for on-chip interconnectsabstractThe increased deployment of System-on-Chip designs has drawn attention to the limitations of on-chip interconnects. As a potential solution to these limitations, Networks-on -Chip (NoC) have been proposed. The NoC routing algorithm significantly influences the performance and energy consumption of the chip. We propose a router architecture which utilizes adaptive routing while maintaining low latency. The two-stage pipelined architecture uses look ahead routing, speculative allocation, and optimal output path selection concurrently. The routing algorithm benefits fromcongestionaware flow control, making better routing decisions. We simulate and evaluate the proposed architecture in terms of network latency and energy consumption. Our results indicate that the architecture is effective in balancing the performance and energy of NoC designs. Jongman Kim, Dongkook Park, Theocharis Theocharides, Narayanan Vijaykrishnan, Chita R. Das |
DAC | 5 |
| 2005 | Rcast: A Randomized Communication Scheme for Improving Energy Efficiency in MANETsabstractIn a typical wireless mobile ad hoc network (MANET) using a shared communicationmedium, every node receives or overhears every data transmission occurring in its vicinity. However, this technique is not applicable when a power saving mechanism (PSM) such as the one specified in IEEE 802.11 is employed, where a packet advertisement period is separated from the actual data transmission period.When a node receives an advertised packet that is not destined to itself, it switches to a low-power state during the data transmission period, and thus, conserves power. However, since some MANET routing protocols such as Dynamic Source Routing (DSR) collect route information via overhearing, they would suffer if they are used with the IEEE 802.11 PSM. Allowing no overhearing may critically deteriorate the performance of the underlying routing protocol, while unconditional overhearing may offset the advantage of using PSM. This paper proposes a new communication mechanism, called RandomCast or Rcast, via which a sender can specify the desired level of overhearing in addition to the intended receiver. Therefore, it is possible that only a random set of nodes overhear and collect route information for future use. Rcast improves not only the energy efficiency, but also the energy balance among the nodes, without significantly affecting the routing efficiency. Extensive simulation using the ns-2 network simulator shows that Rcast is highly energy-efficient compared to the original IEEE 802.11 PSM and On-Demand Power Management (ODPM) protocol in terms of total energy consumption (157% to 236% less than PSM and 28% to 131% less than ODPM) and energy balance (four times less variance than ODPM) among the nodes. Sunho Lim, Chansu Yu, Chita R. Das |
ICDCS | 3 |
| 2005 | Improving Performance of Cluster-based Secure Application Servers with User-level CommunicationabstractIn this paper, we have investigated the performance implications of SSL protocol for providing secure service in a cluster-based application server, and have proposed a backend forwarding scheme for improving server performance through better load balance. The proposed SSL/spl I.bar/with/spl I.bar/bf scheme exploits the underlying user-level communication minimizing the intra-cluster communication overhead. All results in this paper indicate that the proposed backend forwarding scheme is a viable mechanism for improving the performance of application servers in cluster-based data centers. Jin-Ha Kim, Gyu Sang Choi, Chita R. Das |
ICDE | 3 |
| 2005 | A multi-threaded PIPELINED Web server architecture for SMP/SoC machinesabstractDesign of high performance Web servers has become a recent research thrust to meet the increasing demand of network-based services. In this paper, we propose a new Web server architecture, called multi-threaded PIPELINED Web server, suitable for Symmetric Multi-Processor (SMP) or System-on-Chip (SoC) architectures. The proposed PIPELINED model consists of multiple thread pools, where each thread pool consists of five basic threads and two helper threads. The main advantages of the proposed model are global information sharing by the threads, minimal synchronization overhead due to less number of threads, and non-blocking I/O operations, possible with the helper threads.We have conducted an in-depth performance analysis of the proposed server model along with four prior Web server models (Multi-Process (MP), Multi-Thread (MT), Single-Process Event-Driven (SPED) and Asynchronous Multi-Process Event-Driven (AMPED)) via simulation using six Web server workloads. The experiments are conducted to investigate the impact of various factors such as the memory size, disk speed and numbers of clients. The simulation results indicate that the proposed PIPELINED Web server architecture shows the best performance across all system and workload parameters compared to the MP, MT, SPED and AMPED models. Although the MT and AMPED models show competitive performance with less number of processors, the advantage of the PIPELINED model becomes obvious as the number of processors or clients in an SMP/SoC machine increases. The MP model shows the worst performance in most of the cases. The results indicate that the proposed server architecture can be used in future large-scale SMP/SoC machines to boost system performance. Gyu Sang Choi, Jin-Ha Kim, Deniz Ersoz, Chita R. Das |
WWW | 4 |
| 2005 | Performance analysis of a QoS capable cluster interconnect
Eun Jung Kim 0001, Ki Hwan Yum, Chita R. Das |
Perform. Evaluation | 3 |
| 2005 | A Holistic Approach to Designing Energy-Efficient Cluster InterconnectsabstractDesigning energy-efficient clusters has recently become an important concern to make these systems economically attractive for many applications. Since the cluster interconnect is a major part of the system, the focus of this paper is to characterize and optimize the energy consumption in the entire interconnect. Using a cycle-accurate simulator of an InfiniBand Architecture (IBA) compliant interconnect fabric and actual designs of its components, we investigate the energy behavior on regular and irregular interconnects. The energy profile of the three major components (switches, network interface cards (NICs), and links) reveals that the links and switch buffers consume the major portion of the power budget. Hence, we focus on energy optimization of these two components. To minimize power in the links, first we investigate the dynamic voltage scaling (DVS) algorithm and then propose a novel dynamic link shutdown (DLS) technique. The DLS technique makes use of an appropriate adaptive routing algorithm to shut down the links intelligently. We also present an optimized buffer design for reducing leakage energy in 70nm technology. Our analysis on different networks reveals that, while DVS is an effective energy conservation technique, it incurs significant performance penalty at low to medium workload. Moreover, energy saving with DVS reduces as the buffer leakage current becomes significant with 70nm design. On the other hand, the proposed DLS technique can provide optimized performance-energy behavior (up to 40 percent energy savings with less than 5 percent performance degradation in the best case) for the cluster interconnects. Eun Jung Kim 0001, Greg M. Link, Ki Hwan Yum, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, Chita R. Das |
IEEE Trans. Computers | 7 |
| 2004 | Improving Response Time in Cluster-Based Web Servers through CoschedulingabstractSummary form only given. We investigate the feasibility of minimizing the response time of a Web server by exploiting the advantages of both user-level communication and coscheduling. We, thus, propose a coscheduled server model, based on the PRESS design where the remote cache accesses can be coscheduled on different nodes to reduce the response time. We experiment this concept using two known coscheduling techniques, called dynamic coscheduling (DCS) and DCS with immediate blocking. Extensive simulation of four server models (PRESS over TCP/IP, PRESS over VIA, coscheduled PRESS model with DCS, and with DCS and blocking) using 16-node and 32-node cluster configurations indicates that the average response time of a distributed server can be minimized significantly by coscheduling the communicating processes. The use of the DCS scheme reduced the average latency up to 80%, on an average 40%, compared to the PRESS over VIA model that uses only user-level communication, and by order of magnitude compared to the TCP/IP model. The throughput of the three user-level communication models is at least 25% better compared to the PRESS over TCP/IP model. Experiments with varying file size and cache size also confirmed the advantage of using a coscheduling mechanism for improving the response time behavior. Jin-Ha Kim, Gyu Sang Choi, Deniz Ersoz, Chita R. Das |
IPDPS | 4 |
| 2004 | Performance comparison of cache invalidation strategies for Internet-based mobile ad hoc networksabstractInternet-based mobile ad hoc network (IMANET) combines a mobile ad hoc network (MANET) and the Internet to provide universal information accessibility. Although caching frequently accessed data items in mobile terminals (MTs) improves the communication performance in an IMANET, it brings a critical design issue when data items are updated. We analyze several push and pull-based cache invalidation strategies for IMANETS. A global positioning system (GPS) based connectivity estimation (GPSCE) scheme is first proposed to assess the connectivity of an MT for supporting any cache invalidation mechanism. Then, we propose a pull-based approach, called aggregate cache based on demand (ACOD) scheme, to find the queried data items efficiently. In addition, we modify two push-based cache invalidation strategies, proposed for cellular networks, to work in IMANETs. These are a modified timestamp (MTS) scheme, and an MTS with updated invalidation report (MTS+UIR) scheme. Simulation results indicate that our proposed strategy provides high throughput, low query latency, and low communication overhead, and thus, is a viable approach for implementation in IMANETS. Sunho Lim, Wang-Chien Lee, Guohong Cao, Chita R. Das |
MASS | 4 |
| 2004 | Coscheduling in Clusters: Is It a Viable Alternative?abstractIn this paper, we conduct an in-depth evaluation of a broad spectrum of scheduling alternatives for clusters. These include the widely used batch scheduling, local scheduling, gang scheduling, all prior communication-driven coscheduling algorithms (Dynamic Coscheduling (DCS), Spin Block (SB), Periodic Boost (PB), and Co-ordinated Coscheduling (CC)) and a newly proposed HYBRID coscheduling algorithm on a 16-node, Myrinet-connected Linux cluster. Performance and energy measurements using several NAS, LLNL and ANL benchmarks on the Linux cluster provide several interesting conclusions. First, although batch scheduling is currently used in most clusters, all blocking-based coscheduling techniques such as SB, CC and HYBRID and the gang scheduling can provide much better performance even in a dedicated cluster platform. Second, in contrast to some of the prior studies, we observe that blocking-based schemes like SB and HYBRID can provide better performance than spin-based techniques like PB on a Linux platform. Third, the proposed HYBRID scheduling provides the best performance-energy behavior and can be implemented on any cluster with little effort. All these results suggest that blocking-based coscheduling techniques are viable candidates to be used in clusters for significant performance-energy benefits. Gyu Sang Choi, Jin-Ha Kim, Deniz Ersoz, Andy B. Yoo, Chita R. Das |
SC | 5 |
| 2004 | An adaptive power-conserving service discipline for bluetooth (APCB) wireless networks
Hao Zhu 0007, Guohong Cao, George Kesidis, Chita R. Das |
Comput. Commun. | 4 |
| 2004 | Caching and Scheduling in NAD-Based Multimedia ServersabstractMultimedia-on-demand (MOD) applications have grown dramatically in popularity, especially in the domains of education, business, and entertainment. Current MOD servers waste precious resources in performing store-and-forward copying. This excessive overhead increases cost and severely limits the scalability of these servers. In this paper, we propose using the network-attached disk (NAD) architecture to design highly scalable and cost-effective MOD servers. In order to ensure enhanced performance, we propose a scheme, called distributed interval caching (DIG), which utilizes the on-disk buffers for caching intervals between successive streams. We also propose another scheme, called multiobjective scheduling (MOS), which increases the degrees of resource sharing by scheduling the waiting requests for service intelligently. We then integrate the two schemes and study the overall performance benefits through extensive simulation. The results demonstrate that the integrated policy works very well in increasing the number of customers that can be serviced concurrently while decreasing their waiting times for service. The performance benefits vary with several architectural, system workload, and scheduling parameters. We conclude this study by developing an analytical model for ideal DIG in order to estimate the performance limits which may be achieved through various optimizations. Nabil J. Sarhan, Chita R. Das |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2004 | A unified bandwidth reservation and admission control mechanism for QoS provisioning in cellular networksabstractAbstract We propose a unified framework consisting of a differential bandwidth reservation (DBR) algorithm and a Quality of Service (QoS)‐aware admission control scheme to provide QoS guarantees to on‐going connections in cellular networks. The differential bandwidth reservation policy uses a sector of cells in making the bandwidth reservation for accepting a new call. Based on the distance of the target cells in the sectors, two different bandwidth reservation policies are applied to optimize the connection dropping rate (CDR), while maintaining a competitive connection blocking rate (CBR). In addition, two possible mobile terminal (MT) movements are analyzed using the DBR mechanism. In the first case, no knowledge of an MT's moving path is assumed to be known, while in the second case, prior knowledge of a user profile is used in bandwidth reservation, and it is called user profile‐based DBR (UPDBR) algorithm. Using the DBR scheme, we propose an admission control algorithm that uses varying number of cells in a sector to meet admission decisions. Extensive simulation is performed to evaluate our methodology. Comparison of the proposed scheme with two prior schemes shows that our approach is not only capable of providing better QoS guarantees, but is also flexible in terms of using varying number of cells in satisfying the high‐level QoS requirements. Copyright © 2004 John Wiley & Sons, Ltd. Sunho Lim, Guohong Cao, Chita R. Das |
Wirel. Commun. Mob. Comput. | 3 |
| 2003 | Co-Ordinated Coscheduling in Time-Sharing Clusters through a Generic FrameworkabstractIn this paper, we attempt to address several key issues in designing coscheduling algorithms for clusters. First, we propose a generic framework for deploying coscheduling techniques by providing a reusable and dynamically loadable kernel module. Second, we implement all prior dynamic coscheduling algorithms (dynamic coscheduling (DCS), spin block (SB) and periodic boost (PB)) and a new coscheduling technique, called co-ordinated coscheduling (CC), using the above framework. Third, with exhaustive experimentation using mixed workloads, we observe that unlike PB, which provided the best performance on a Solaris platform (followed by SB and DCS), the proposed CC scheme outperforms all other techniques on a Linux platform, followed by SB, PB and DCS, in that order. Finally, we argue that due to its modular design, portable implementation on a standard platform, high performance and tolerance to workload mixes, the proposed CC scheme can be a viable scheduling option for time-sharing clusters. Gyu Sang Choi, Chita R. Das, Andy B. Yoo, Shailabh Nagar |
CLUSTER | 3 |
| 2003 | Impact of Job Allocation Strategies on Communication-Driven Coscheduling in Clusters
Gyu Sang Choi, Jin-Ha Kim, Andy B. Yoo, Chita R. Das |
Euro-Par | 5 |
| 2003 | A control theoretic approach for designing adaptive AQM schemesabstractIn this paper, we use a control theoretic approach to develop a generic framework for analyzing various active queue management (AQM) schemes as proportional-integral-derivative (PID) controllers. Based on this PID model, we propose an adaptive control mechanism to improve the system stability and performance under changing network conditions. We then present a generic implementation of the PID controller by introducing a derivative control into a PI controller. In addition, we propose an improved adaptive virtual queue (AVQ) scheme with explicit queue length control. A simulation study under a wide range of traffic conditions suggests that the proposed algorithms outperform the existing AQM schemes in achieving better system performance and stability. Xidong Deng, Sungwon Yi, George Kesidis, Chita R. Das |
GLOBECOM | 4 |
| 2003 | Performance Enhancement Techniques for InfiniBand? ArchitectureabstractThe InfiniBand/sup TM/ Architecture (IBA) is envisioned to be the default communication fabric for future system area networks (SAN). However, the released IBA specification outlines only higher level functionalities, leaving it open for exploring various design alternatives. In this paper we investigate four co-related techniques to provide high and predictable performance in IBA. These are: (i) using the shortest path first (SPF) algorithm for deterministic packet routing; (ii) developing a multipath routing mechanism for minimizing congestion; (iii) developing a selective packet dropping scheme to handle deadlock and congestion; and (iv) providing multicasting support for customized applications. These designs are evaluated using an integrated workload on a versatile IBA simulation testbed. Simulation results indicate that the SPF routing, multipath routing, packet dropping, and multicasting schemes are quite effective in delivering high and assured performance in clusters. One of the major contributions of this research is the IBA simulation testbed, which is an essential tool to evaluate various design tradeoffs. Eun Jung Kim 0001, Ki Hwan Yum, Chita R. Das, Mazin S. Yousif, José Duato |
HPCA | 3 |
| 2003 | A novel caching scheme for Internet based mobile ad hoc networksabstractInternet based mobile ad hoc network (IMANET) is an emerging technique that combines a wired network (e.g. Internet) and a mobile ad hoc network (manet) for developing a ubiquitous communication infrastructure. However, imanet has several limitations to fulfill users' demands to access various kinds of information such as limited accessibility to the wired Internet, insufficient wireless bandwidth, and longer message latency. In this paper, we address the issues involved in information search and access in IMANET. A broadcast based simple search (SS) algorithm and an aggregate caching mechanism are proposed for improving the information accessibility and reducing average communication latency in imanet. As part of the aggregate cache, a cache admission control policy and a cache replacement policy, called time and distance sensitive (TDS) replacement, are developed to reduce the cache miss ratio and improve the information accessibility. We evaluate the impact of caching, cache management, and access points, which are connected to the Internet, through extensive simulation. The simulation results indicate that the proposed aggregate cache can significantly improve an imanet performance in terms of throughput and average number of hops to access data. In particular, with aggregate caching, more than 200% improvement in throughput is achieved compared to the imanet with no cache case, when the access pattern follows a Zipf distribution. Sunho Lim, Wang-Chien Lee, Guohong Cao, Chita R. Das |
ICCCN | 4 |
| 2003 | An Integrated Resource Sharing Policy for Multimedia Storage Servers Based on Network-Attached DisksabstractIn this paper we propose using the network-attached disk (NAD) architecture to design highly scalable and cost-effective multimedia-on-demand (MOD) servers. In order to ensure enhanced performance, we propose two schemes, called distributed interval caching (DIC) and multi-objective scheduling (MOS). The DIC scheme utilizes the on-disk buffers for caching intervals between successive streams, while the MOS scheme improves resource sharing by scheduling requests for service intelligently based on four predefined criteria. We then integrate the two schemes and study the overall performance benefits through extensive simulation. We also study the effectiveness of the proposed DIC scheme by developing an analytical model that estimates the performance limit of DIC The results demonstrate that the integrated policy works very well in increasing the number of customers that can be serviced concurrently while decreasing their waiting times, and that the performance improvements scale with the number of disks in the server. Nabil J. Sarhan, Chita R. Das |
ICDCS | 2 |
| 2003 | Energy optimization techniques in cluster interconnectsabstractDesigning energy-efficient clusters has recently become an important concern to make these systems economically attractive for many applications. Since the links and switch buffers consume the major portion of the power budget of the cluster, the focus of this paper is to optimize the energy consumption in these two components. To minimize power in the links, we propose a novel dynamic link shutdown (DLS) technique. The DLS technique makes use of an appropriate adaptive routing algorithm to shutdown the links intelligently. We also present an optimized buffer design for reducing leakage energy. Our analysis on different networks using a complete system simulator reveals that the proposed DLS technique can provide optimized performance-energy behavior (up to 40% energy savings with less than 5% performance degradation in the best case) for the cluster interconnects. Eun Jung Kim 0001, Ki Hwan Yum, Greg M. Link, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, Mazin S. Yousif, Chita R. Das |
ISLPED | 8 |
| 2003 | An End-to-End Resource Scheduling Scheme for the Presentation of Composite Multimedia Information in a Networked Environment
Suneuy Kim, Chita R. Das |
MMM | 2 |
| 2002 | Integrated Admission and Congestion Control for QoS Support in ClustersabstractAdmission and congestion control mechanisms are integral parts of any Quality of Service (QoS) design for networks that support integrated traffic. In this paper we propose an-admission control algorithm and a congestion control algorithm for clusters, which are increasingly being used in a diverse set of applications that require QoS guarantees. The uniqueness of our approach is that we develop these algorithms for wormhole-switched networks. We use QoS-capable wormhole routers and QoS-capable network interface cards (NICs), referred to as Host Channel Adapters (HCAs) in InfiniBand/spl trade/ Architecture (IBA), to evaluate the effectiveness of these algorithms. The admission control is applied at the HCAs and the routers, while the congestion control is deployed only at the HCAs. Simulation results indicate that the admission and congestion control algorithms are quite effective in delivering the assured performance. The proposed credit-based congestion control algorithm is simple and practical in that it relies on hardware already available in the HCA to regulate traffic injection. Ki Hwan Yum, Eun Jung Kim 0001, Chita R. Das, Mazin S. Yousif, José Duato |
CLUSTER | 3 |
| 2002 | Stabilized virtual buffer (SVB) - an active queue management scheme for Internet quality-of-serviceabstractWe present a virtual queue-based active queue management (AQM) scheme, called stabilized virtual buffer (SVB). The SVB scheme uses the packet arrival rate and queue length information to drop/mark packets probabilistically in a congested Internet router. System goodput, packet loss rate, average queue length, and stability of the queue are used to compare the proposed SVB scheme with prior AQM schemes (RED - random early detection; REM - random exponential marking; AVQ - adaptive virtual queue). Simulation results indicate that the SVB algorithm can provide better goodput and lower loss rate than the other three AQMs. The most striking feature of the proposed scheme is its robustness to workload fluctuations in maintaining a stable queue for different workload mixes (short and long flows) and parameter settings. Xidong Deng, Sungwon Yi, George Kesidis, Chita R. Das |
GLOBECOM | 4 |
| 2002 | Providing fairness in DiffServ architectureabstractThe Differentiated Service (DiffServ) architecture does not specify any priority scheme between assured forwarding (AF) out-profile packets and best-effort (BE) packets. Therefore, a misbehaving AF flow can penalize many BE flows unless a fair bandwidth sharing mechanism is employed in the routers. In this paper, we propose two different techniques for solving the inter- and intra-class fairness problems at the core and edge routers, respectively. For the core routers, we propose a fair weighted round robin (FWRR) scheduler that protects BE packets from monopolizing AF out-profile packets by dynamically adjusting the service weights and buffer spaces according to the traffic changes. For the edge routers, we propose a scheme, called fair dropper (FD), that provides intra-class fairness by penalizing the greedy flows. Simulation results indicate that both these techniques are quite effective in providing inter- and intra-class fairness, while maintaining a low packet loss rate. Sungwon Yi, Xidong Deng, George Kesidis, Chita R. Das |
GLOBECOM | 4 |
| 2002 | An adaptive power-conserving service discipline for BluetoothabstractBluetooth is a new short-range radio technology to form a small wireless system. In most of the current Bluetooth products, the master polls the slaves in a round robin manner and it may waste a significant amount of power. We propose an adaptive power conserving scheme to address this problem. The proposed solution schedules each flow based on its predictive rate and achieves power optimization based on a low-power mode existing in Bluetooth standard. Unlike other research work related to low-power, we also consider QoS of each flow. Theoretical analyses verify that our scheme can achieve throughput guarantees, delay guarantees, and fairness guarantees. Simulation results demonstrate that our scheme can save a significant amount of power compared to the round robin scheme and it shows that there exists a tradeoff between power and delay under various traffic models. Hao Zhu 0007, Guohong Cao, George Kesidis, Chita R. Das |
ICC | 4 |
| 2002 | Power-Aware Prefetch in Mobile EnvironmentsabstractMost of the prefetch techniques used in the current cache management schemes do not consider the power constraints of the mobile clients and other factors such as the size of the data items, the data access rate, and the data update rate. We address these issues by proposing a power-aware prefetch scheme, called the value-based adaptive prefetch (VAP) scheme. The VAP scheme defines a value function which can optimize the prefetch cost to achieve better performance. Also, VAP dynamically adjusts the number of prefetches based on the current energy level to prolong the system running time. As stretch is widely adopted as a performance metric for variable-size data requests, we show by analysis that the proposed algorithm can indeed achieve the optimal performance in terms of stretch when power consumption is considered. Simulation results demonstrate that our algorithm significantly outperforms existing prefetching algorithms under various scenarios. Liangzhong Yin, Guohong Cao, Chita R. Das, Ajeesh Ashraf |
ICDCS | 3 |
| 2002 | An admission control scheme for QoS-sensitive cellular networksabstractWe propose an admission control scheme to guarantee a certain level of QoS to on-going connections in cellular networks. This admission control scheme is based on a differential bandwidth reservation policy that uses a sector of cells in making bandwidth reservation for accepting the new call. The sector of cells, which are located along the way to which the MT might move, is further divided into two regions depending on whether they have an immediate impact on the handoff or not. Two different bandwidth reservation policies are applied to cells in the two regions to optimize the connection dropping rate (CDR) while maximizing the connection blocking rate (CBR). In contrast to most prior policies, the proposed admission control scheme uses the varying number of cells in the sector to make the admission decision. Depending on the currently measured average CDR of the cells in the sector and the current cell where a new connection is generated, the number of cells involved in admission control can be changed dynamically to satisfy the target QoS (CDR) parameter. Simulation results indicate that our admission control policy guarantees the required CDR over the entire workload, while maintaining a competitive CBR. Comparison of the proposed scheme with two prior schemes shows that our approach is not only capable of providing better QoS guarantees, but also is more flexible in terms of using varying number of cells in satisfying a certain QoS requirement. Sunho Lim, Guohong Cao, Chita R. Das |
WCNC | 3 |
| 2002 | A Fast and Efficient Processor Allocation Scheme for Mesh-Connected MulticomputersabstractEfficient processor allocation is crucial for obtaining high performance in space-shared parallel computers. A good processor allocation algorithm should find available processors for incoming jobs, if they exist, with minimum overhead. In this paper, we propose such a fast and efficient processor allocation scheme for mesh-connected multicomputers. By using simple coordinate calculation and spatial subtraction, the proposed scheme reduces the search space drastically and, hence, can locate a free submesh very quickly. The algorithm is implemented efficiently using a stack and therefore is called the stack-based allocation (SBA) algorithm. Extensive simulation reveals that our scheme incurs much less allocation overhead than all of the existing allocation algorithms, while delivering competitive performance. Byung S. Yoo, Chita R. Das |
IEEE Trans. Computers | 2 |
| 2002 | MediaWorm: A QoS Capable Router Architecture for ClustersabstractWith the increasing use of clusters in real-time applications, it has become essential to design high-performance networks with quality-of-service (QoS) guarantees. We explore the feasibility of providing QoS in wormhole switched routers, which are widely used in designing scalable, high-performance cluster interconnects. In particular, we are interested in supporting multimedia video streams with CBR and VBR traffic, in addition to the conventional best-effort traffic. The proposed MediaWorm router uses a rate-based bandwidth allocation mechanism, called Fine-Grained VirtualClock (FGVC), to schedule network resources for different traffic classes. Our simulation results on an 8-port router indicate that it is possible to provide jitter-free delivery to VBR/CBR traffic up to an input load of 70-80 percent of link bandwidth and the presence of best-effort traffic has no adverse effect on real-time traffic. Although the MediaWorm router shows a slightly lower performance than a pipelined circuit switched (PCS) router, commercial success of wormhole switching, coupled with simpler and cheaper design, makes it an attractive alternative. Simulation of a (2/spl times/2) fat-mesh using this router shows performance comparable to that of a single switch and suggests that clusters designed with appropriate bandwidth balance between links can provide required performance for different types of traffic. Ki Hwan Yum, Eun Jung Kim 0001, Chita R. Das, Aniruddha S. Vaidya |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2001 | Selective Checkpointing and Rollbacks in Multithreaded Distributed SystemsabstractModern distributed systems are often multithreaded and object-oriented in their design. They require efficient techniques to checkpoint and restore their state for improving fault-tolerance properties. The traditional process-based techniques of distributed checkpointing and rollback algorithms suffer from the problem of false dependencies, which makes them very rigid and inefficient for use with modern systems. In this paper, we develop protocols that can selectively checkpoint (and rollback) some threads of a distributed system while leaving others untouched, and yet ensuring the consistency of state resulting from such a partial rollback. Mangesh Kasbekar, Chita R. Das |
ICDCS | 2 |
| 2001 | Adaptive Block Rearrangement Algorithms for Video-On-Demand ServersabstractVideo-on-demand (VOD) is increasingly becoming one of the most important and successful services due to the recent advances in storage subsystems, compression technology and, networking. Therefore, the investigation of various alternatives to improve the performance of VOD servers has become a major research focus. The reduction of disk access time through intelligent data placement strategies is one such avenue and is the theme of this paper: Movie rental patterns indicate that accesses to movies are highly localized with only a small number of movies receiving most of the accesses. In this paper we exploit the access patterns and propose an adaptive rearrangement of the blocks on each disk within the server. With this approach, the blocks of the movies with comparable access frequencies are kept closer to each other We analyze two rearrangement schemes, called centered and sequential. In the centered layout, blocks are placed according to their access patterns starting with the most popular movie at the center. The sequential layout places movies in the order of their popularity starting at the edge of the disk. We compare and evaluate, through an intensive simulation study, the effectiveness of these layouts with respect to arbitrary layouts. The simulation results indicate that significant disk improvements could be attained by adopting the proposed schemes, and that the centered layout is the best performer. Nabil J. Sarhan, Chita R. Das |
ICPP | 2 |
| 2001 | QoS provisioning in clusters: an investigation of Router and NIC designabstractDesign of high performance cluster networks (routers) with Quality-of-Service (QoS) guarantees is becoming increasingly important to support a variety of multimedia applications, many of which have real-time constraints. Most commercial routers, which are based on the wormhole-switching paradigm, can deliver high performance, but lack QoS provisioning. In this paper, we present a pipelined wormhole router architecture that can provide high and predictable performance for integrated traffic in clusters. We consider two different implementations—a non-preemptive model and a more aggressive preemptive model. We also present the design of a network interface card (NIC) based on the Virtual Interface Architecture (VIA) design paradigm to support QoS in the NIC. The QoS capable router and NIC designs are evaluated with a mixed workload consisting of best-effort traffic, multimedia streams, and control traffic. Ki Hwan Yum, Eun Jung Kim 0001, Chita R. Das |
ISCA | 3 |
| 2001 | Calculation of Deadline Missing Probability in a QoS Capable Cluster InterconnectabstractThe growing use of clusters in diverse applications, many of which have real-time constraints, requires Quality-of-Service (QoS) support from the underlying cluster interconnect. In this paper we propose an analytical model that captures the characteristics of a QoS capable wormhole router which is the basic building block of cluster networks. The model captures the behavior of integrated traffic in a cluster and computes the average deadline missing probability for real-time traffic. The cluster interconnect, considered here, is a hypercube network. Comparison of Deadline Missing Probability (DMP) using the proposed model with that of the simulation shows that our analytical model is accurate and useful. Eun Jung Kim 0001, Ki Hwan Yum, Chita R. Das |
NCA | 3 |
| 2001 | On the Effectiveness of a Counter-Based Cache Invalidation Scheme and Its Resiliency to Failures in Mobile EnvironmentsabstractCaching frequently accessed data items on the client side is an effective technique to improve the performance of data dissemination in mobile environments. Classical cache invalidation strategies are not suitable for mobile environments due to the disconnection and mobility of the mobile clients. One attractive cache invalidation technique is based on invalidation reports (IRs). However, IR-based approach suffers from long query latency and it cannot efficiently utilize the broadcast bandwidth. In this paper, we propose techniques to address these problems. We first extend the UIR-based approach to reduce the query latency. Then, we propose techniques to efficiently utilize the broadcast bandwidth based on counters associated with each data item. Novel techniques are designed to maintain the accuracy of the counter in case of server failures, client failures, and disconnections. Extensive simulations are provided and used to evaluate the proposed methodology. Compared to previous IR-based algorithms, the proposed solution can significantly reduce the query latency, improve the bandwidth utilization, and effectively deal with disconnections and failures. Guohong Cao, Chita R. Das |
SRDS | 2 |
| 2001 | Efficient processor management schemes for mesh-connected multicomputers
Byung S. Yoo, Chita R. Das |
Parallel Comput. | 2 |
| 2001 | Impact of Virtual Channels and Adaptive Routing on Application PerformanceabstractResearch on multiprocessor interconnection networks has primarily focused on wormhole switching, virtual channel flow control, and routing algorithms to enhance their performance. The rationale behind this research is that by alleviating the network latency for high network loads, the overall system performance would improve; many studies have used synthetic workloads to support this claim. However, such workloads may not necessarily capture the behavior of real applications. In this paper, we have used parallel applications for a closer examination of the network behavior. In particular, the performance benefit from enhancing a 2D mesh with virtual channels (VCs) and a fully adaptive routing algorithm is examined with a set of shared-memory and message passing applications. Execution time and average message latency of shared memory applications are measured using execution-driven simulation and by varying many architectural attributes that affect the network workload. The communication traces of message passing applications, collected on an IBM-SP2, are used to run a trace-driven simulation of the mesh architecture to obtain message latency. Simulation results show that VCs and adaptive routing can reduce the network latency to varying degrees depending on the application. However, these modest benefits do not translate to significant improvements in the overall execution time because the load on the network is not high enough to exploit the advantages of the network enhancements. Moreover, this benefit may be negated if the architectural enhancements increase the network cycle time. Rather, emphasis should be placed on improving the raw network bandwidth and faster network interfaces. Aniruddha S. Vaidya, Anand Sivasubramaniam, Chita R. Das |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2000 | Investigating QoS Support for Traffic Mixes with the MediaWorm RouterabstractWith the increasing use of clusters in real-time applications, it has become essential to design high performance networks with quality of service (QoS) guarantees. In this paper, we explore the feasibility of providing QoS in worm-hole switched routers, which are otherwise well known for designing high performance interconnects. In particular, we are interested in supporting multimedia video streams, in addition to the conventional best-effort traffic. The proposed MediaWorm router uses a rate-based bandwidth allocation mechanism, called Virtual Clock, to schedule network resources for different traffic classes. Our simulation results on an 8-port router indicate that it is possible to provide jitter-free delivery to VBR/CBR traffic up to an input load of 70-80% of link bandwidth, and the presence of best effort traffic has no adverse effect on the real-time traffic. Although the MediaWorm router shows a slightly lower performance than a pipelined circuit switched (PCS) router, commercial success of worm-hole switching coupled with the simpler and cheaper design makes it an attractive alternative. Simulation of a (2/spl times/2) fat-mesh using this router suggests that clusters designed with appropriate bandwidth balance between links can provide good performance for different types of traffic. Ki Hwan Yum, Aniruddha S. Vaidya, Chita R. Das, Anand Sivasubramaniam |
HPCA | 3 |
| 2000 | A Reliable Statistical Admission Control Strategy for Interactive Video-on-Demand Servers with Interval CachingabstractAn admission control algorithm is a key component of a video server that supports quality of service (QoS). Such an algorithm determines whether or not a new request can be admitted to the server without compromising the required performance of the in-service users. All prior admission control policies for video servers have focussed on improving the underlying storage system utilization such that a VOD server can accept more requests while satisfying the QoS requirements. In this paper, we examine admission control for an interactive video server that uses a resource sharing mechanism, called interval caching. The primary motivation of this work is to develop an efficient admission policy to optimize the server throughput with acceptable QoS. We analyze four admission control algorithms, called deterministic, predictive, statistical average, and reliable statistical admission control (RSAC), for an interactive video server and find that the proposed RSAC scheme is quite efficient compared to the other three schemes in terms of throughput and jitter. The RSAC policy attempts to optimize the disk and cache utilization while reserving a certain amount of disk bandwidth for streams that are evicted from the cache, and are likely to cause jitter due to unavailable disk space. We estimate this reserved disk bandwidth as a function of the average number of cached steams, which in turn indicates improvement in server capacity with interval caching. The average number of cached streams and the subsequent reserved bandwidth are computed wing a simple, yet accurate mathematical model. Extensive performance evaluation indicates that the RSAC is an attractive scheme for an interactive or non-interactive video server and can be implemented with other performance enhancement techniques. Suneuy Kim, Chita R. Das |
ICPP | 2 |
| 1999 | LAPSES: A Recipe for High Performance Adaptive Router DesignabstractEarlier research has shown that adaptive routing can help in improving network performance. However, it has not received adequate attention in commercial routers mainly due to the additional hardware complexity, and the perceived cost and performance degradation that may result from this complexity. These concerns can be mitigated if one can design a cost-effective router that can support adaptive routing. This paper proposes a three step recipe-Look-Ahead routing, intelligent Path Selection, and an Economic Storage implementation, called the LAPSES approach-for cost-effective high performance pipelined adaptive router design. The first step, look-ahead routing, reduces a pipeline stage in the router by making table lookup and arbitration concurrent. Next, three new traffic-sensitive path selection heuristics (LRU, LFU and MAX-CREDIT) are proposed to select one of the available alternate paths. Finally, two techniques for reducing routing table size of the adaptive router are presented. These are called meta-table routing and economical storage. The proposed economical storage needs a routing table with only 9 and 27 entries for two and three dimensional meshes, respectively. All these design ideas are evaluated on a (16/spl times/16) mesh network via simulation. A fully adaptive algorithm and various traffic patterns are used to examine the performance benefits. Performance results show that the look-ahead design as well as the path selection heuristics boost network performance, while the economical storage approach turns out to be an ideal choice in comparison to full-table and meta-table options. We believe the router resulting from these three design enhancements can make adaptive routing a viable choice for interconnects. Aniruddha S. Vaidya, Anand Sivasubramaniam, Chita R. Das |
HPCA | 3 |
| 1999 | A Parallel Optimal Branch-and-Bound Algorithm for MIN-Based MultiprocessorsabstractA parallel Optimal Best-First search Branch-and-Bound (B&B) algorithm (obs) is proposed and evaluated for MIN-based multiprocessor systems. The proposed algorithm decomposes a problem into a number of subproblems and each subproblem is processed on a small group of processors. A performance analysis is conducted to estimate the speed-up of the proposed parallel B&B algorithm. It considers both the computation and communication times to evaluate the realistic performance. Simulation data are given, along with analysis results for model validation. It is shown that the proposed algorithm performs better than other reported schemes with its various advantageous features such as: less subproblem evaluations, proper load balancing, and limited scope of remote communication through the network. Myung K. Yang, Chita R. Das |
ICPP | 2 |
| 1999 | A Closer Look at Coscheduling Approaches for a Network of WorkstationsabstractEfficient scheduling of processes on processors of a Network of Workstations (NOW) is essential for good system performance.However, the design of such schedulers is challenging because of the complex interaction between several system and workload parameters.Coscheduling, though desirable, is impractical for such a loosely coupled environment.71vo operations, waiting for a message and arrival of a message, can be used to take remedial actions that can guide the behavior of the system towards coscheduling using local information.We present a taxonomy of three possibilities for each of these two operations, leading to a design space of 3 x 3 scheduling mechanisms.This paper presents an extensive implementation and evaluation exercise in studying these mechanisms.Adhering to the philosophy that scheduling and communication are intertwined and should be studied in conjunction, a complete communication substrate for UltraSPARC workstations, connected by Myrinet and running Solaris 2.5.1, has been developed.This platform provides the entire Message Passing Interface (MPI) to readily run off-the-shelf MPI applications by employing protected low-latency user-level messaging.Several applications can concurrently use this interface.This platform has been used to design, implement, and uniformly evaluate nine scheduling strategies with a mixture of concurrent real applications with varying communication intensities.This includes four new schemes (Periodic Boost, Periodic Boost with Spin Block, Spin Yield, Periodic Boost with Shailabh Nagar, Ajit Banerjee, Anand Sivasubramaniam, Chita R. Das |
SPAA | 4 |
| 1999 | Issues in the Design of a Reflective Library for Checkpointing C++ ObjectsabstractObject Persistence is an important feature of Object-oriented languages. The C++ language specification does not include or discuss any method of providing persistence for C++ objects. Several schemes have been developed for adding persistence to C++. Some of them require persistent objects to be allocated and treated differently than non-persistent objects, while some others require the programmer to provide vital parts of the persistence mechanism. It is desirable to make the persistence feature transparent, but the nature of C++ makes it difficult. This paper discusses in detail the various interesting language issues to be considered for adding persistence to C++ and how they lead to the design of the reflective object-checkpointing library, MemberAnalyzer. Mangesh Kasbekar, Chita R. Das, Shalini Yajnik, Reinhard Klemm, Yennun Huang |
SRDS | 2 |
| 1999 | Alternatives to Coscheduling a Network of Workstations
Shailabh Nagar, Ajit Banerjee, Anand Sivasubramaniam, Chita R. Das |
J. Parallel Distributed Comput. | 4 |
| 1999 | A Testbed for Evaluation of Fault-Tolerant Routing in Multiprocessor Interconnection NetworksabstractThis paper presents a comprehensive evaluation testbed for interconnection networks and routing algorithms using real applications. The testbed is flexible enough to implement any network topology and fault-tolerant routing algorithm, and allows the system architect to study the cost versus performance trade-offs for a range of network parameters. We illustrate its use with one fault-tolerant algorithm and analyze the performance of four shared memory applications with different fault conditions. We also show how the testbed can be used to drive future research in fault-tolerant routing algorithms and architectures by proposing and evaluating novel architectural enhancements to the network router, called path selection heuristics (PSH). We propose three such schemes and the Least Recently Used (LRU) PSH is shown to give the best performance in the presence of faults. Aniruddha S. Vaidya, Chita R. Das, Anand Sivasubramaniam |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1998 | Virtual channel multiplexing in networks of workstations with irregular topologyabstractNetworks of workstations are becoming a cost-effective alternative for small-scale parallel computing. Although they may not provide the closely coupled environment of multicomputers and multiprocessors, they meet the needs of a great variety of parallel computing problems at a lower cost. However in order to achieve a high efficiency, the interconnects used to build the network of workstations must provide a very high bandwidth and low latencies, making their design a critical issue. Recently, a very efficient flow control protocol for networks of workstations has been proposed by the authors. This protocol multiplexes physical channels between several virtual channels and minimizes the use of control flits by transmitting several data flits each time a virtual channel gets the link. In this protocol, a virtual channel sends data flits until the message blocks or is completely transmitted. However it can reduce network throughput, by increasing short message latency, due to long messages monopolizing channels and hindering the progress of short messages. In this paper, we analyze the impact of limiting the number of flits (block size) that a virtual channel can send once it gets the link. We propose a new version of the previous flow control protocol that is easily, implementable on hardware. Simulation results show that limiting the maximum block size is not a good design decision, because the overall network performance decreases. Only when short message latency is crucial is it is acceptable to limit the block size. Federico Silla, José Duato, Anand Sivasubramaniam, Chita R. Das |
HiPC | 4 |
| 1998 | The Penn State Computing Condominium Scheduling SystemabstractThe Penn State RS/6000 SP is a uniquely acquired and operated computing facility. This 143 CPU machine, centrally located and jointly owned, is a result of collaboration between academic departments, research groups, and the central academic computing facility. It is the largest on- campus resource at Penn State for meeting the high performance computing needs. Due to the joint ownership structure of the machine, the job scheduling requirements are significantly different from the usual methods of job processor allocation in distributed memory parallel machines. After several years of adapting different queuing systems, primarily the Distributed Queuing System, to our needs, it became obvious that the conventional scheduling systems did not serve the machine scheduling requirements unique to the Penn State SP. We concluded that a robust and easily configurable system needs to be developed to meet our unique needs. We have drawn inspiration from and modeled our system on EASY. As with EASY, we use the application programming interface of LoadLeveler to implement our scheduler. Our scheduler is named Penn State Condominium Scheduler (PSCS). PSCS does policy implementation and job execution on the machine is done by LoadLeveler. PSCS is written to facilitate easier configuration and administration. It does not have any processor architecture dependence. It is similar to the native scheduler in LoadLeveler in this regard. PSCS has incorporated three unique features: (i) node owner affinity which ensures fairness by allocation based on ownership, (ii) backfilling which ensures efficient utilization of resources, and (iii) affinity for services provided which ensures proper matching of jobs to the processors based on memory, software and other requirements. Jobs from users who own nodes in the SP complex have affinity to those particular processors owned by them. They also have preferences granted to them depending on their ownership level. Once the demand from the node owners is met, the next important goal is to keep the machine as fully occupied with running jobs as possible. This is accomplished by backfilling. This scheduler incorporates these features which are most important to successful implementation of multi-owner, centrally located, heterogeneous computing facilities. Pawan Agnihotri, Vijay K. Agarwala, Jeffrey J. Nucciarone, Kevin M. Morooney, Chita R. Das |
SC | 5 |
| 1998 | A Fast and Efficient Processor Management Scheme for k-ary n-cubes
Byung S. Yoo, Chita R. Das |
J. Parallel Distributed Comput. | 2 |
| 1997 | Towards a Communication Characterization Methodology for Parallel ApplicationsabstractThe interconnection network (ICN) is a vital component of a parallel machine and is often the limiting factor in the performance of several parallel applications. While ICN performance evaluation has been a widely researched topic, there have been very few studies that have used real applications to drive this research. In this paper we develop a framework for characterizing the communication properties of parallel applications. Message generation frequency, spatial distribution of messages and message length are the three attributes that quantify any communication. We develop a methodology to quantify these attributes, in particular the first two attributes. We employ two strategies, namely dynamic and static, in our methodology. In the former, the applications are executed on an execution-driven simulator called SPASM, while in the latter they are executed on a parallel machine, IBM SP2. We gather communication events from these executions and feed them to a 2-D mesh network simulator. The log of the network activity is then analyzed using a statistical analysis package (SAS) to find the message inter-arrival time distribution and spatial distribution via regression analysis. Five shared memory applications and two message passing applications are analyzed to quantify their communication workloads. It is shown that it is possible to express the message generation and spatial distribution of an application in terms of commonly used distributions. These distributions can be used in the analysis of ICNs for developing realistic performance models. Sucheta Chodnekar, Vijayalakshmi Srinivasan, Aniruddha S. Vaidya, Anand Sivasubramaniam, Chita R. Das |
HPCA | 5 |
| 1997 | A Performance Modeling Technique for Mesh-Connected MulticomputersabstractModeling the perfomance of space-shared multicomputers is a non-trivial task mainly due to difficulty in modeling the effect of external fragmentation on system performance. Mesh-connected multicomputers are hard to model in particular because of great variance in job sizes. Therefore, researchers have relied on simulation method to evaluate the mesh performance. We propose a novel modeling technique called hybrid method in this paper. The proposed technique utilizes simulation method to estimate the capacity of a system. Then, a queueing model with multiple servers is constructed using the system capacity as the number of servers in the queueing system. The technique is validated through simulation experiments. The results reveal that the hybrid method provides very close estimation of the mesh performance with very little overhead. The proposed technique can also be used for performance modeling of other multicomputers with different topologies. Byung S. Yoo, Chita R. Das, Jong Kim 0001 |
ICPADS | 2 |
| 1997 | Communication in Parallel Applications: Characterization and Sensitivity AnalysisabstractCommunication characterization of parallel applications is essential to understand the interplay between architectures and applications in determining the maximum achievable performance. Although a significant amount of research has been conducted on execution-based architectural evaluations, very little effort has gone into capturing the communication behavior of an application mathematically. In this paper, we attempt to characterize the communication behavior of applications by temporal, spatial and volume attributes. We also study the impact of variation in application and architectural parameters on the communication behavior in terms of the three attributes. Our results show that for the chosen suite of applications, the message arrival and spatial distributions can be closely approximated by known statistical distributions and that the temporal as well as spatial distributions of all applications remain unchanged with respect to four parameters considered in this study. These results lead us closer to the belief that it is possible to abstract the communication properties of parallel applications in convenient mathematical forms that have wide applicability. Dale Seed, Anand Sivasubramaniam, Chita R. Das |
ICPP | 3 |
| 1997 | Good Processor Management = Fast Allocation + Efficient SchedulingabstractFast and efficient processor allocation and job scheduling algorithms are essential components of a multi-user multicomputer operating system. In this paper we propose two novel processor management schemes which meet such demands for mesh-connected multicomputers. A stack-based allocation algorithm that can locate a free sub-mesh for a job very quickly using simple coordinate calculation and spatial subtraction is proposed. Simulation results show that the stack-based allocation algorithm outperforms all the existing allocation policies in terms of allocation overhead while delivering competitive performance. Another technique, called group scheduling, schedules jobs in such a way that the jobs belonging to the same group do not block each other. The groups are scheduled in an FCFS order to prevent starvation. This simple but efficient scheduling policy reduces the response rime significantly by minimizing the queueing delay for the jobs in the same group. These two schemes, when used together can provide faster service to users with very little overhead. Byung S. Yoo, Chita R. Das |
ICPP | 2 |
| 1997 | Performance Benefits of Virtual Channels and Adaptive Routing: An Application-Driven Study
Aniruddha S. Vaidya, Anand Sivasubramaniam, Chita R. Das |
International Conference on Supercomputing | 3 |
| 1997 | Performance Analysis of Buffering Schemes on Wormhole RoutersabstractWormhole switched input-buffered and middle-buffered routers with virtual channels are analyzed in this paper. Middle buffering refers to the placement of virtual channels between the demultiplexers and multiplexers of a crossbar switch. An analytical model for multistage interconnection networks using middle-buffered switches is developed. In addition, extensive simulation is conducted to assess the performance of the two buffering techniques in different network topologies. The study demonstrates that middle buffering with virtual channels provides better performance than input buffering with virtual channels in multistage interconnection networks, two-dimensional meshes, and hypercubes. Younes M. Boura, Chita R. Das |
IEEE Trans. Computers | 2 |
| 1996 | Parallel Simulation of Mesh Routing AlgorithmsabstractPerformance of a network depends primarily on the network topology, switching mechanism, flow control protocol and the underlying routing algorithm. While many routing algorithms have been proposed recently for direct networks, there is no time efficient technique to evaluate and compare all of them. A conventional routing algorithm simulation of a network on a uniprocessor takes unacceptably large computing time. The simulation can be made very time efficient by parallelizing it and running on a parallel test bed. This research is focussed on designing a parallel routing algorithm simulator for n-dimensional mesh connected networks with wormhole switching and virtual channel flow control. The research addresses partitioning mapping, synchronization issues, and implementation of various routing algorithms for 2-D and 3-D mesh architectures. Experimental results show that the parallel simulator can provide significant speedup compared do a uniprocessor environment. S. Rahman, Chita R. Das |
ICDCS | 2 |
| 1996 | Allocation and Mapping Based Reliability Analysis of Multistage Interconnection NetworksabstractTask allocation using cubic partitioning of multistage interconnection networks (MINs) offers several advantages over random allocation of resources. The objective of this paper is to analyze MIN reliability considering the cubic allocation algorithm. A comprehensive analytical model is derived for predicting reliability of MIN-based systems where tasks are allocated using the buddy strategy. System reliability with the free list allocation policy is computed via simulation. It is shown that the system reliability is dependent on the allocation algorithm and the free list policy is superior to the buddy scheme in this respect. Two types of mapping algorithms known as conventional and bit reversal are used on a baseline MIN to show that the same allocation algorithm can result in different reliability and performance. A performance-related reliability measure is analyzed using probability of acceptance as the performance measure to demonstrate the trade-offs between performance and reliability. Prasant Mohapatra, Chansu Yu, Chita R. Das |
IEEE Trans. Computers | 3 |
| 1996 | A probabilistic model for the fault tolerance of multilayer perceptronsabstractThis paper presents a theoretical approach to determine the probability of misclassification of the multilayer perceptron (MLP) neural model, subject to weight errors. The type of applications considered are classification/recognition tasks involving binary input-output mappings. The analytical models are validated via simulation of a small illustrative example. The theoretical results, in agreement with simulation results, show that, for the example considered, Gaussian weight errors of standard deviation up to 22% of the weight value can be tolerated. The theoretical method developed here adds predictability to the fault tolerance capability of neural nets and shows that this capability is heavily dependent on the problem data. Najwa Sara Merchawi, Soundar R. T. Kumara, Chita R. Das |
IEEE Trans. Neural Networks | 3 |
| 1996 | Performance Analysis of Finite-Buffered Asynchronous Multistage Interconnection NetworksabstractWe present a queueing model for performance analysis of finite-buffered multistage interconnection networks. The proposed model captures network behaviour in an asynchronous communication mode and is based on realistic assumptions. A uniform traffic model is developed first and then extended to capture nonuniform traffic in the presence of a hot-spot. Throughput and delay are computed using the proposed model and the results are validated via simulation. The analysis is extended to predict performance of MIN-based multiprocessors. The effects of buffer length, switch size, and the maximum allowable outstanding requests on the system performance are discussed. Various design decisions using this model are drawn with respect to delay, throughput, and system power. Prasant Mohapatra, Chita R. Das |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1995 | Modeling Virtual Channel Flow Control in HypercubesabstractAn analytical model for virtual channel flow control in n-dimensional hypercubes using the e-cube routing algorithm is developed. The model is based on determining the values of the different components that make up the average message latency. These components include the message transfer time, the blocking delay at each dimension, the multiplexing delay at each dimension, and the waiting delay at the source node. The first two components are determined using a probabilistic analysis. The average degree of multiplexing is determined using a Markov model, and the waiting delay at the source node is determined using an M/M/m queueing system. The model is fairly accurate in predicting the average message latency for different message sizes and a varying number of virtual channels per physical channel. It is demonstrated that wormhole switching along with virtual channel flow control make the average message latency insensitive to the network size when the network is relatively lightly loaded (message arrival rate is equal to 40% of channel capacity), and that the average message latency increases linearly with the average message size. The simplicity and accuracy of the analytical model make it an attractive and effective tool for predicting the behavior of n-dimensional hypercubes.> Younes M. Boura, Chita R. Das |
HPCA | 2 |
| 1995 | Fault-Tolerant Routing in Mesh Networks
Younes M. Boura, Chita R. Das |
ICPP (1) | 2 |
| 1995 | Processor Management Techniques for Mesh-Connected Multiprocessors
Byung S. Yoo, Chita R. Das, Chansu Yu |
ICPP (2) | 2 |
| 1995 | Experimenting with a Shared Virtual Memory Environment for Hypercubes
Amit Agarwala, Chita R. Das |
J. Parallel Distributed Comput. | 2 |
| 1995 | A Lazy Scheduling Scheme for Hypercube Computers
Prasant Mohapatra, Chansu Yu, Chita R. Das |
J. Parallel Distributed Comput. | 3 |
| 1995 | On Dependability Evaluation of Mesh-Connected ProcessorsabstractAnalytical techniques for reliability and availability prediction of mesh-connected systems are proposed. The models are based on the submesh requirements. First, a reliability model is proposed assuming that a submesh can be always recognized if it exits. Analysis of the linear consecutive n-out-of-N system is extended using an expanding row/column technique to evaluate the submesh reliability. An alternative approach called row folding is also discussed. Due to the high complexity involved in computing the exact reliability, both of these techniques use approximation to estimate lower bounds. Next, the submesh reliability is computed based on two different allocation policies, known as the two-dimensional buddy system (TDBS), and the frame sliding (FS). The model with the TDBS is further extended to estimate the reliability of multiple working submeshes, which is useful in a multiuser environment. Availability analysis for a submesh of the required size is conducted using a Markov chain (MC). State truncation is used to reduce the computation time, and the MC is solved using a software package called HARP. Validation of the analytical models is done through extensive simulation. Issues, such as reliability comparison based on allocation policies, and methods for improving system reliability are addressed using the analytical models.> Prasant Mohapatra, Chita R. Das |
IEEE Trans. Computers | 2 |
| 1995 | Distributed Fault Diagnosis in Multistage Network-Based MultiprocessorsabstractThis paper is concerned with a distributed, system level fault diagnosis scheme for multistage network-based multiprocessors. The target system, which we choose as a representative, employs a multistage interconnection network with 4/spl times/4 switching elements. We propose a fast diagnostic method which uses a quadtree and its coupler structure. These two quadtree structures partition the system into a number of link-independent groups. This partitioning provides an important diagnostic property; the communication paths in each link-independent group are either identical or disjoint. Several previous works in fault diagnosis investigated the multistage interconnection network only. This paper presents an entire multiprocessor diagnosis, including the detection and location of single faults caused by processor nodes, switching elements, and communication links. In addition, the diagnosis of a group of multiple faults partitioned by the tree structures is also discussed.> Tsang-Ling Sheu, Woei Lin, Chita R. Das |
IEEE Trans. Computers | 3 |
| 1995 | Disjoint Task Allocation Algorithms for MIN Machines with Minimal ConflictsabstractThis paper addresses task allocation schemes for MIN-based multiprocessors. Two types of allocation policies, cubic and noncubic, are discussed here. Conflicts through the network and inability to partition the system effectively are the main bottlenecks in a MIN-based system. To solve both the problems, a renaming scheme for input and output ports of a MIN is proposed. We use the baseline MIN as an example in this work and call the renaming scheme as bit reversal (BR) matching pattern. Allocation with the new matching pattern minimizes conflicts and partitions the system completely into independent subsystems. The novelty of this matching pattern is that we can use any dynamic cubic allocation and/or scheduling scheme developed for the hypercubes also for the MIN machines. The BR matching pattern can be used with any kind of MIN. An allocation policy for noncubic tasks is also presented with this matching pattern. Various performance measures with different allocation algorithms are compared via simulation. The advantages of the algorithms with the proposed matching pattern are shown in terms of system efficiency, delay and task miss ratio.> Chansu Yu, Chita R. Das |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1994 | Efficient Fully Adaptive Wormhole Routing in n-Dimensional MeshesabstractAn efficient fully adaptive wormhole routing algorithm for n-dimensional meshes is developed. The routing algorithm provides full adaptivity at a cost of one additional virtual channel per physical channel irrespective of the number of dimensions of the network. The algorithm is based on dividing the network graph into two acyclic graphs that contain all of the physical channels in the system. Virtual channels are classified as either waiting or nonwaiting channels. Busy channels that a message waits for to become available are classified as waiting channels, otherwise they are classified as nonwaiting channels. Thus, a message considers nonwaiting channels first to reach its destination. If all non-waiting channels are busy, the message considers waiting channels. Messages acquire waiting channels in two phases. In each phase, waiting channels belonging to one acyclic network graph are traversed. This 2-phase routing algorithm could be either minimal or nonminimal. However, we concentrate on minimal routing. It is demonstrated that this adaptive routing algorithm can utilize the virtual paths (channels) between any two nodes more efficiently than any of the present algorithms with the same hardware requirement.> Younes M. Boura, Chita R. Das |
ICDCS | 2 |
| 1994 | A Shared Memory Environment for HypercubesabstractThis paper describes the design and implementation of a shared virtual memory (SVM) system for the nCUBE 2 hypercube multicomputer. The SVM system provides the user a single coherent address space across all nodes. It is implemented at the user level in a C programming environment using high level constructs to support data sharing. Shared variables are treated as objects rather than pages. We have improved upon an existing algorithm for maintaining coherency in the SVM system, thus achieving a reduction in the number of inter-node messages required in coherency maintenance. Detailed timing analysis is conducted to analyze the feasibility of this shared environment. Experimental results indicate that parallel programs running under an SVM system show linear speedup, suggesting that SVM systems could provide an effective programming environment for the next generation of distributed memory parallel computers. A bottleneck of this implementation seems to be the expensive interrupt handling by the nCUBE 2 kernel. Amit Agarwala, Chita R. Das |
ICPP (1) | 2 |
| 1994 | Performance Analysis of Combining Multistage Interconnection NetworksabstractConcurrent access to a shared variable may cause network saturation in parallel computers. This problem, commonly termed as hot spot contention, can be alliviated by combining requests destined to the hot memory module. In this paper, we propose an analytical model to predict performance of combining multistage interconnection networks. The model considers realistic assumptions like finite length buffers in the switches, deterministic service time, finite degree of combining. Simulation results are used to validate the analytical model. Prasant Mohapatra, Sheldon Wong, Chita R. Das |
ICPP (1) | 3 |
| 1994 | Limit Allocation: An Efficient Processor Management Scheme for HypercubesabstractEfficient task management in a hypercube multi-processor becomes difficult due to system overflow, where an incoming job cannot be allocated in spite of a sufficient number of free processors. Overflow occurs either due to the inability of recognizing a free subcube or due to external fragmentation. In this paper, we propose an allocation strategy that tries to scale down an incoming job size if it cannot fit into a fragmented hypercube. We call it limit allocation. We discuss three simple schemes, Limit-k, Greedy and Average. We conduct both analysis and simulation to characterize and compare various allocation policies. An M/M/m queueing model is developed to predict the behavior of buddy, free list and limit-k policies. The simulation study shows that the two adaptive schemes, greedy and average, outperform all other schemes reported so far in the literature. Chansu Yu, Chita R. Das |
ICPP (2) | 2 |
| 1994 | Hypercube Communication Delay with Wormhole RoutingabstractWe present an analytical model for the performance evaluation of hypercube computers. This analysis is aimed at modeling a deadlock-free wormhole routing scheme prevalent on second generation hypercube systems. Probability of blocking and average message delay are the two performance measures discussed. We start with the communication traffic to find the probability of blocking. The traffic analysis can capture any message destination distribution. Next, we find the average message delay that consists of two parts. The first part is the actual message transfer delay between any source and destination nodes. The second part of the delay is due to blocking caused by the wormhole routing scheme. The analysis is also extended to virtual cut-through routing and random wormhole routing techniques. The validity of the model is demonstrated by comparing analytical results with those from simulation.> Jong Kim 0001, Chita R. Das |
IEEE Trans. Computers | 2 |
| 1994 | Performance Analysis of Cluster-Based MultiprocessorsabstractA queueing model for performance evaluation of cluster-based multiprocessors is proposed. Most system components are modeled as M/D/1/L queues to capture deterministic service time and finite buffer behavior. Various subsystems are analyzed independently and then integrated for the system level analysis. Average delay, throughput, and processor utilization are the performance parameters studied in this analysis. The analytical results are first validated via simulation. Next, several design alternatives are discussed using the model. These include the effect of buffer length and identification of bottleneck centers for various design configurations.> Prasant Mohapatra, Chita R. Das, Tse-Yun Feng |
IEEE Trans. Computers | 2 |
| 1994 | A Cache coherence protocol for MIN-based multiprocessors
Mazin S. Yousif, Chita R. Das, Matthew J. Thazhuthaveetil |
J. Supercomput. | 2 |
| 1994 | Evaluation of a Parallel Branch-and-Bound Algorithm on a Class of MultiprocessorsabstractWe propose and evaluate a parallel "decomposite best-first" search branch-and-bound algorithm (dbs) for MIN-based multiprocessor systems. We start with a new probabilistic model to estimate the number of evaluated nodes for a serial best-first search branch-and-bound algorithm. This analysis is used in predicting the parallel algorithm speed-up. The proposed algorithm initially decomposes a problem into N subproblems, where N is the number of processors available in a multiprocessor. Afterwards, each processor executes the serial best-first search to find a local feasible solution. Local solutions are broadcasted through the network to compute the final solution. A conflict-free mapping scheme, known as the step-by-step spread, is used for subproblem distribution on the MIN. A speedup expression for the parallel algorithm is then derived using the serial best-first search node evaluation model. Our analysis considers both computation and communication overheads for providing realistic speed-up. Communication modeling is also extended for the parallel global best-first search technique. All the analytical results are validated via simulation. For large systems, when communication overhead is taken into consideration, it is observed that the parallel decomposite best-first search algorithm provides better speed-up compared to other reported schemes.> Myung K. Yang, Chita R. Das |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1993 | A Class of Partially Adaptive Routing Algorithms for n_dimensional MeshesabstractA simple model, called the direction restriction model, for developing partially adaptive routing algorithms for n_dimensional meshes is introduced in this paper. This model is based on dividing a system into two unidirectional networks that contain all physical channels of the system. Younes M. Boura, Chita R. Das |
ICPP (3) | 2 |
| 1993 | A Queuing Model for Finite-Buffered Multistage Interconnection NetworksabstractIn this paper, we present a queueing model for per formance analysis of finite-buffered multistage inter connection networks. The model captures network be havior in an asynchronous communication mode and is based on realistic assumptions. Throughput and de lay are computed using the proposed model and the results are validated via simulation. Various design decisions using this mode! are drawn with respect to delay, throughput, and system power. Prasant Mohapatra, Chita R. Das |
ICPP (1) | 2 |
| 1993 | A Lazy Scheduling Scheme for Improving Hypercube PerformanceabstractProcessor allocation and job scheduling are com plementary techniques to improve the performance of multiprocessors. It has been observed that all the hypercube allocation policies with the FCFS schedul ing show little performance difference. A greater im pact on the performance can be obtained by efficient job scheduling. This paper presents an effort in that direction by introducing a new scheduling algorithm called lazy scheduling for hypercubes. The motivation of this scheme is to eliminate the limitations of the FCFS scheduling. This is done by maintaining sep arate queues for different job sizes and delaying the allocation of a job if any other job(s) of the same di mension is(are) running in the system. Simulation studies show that the hypercube performance is dra matically enhanced by using the lazy scheme as com pared to the FCFS scheduling. Comparison with a re cently proposed scheme called scan indicates that the lazy scheme performs better than scan under a wide range of workloads. Prasant Mohapatra, Chansu Yu, Chita R. Das, Jong Kim 0001 |
ICPP (1) | 3 |
| 1993 | A Cache Coherence Protocol for MIN-Based Multprocessors With Limited InclusionabstractIn this paper, we look into a feasible approach to incorporating caches into selected switching ele ments of a multistage interconnection network (MIN)- based multiprocessor. Along with the processor private caches, these switch caches form a two-level cache hi erarchy. Selected switch caches within a particular stage of the MIN are connected by a coherence con trol bus, through which a write-invalidate cache coher ence protocol is maintained. Considering scalability and practicality issues, only limited inclusion between the two cache levels is enforced. A simulation-based performance study is conducted to analyze the impact of the protocol on system performance. Comparison between limited and strict inclusion shows that system performance declines with limited inclusion. Mazin S. Yousif, Chita R. Das, Matthew J. Thazhuthaveetil |
ICPP (1) | 2 |
| 1993 | An Availability Model for MIN-Based MultiprocessorsabstractSystem decomposition is a novel technique for modeling the dependability of complex systems without constructing a single-level Markov Chain (MC). This is demonstrated in this paper for the availability computation of a class of multiprocessors that uses 4*4 switching elements for the multistage interconnection network (MIN). The availability model is known as task-based availability, where a system is considered operational as long as the task requirements are satisfied. The authors develop two simple MC's for the processors and memories and solve them using a software package, called HARP. The probabilities of i processing elements (PE's) and j memory modules (MM's) working at any time t, denoted as Pi(t) and Pj(t), are obtained from their corresponding MC's. The effect of the MIN is captured in the model by finding the number of switches required for the connection of i PE's and j MM's. A third MC is then developed for the switches to find the probability that the MIN provides the required (i*j) connection. Multiplying this term with Pi(t) and Pj(t), the probability of an (i*j) working group is obtained. The methodology is generalized to model arbitrary as well as larger size systems. Transient and steady state availabilities are computed for a variety of MIN configurations and the results are validated through simulation.> Chita R. Das, Prasant Mohapatra, Lei Tien, Laxmi N. Bhuyan |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1992 | Multitasking in Multistage Interconnection Network MachinesabstractCubic and noncubic task allocation algorithms for multistage-interconnection-network (MIN)-based multiprocessors are presented. Conflicts in passage through the network and inability to partition the system effectively are the main bottlenecks in a MIN-based system. To solve both problems, a renaming scheme called bit reversal (BR) matching pattern is proposed. This matching pattern minimizes conflicts and partitions the system completely into subsystems. Simulation results that show the advantage of allocation algorithms using the proposed matching pattern in terms of system efficiency, delay, and task miss ratio are presented.> Chansu Yu, Chita R. Das |
ICDCS | 2 |
| 1992 | A Unified Task-Based Dependability Model for Hypercube ComputersabstractA unified analytical model for computing the task-based dependability (TDB) of hypercube architectures is presented. A hypercube is deemed operational as long as a task can be executed on the system. The technique can compute both reliability and availability for two types of task requirements-I-connected model and subcube model. The I-connected TBD assumes that a connected group of at least I working nodes is required for task execution. The subcube TBD needs at least an m-cube in an n-cube, mor=I or x>or=2/sup m/) are working in an n-cube at time t by the conditional probability that the hypercube can satisfy any one of the two task requirements from x working nodes. Recursive models are proposed for the two types of task requirements to find the connection probability. The subcube requirement is extended to find multiple subcubes for analyzing multitask dependability. The analytical results are validated through extensive simulation.> Chita R. Das, Jong Kim 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1991 | Modeling wormhole routing in a hypercubeabstractAn analytical model for the performance evaluation of asynchronous hypercubes is presented. This analysis is aimed at modeling a deadlock-free wormhole routing scheme prevalent on second-generation hypercube systems. Probability of blocking and average message delay are discussed. The communication traffic to find the probability of blocking is the starting point. The traffic analysis can capture any message destination distribution. The average message delay that consists of two parts is found. The analysis is extended to virtual cut-through routing and random wormhole routing techniques. The validity of the model is demonstrated.> Jong Kim 0001, Chita R. Das |
ICDCS | 2 |
| 1991 | A Cache-Based Checkpointing Scheme for MIN-Based Multiprocessors
Mazin S. Algudady, Chita R. Das, Matthew J. Thazhuthaveetil |
ICPP (1) | 2 |
| 1991 | On Subcube Dependability in a HypercubeabstractIn this paper, we present an analytical model for computing the dependability of hypercube systems. The model, referred to as task-based dependability (TBD), is developed under the assumption that a task needs at least an m-cube (m < n) in an n-cube for its execution. Two probabilistic terms are required for computing this dependability. The first is the probability of any x nodes working out of 2n nodes. The second term is a conditional probability that at least a connected m-cube exists among those x working nodes. This term is computed using a recursive expression. Two dependability measures, reliability and availability, are analyzed in this paper. A combinatorial enumeration is used in the reliability analysis, and a machine repairman model is used in the availability analysis to find the first probability. The machine repairman model is modified to capture imperfect coverage and imprecise repair. The TBD model is also extended to find multitask dependability. Numerical results are presented for n-cubes with different task requirements and are validated through extensive simulation. It is observed that an m-cube requirement is highly restrictive compared to the simple 2m-connected node requirement. Jong Kim 0001, Chita R. Das |
SIGMETRICS | 2 |
| 1991 | A Parallel Branch-and Bound Algorithm for MIN-Based MultiprocessorsabstractA parallel "Decomposite Best-First" search Branch-and-Bound algorithm (pdbsbb) for MIN-based multiprocessor systems is proposed in this paper. A conflict free mapping scheme, known as step-by-step spread, is used to map the algorithm efficiently on to a MIN-based system for reducing communication overhead. It is shown that the proposed algorithm provides better speed-up than other reported schemes when communication overhead is taken into consideration. Myung K. Yang, Chita R. Das |
SIGMETRICS | 2 |
| 1991 | A Top-Down Processor Allocation Scheme for Hypercube ComputersabstractAn efficient processor allocation policy is presented for hypercube computers. The allocation policy is called free list since it maintains a list of free subcubes available in the system. An incoming request of dimension k (2/sup k/ nodes) is allocated by finding a free subcube of dimension k or by decomposing an available subcube of dimension greater than k. This free list policy uses a top-down allocation rule in contrast to the bottom-up approach used by the previous bit-map allocation algorithms. This allocation scheme is compared to the buddy, gray code (GC), and modified buddy allocation policies reported for the hypercubes. It is shown that the free list policy is optimal in a static environment, as are the other policies, and it also gives better subcube recognition ability compared to the previous schemes in a dynamic environment. The performance of this policy, in terms of parameters such as average delay, system utilization, and time complexity, is compared to the other schemes to demonstrate its effectiveness. The extension of the algorithm for parallel implementation, noncubic allocation, and inclusion/exclusion allocation is also given.> Jong Kim 0001, Chita R. Das, Woei Lin |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1990 | Fault-Tolerant Task Mapping Algorithms for MIN-Based Multiprocessors
Mazin S. Algudady, Chita R. Das, Woei Lin |
ICPP (1) | 2 |
| 1990 | A write update cache coherence protocol for MIN-based multiprocessors with accessibility-based split cachesabstractThe authors present a cache coherence protocol for MIN-based multiprocessors with two distinct private caches: private-block caches containing information private to a processor and shared-block caches containing data accessible by all processors. The protocol utilizes a coherence control bus (snooping) for connecting all shared-block cache controllers. Timing problems due to variable transit delay through the MIN are dealt with by introducing transient states in the protocol. Assuming homogeneity of all nodes, a single-node queuing model is developed to analyze the system performance. This model is solved using the mean-value-analysis technique with protocol state probabilities, and few communication delays as input parameters. System performance measures are verified through simulation.> Mazin S. Algudady, Chita R. Das, Matthew J. Thazhuthaveetil |
SC | 2 |
| 1989 | A Processor Allocation Scheme for Hypercube Computers
Jong Kim 0001, Chita R. Das, Woei Lin |
ICPP (2) | 2 |
| 1989 | Distributed Fault Diagnosis in the Butterfly Parallel Processor
Tsang-Ling Sheu, Woei Lin, Chita R. Das, Mary Jane Irwin |
ICPP (1) | 3 |
| 1989 | A Conflict-Free Routing Scheme on Multistage Interconnection NetworksabstractA conflict-free routing scheme is presented for a class of parallel and distributed computing systems. The core of the scheme is a quadtree communication structure. The quadtree structure suggests a general approach to mapping a class of parallel algorithms with intensive communication requirements for selecting data from many different sources and distributing data from a single source. By properly merging messages and efficiently replicating data, the quadtree structure can complete required communications in O(log/sub 4/ M) parallel steps, where M is the network size. It is shown that the size of a quadtree communication structure can be contracted and stretched by adjusting the number of descendent nodes without affecting its conflict-free property. The relationship between the computation/communication ratio of various parallel algorithms and the number of tree levels is presented, and finally, their joint effect on the response time of combining and distributing data messages is examined. This analysis helps determine the optimal adaptation of the quadtree for minimizing the overall algorithm execution time.> Woei Lin, Tsang-Ling Sheu, Chita R. Das, Tse-Yun Feng, Chuan-lin Wu |
IEEE Trans. Computers | 3 |
| 1988 | A quadtree communication structure for fast data searching and distributionabstractPresents a quadtree communication structure and two associated procedures for efficient, contention-free data searching and distribution on the BBN Butterfly parallel processor and its family. The proposed quadtree structure suggests a general approach to mapping a class of parallel algorithms with intensive communication requirements for performing two primitive operations: selecting data from many different sources and distributing data from a single source. While performing these two operations through the quadtree structure, the 'ascend' and 'descend' procedures incur no link conflicts in the Butterfly network. A concrete example of mapping the linear programming algorithm is given to show the effectiveness of the proposed quadtree communication structure.> Woei Lin, Tsang-Ling Sheu, Chita R. Das |
COMPSAC | 3 |
| 1988 | A Reliability Predictor for MIN-connected Multiprocessor Systems
John J. Macaluso, Chita R. Das, Woei Lin |
ICPP (1) | 2 |
| 1987 | Dependability evaluation of interconnection networks
Chita R. Das, Laxmi N. Bhuyan |
Inf. Sci. | 1 |
| 1986 | Dependability Evaluation of Multicomputer Networks
Laxmi N. Bhuyan, Chita R. Das |
ICPP | 2 |
| 1985 | Reliability Simulation of Multiprocessor Systems
Chita R. Das, Laxmi N. Bhuyan |
ICPP | 1 |
| 1985 | Computation Availability of Multiple-Bus Multiprocessors
Chita R. Das, Laxmi N. Bhuyan |
ICPP | 1 |
| 1985 | Bandwidth Availability of Multiple-Bus MultiprocessorsabstractThe effect of failures on the performance of multiple-bus multiprocessors is considered. Bandwidth expressions for this architecture are derived for uniform and nonuniform memory references. Mathematical models are developed to compute the reliability and the performance-related bandwidth availability (BA). The results obtained for the multiple-bus interconnection are compared with those of a crossbar. The models are also extended to analyze the partial bus structure, where the memories are divided into groups and each group is connected to a subset of buses. The reliability and the BA of the multiple-bus and partial bus architectures are compared. Chita R. Das, Laxmi N. Bhuyan |
IEEE Trans. Computers | 1 |