EDBT 2026 Demo / reviewers in the wild / expert
Smruti R. Sarangi
dblp:s/SmrutiRSarangi · also Smruti Ranjan Sarangi, Smruti Sarangi
· DBLP profile ↗
83ranked-venue papers
6as first author
30since 2021 · last 2026
0000-0002-1657-8523ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 63 · 5 first-author · 16 since 2021Software engineering, systems software and programming languages · 16 · 6 since 2021Security and privacy · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorComputer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | StealthDev: Side-Channel-Resistant Forensic Framework for Investigating Websites with Anti-DebuggingabstractThe analysis of JavaScript within websites is a foundational step in web security analysis. One of the most powerful tools available for such an analysis is the browser's integrated developer tools (DevTools), which provide deep visibility into the client-side code execution. However, sophisticated malicious websites often employ anti-debugging techniques to detect the presence of such tools and alter their behavior to appear benign. Prior research into such polymorphic sites has focused mainly on isolated client-side detection vectors, those that operate independently within the browser without requiring any coordination from the server. In this work, we argue that collusive vectors, which involve explicit communication between the client and an attacker-controlled server, are common, highly effective, and stealthy. These collusive client-attacker techniques evade detection by minimizing observable client-side artifacts, making them particularly challenging to analyze. Furthermore, despite the growing prevalence of such tactics, the security community still lacks a comprehensive framework to counter these anti-debugging tactics effectively.In this paper, we examine the prevalence of isolated and collusive anti-debugging vectors and then develop a solution that enables stealthy analysis of websites. First, we introduce a novel class of anti-debugging techniques in which adversaries exploit debugger-related resource requests, such as those for source maps or debugging symbols, initiated by the browser to detect DevTools. This mechanism allows attackers to detect the presence of active debuggers with a minimal client-side footprint. Through a large-scale analysis of the top 100,000 websites and their subpages, we find that approximately 1 in 17 websites employing severe anti-debugging measures leverage such collusive vectors to identify and obstruct DevTools-based analysis. Second, to defend against these threats, we present StealthDev, a policy-driven debugging framework built on a modified Chromium browser. StealthDev is specifically designed to remain stealthy and resilient against isolated and collusive anti-debugging techniques, including those relying on timing side channels. Our evaluation shows that StealthDev successfully bypasses anti-debugging defenses on over 97% of the most evasive websites, while incurring a modest performance overhead of just 5% on average page load times. Rahul Kanyal, Smruti R. Sarangi |
AsiaCCS | 2 |
| 2026 | PatchEX: High-Quality Real-Time Temporal Supersampling through Patch-based Parallel ExtrapolationabstractHigh-refresh rate displays have become very popular in recent years due to the need for superior visual quality in gaming, professional displays and specialized applications such as medical imaging. However, high-refresh rate displays alone do not guarantee a superior visual experience; the GPU needs to render frames at a matching rate. Otherwise, we observe disconcerting visual artifacts such as screen tearing and stuttering. Real-time frame generation is an effective technique to increase frame rates by predicting new frames from other rendered frames. There are two methods in this space: interpolation and extrapolation. Interpolation-based methods provide good image quality at the cost of a higher runtime because they also require the next rendered frame. On the other hand, extrapolation methods are much faster at the cost of quality. This article introduces PatchEX , a novel frame extrapolation method that aims to provide the quality of interpolation at the speed of extrapolation. It smartly segments each frame into foreground and background regions and employs a novel neural network to generate the final extrapolated frame. Additionally, a wavelet transform (WT)-based filter pruning technique is applied to compress the network, significantly reducing the runtime of the extrapolation process. Our results demonstrate that PatchEX achieves a 61.32% and 49.21% improvement in PSNR over the latest extrapolation methods ExtraNet and ExtraSS, respectively, while being 3× and 2.6× faster, respectively. Akanksha Dixit 0002, Smruti R. Sarangi |
ACM Trans. Graph. | 2 |
| 2026 | Styx: An Efficient Workflow Engine for Serverless PlatformsabstractServerless platforms are widely adopted for deploying applications due to their autoscaling capabilities and pay-asyou- go billing models. These platforms execute an application's functions inside ephemeral containers and scale the number of containers based on incoming request rates. To meet service level objectives (SLOs), they often over-provision resources by maintaining warm containers or rapidly spawning new ones during traffic bursts. However, this strategy frequently leads to inefficient resource utilization, especially during periods of low activity. Prior research addresses this issue through intelligent scheduling, lightweight virtualization, and containersharing mechanisms. More recent work aims to improve resource utilization by remodeling the execution of a function within a container to better separate compute and I/O stages. Despite these improvements, existing approaches often introduce delays during execution and induce memory pressure under traffic bursts. In this paper, we present Styx, a novel workflow engine that enhances resource utilization by intelligently decoupling compute and I/O stages. Styx employs a fetch latency predictor that uses real-time system metrics from both the serverless node and the remote storage server to accurately estimate prefetch operations, ensuring input data is available exactly when needed. Furthermore, it offloads the output data upload operation from a container to a host-side data service, thereby efficiently managing provisioned memory. Our approach improves the overall memory allocation by 32.6% when running all the serverless workflows simultaneously when compared to Dataflower + Truffle. Additionally, this method improves the tail latency and the mean latency of a workflow by an average of 26.3% and 21%, respectively. Abhisek Panda, Smruti R. Sarangi |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | FaaSImage: An Efficient Image Manager for FaaSabstractThe cold start latency in serverless systems is a matter of great concern. It militates against its basic foundation, which is fast millisecond-level execution of mostly stateless functions. Over the last five years, a lot of work has been done in academia and industry to mitigate the overheads caused by long cold start times. In this paper, we focus on a specific line of work that proposes to modify the Docker container's architecture to address this problem. We observe that state-of-the-art work has either fused Docker layers or used smart on-demand fetching of data. Abhisek Panda, Smruti R. Sarangi |
Middleware | 2 |
| 2025 | SecScale : A Scalable and Secure Trusted Execution Environment for Servers
Ani Sunny, Nivedita Shrivastava, Smruti R. Sarangi |
J. Syst. Archit. | 3 |
| 2025 | VoxDepth: Rectification of Depth Images on Edge DevicesabstractAutonomous mobile robots like self-flying drones and industrial robots heavily depend on depth images to perform tasks such as 3D reconstruction and visual SLAM. However, the presence of inaccuracies in these depth images can greatly hinder the effectiveness of these applications, resulting in sub-optimal results. Depth images produced by commercially available cameras frequently exhibit noise, which manifests as flickering pixels and erroneous patches. Machine Learning (ML)-based methods to rectify these images are unsuitable for edge devices that have very limited computational resources. Non-ML methods are much faster but have limited accuracy, especially for correcting errors that are a result of occlusion and camera movement. We propose a scheme called VoxDepth that is fast, accurate, and runs very well on edge devices such as the NVIDIA Jetson Nano board. It relies on a host of novel techniques: 3D point cloud construction and fusion, and using it to create a 2D template to fix erroneous depth images. VoxDepth shows superior results on both synthetic and real-world datasets. We specifically demonstrate a 31% improvement in quality as compared with state-of-the-art methods on real-world depth datasets, while maintaining a competitive frame rate of 27 FPS (frames per second). Yashashwee Chakrabarty, Akanksha Dixit 0002, Smruti R. Sarangi |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2025 | WaWoR: Wasted Work Reduction During Snapshotting in EH-WSNsabstractEnergy harvesting wireless sensor networks (EHWSNs) are useful for ambient monitoring, especially in hazardous and hard-to-reach environments, where sensor nodes sense and transmit data to a remote sink via multi-hop communication, enabling informed decision making. Accurate decisions require a global snapshot that comprises environmental parameters simultaneously sensed from all nodes. To realize an efficient snapshot collection scheme, all nodes should send an equal number of messages where each message contributes to the snapshot; excessive messages from some nodes are ineffective and in-turn increase the network traffic and contribute to the wasteful work done by nodes. Reducing such ineffectual messages is challenging due to the variable ambient energy supply at nodes, network congestion, limited information about the overall state of the system and varying node distances from the sink. To achieve this objective, we introduce WaWoR, a novel distributed system where nodes dynamically decide when to sense and transmit snapshot messages based on their physical locations, energy availability and perceived network congestion. WaWoR outperforms two theoretical hypothetical baselines and two state-of-the-art-systems. It captures [90−175]% of the snapshots compared to the theoretical baselines and [1.02 − 4.26]× more snapshots than state-of-the-art systems. Additionally, it reduces wasted messages, which further reduces the total energy consumption by [1.6 − 9.9]×. Priyanka Singla 0001, Smruti R. Sarangi |
IEEE Trans. Sustain. Comput. | 2 |
| 2024 | HybMT: Hybrid Meta-Predictor based ML Algorithm for Fast Test Vector GenerationabstractML models are increasingly being used to increase the test coverage and decrease the overall testing time. This field is still in its nascent stage and up till now there were no algorithms that could match or outperform commercial tools in terms of speed and accuracy for large circuits. We propose an ATPG algorithm HybMT in this paper that finally breaks this barrier Like sister methods, we augment the classical PODEM algorithm that uses recursive backtracking. We design a custom 2-level predictor that predicts the input net of a logic gate whose value needs to be set to ensure that the output is a given value (0 or 1). Our predictor chooses the output from among two first-level predictors, where the most effective one is a bespoke neural network and the other is an SVM regressor. As compared to a popular, state-of-the-art commercial ATPG tool, HybMT shows an overall reduction of 56.6% in the CPU time without compromising on the fault coverage for the EPFL benchmark circuits. HybMT also shows a speedup of 126.4% over the best ML-based algorithm while obtaining an equal or better fault coverage for the EPFL benchmark circuits. Shruti Pandey, Jayadeva, Smruti R. Sarangi |
ASPDAC | 3 |
| 2024 | Semantic-Aided Image Transmission System with Unequal Error Protection for Next-Generation Communication NetworksabstractSemantic communication (SC) aims to convey the meaning of data instead of focusing on its bit-by-bit reconstruction. SC finds applications in beyond 5G and 6G networks for artificial intelligence-empowered multimedia content delivery. In this paper, we propose a novel semantic-aided autoencoder-based image transmission system that leverages semantic information in the form of the segmentation map of an image. We demonstrate up to 23% and 18% improvement (in terms of mean square error and peak signal-to-noise ratio, respectively) in the quality of the received image with only 2% extra bandwidth over a traditional autoencoder-based image transmission system. The study also explores channel coding strategies for our proposed system. We focus on the intrinsically robust nature of semantic data, as compared to traditional data, to design low-density parity check code, Hamming code, and polar code-based unequal error protection (UEP) schemes. Comparative evaluations between UEP and equal error protection schemes show that while both approaches yield similar performance, UEP schemes are more efficient. Nargis Fayaz, Aman Shreshtha, Smruti R. Sarangi, Ranjan K. Mallik, Brejesh Lall |
WCNC | 3 |
| 2024 | PanoptiChrome: A Modern In-browser Taint Analysis FrameworkabstractTaint tracking in web browsers is a problem of profound interest because it allows developers to accurately understand the flow of sensitive data across JavaScript (JS) functions. Modern websites load JS functions from either the web server or other third-party sites, hence this problem has acquired a much more complex and pernicious dimension. Sadly, for the latest version of the Chromium browser (used by 75% of users), there is no dynamic taint propagation engine primarily because it is incredibly complex to build one. The nearest contending work in this space was published in 2018 for version 57; at the time of writing, we are at Chromium version 117, and the current version is very different from the 2018 version. We outline the details of a multi-year effort in this paper that led to PanoptiChrome, which accurately tracks information flow across an arbitrary number of sources and sinks and is, to a large extent, portable across platforms. As an example use case of the platform, we experimentally show that we can discover fingerprinting APIs that can uniquely identify the browser and sometimes the user, which are missed by state-of-the-art tools, owing to our comprehensive dynamic analysis methodology. For the top 20,000 most popular websites, we discovered a total of 362 APIs that have the potential to be used for fingerprinting -- out of these, 208 APIs were previously not reported by state-of-the-art tools. Rahul Kanyal, Smruti R. Sarangi |
WWW | 2 |
| 2024 | FaaSCtrl: A Comprehensive-Latency Controller for Serverless PlatformsabstractServerless computing systems have become very popular because of their natural advantages with respect to auto-scaling, load balancing and fast distributed processing. As of today, almost all serverless systems define two QoS classes: best-effort ($BE$) and latency-sensitive ($LS$). Systems typically do not offer any latency or QoS guarantees for$BE$jobs and run them on a best-effort basis. In contrast, systems strive to minimize the processing time for$LS$jobs. This work proposes a precise definition for these job classes and argues that we need to consider a bouquet of performance metrics for serverless applications, not just a single one. We thus propose the comprehensive latency ($CL$) that comprises the mean, tail latency, median and standard deviation of a series of invocations for a given serverless function. Next, we design a systemFaaSCtrl, whose main objective is to ensure that every component of the$CL$is within a prespecified limit for an LS application, and for BE applications, these components are minimized on a best-effort basis. Given the sheer complexity of the scheduling problem in a large multi-application setup, we use the method of surrogate functions in optimization theory to design a simpler optimization problem that relies on performance and fairness. We rigorously establish the relevance of these metrics through characterization studies. Instead of using standard approaches based on optimization theory, we use a much faster reinforcement learning (RL) based approach to tune the knobs that govern process scheduling in Linux, namely the real-time priority and the assigned number of cores. RL works well in this scenario because the benefit of a given optimization is probabilistic in nature, owing to the inherent complexity of the system. We show using rigorous experiments on a set of real-world workloads thatFaaSCtrlachieves its objectives for both LS and BE applications and outperforms the state-of-the-art by 36.9% (for tail response latency) and 44.6% (for response latency's std. dev.) for LS applications. Abhisek Panda, Smruti R. Sarangi |
IEEE Trans. Cloud Comput. | 2 |
| 2024 | PredATW: Predicting the Asynchronous Time Warp Latency For VR SystemsabstractWith the advent of low-power ultra-fast hardware and GPUs, virtual reality (VR) has gained a lot of prominence in the past few years and is being used in various areas, such as education, entertainment, scientific visualization, and computer-aided design. VR-based applications are highly interactive, and one of the most important performance metrics for these applications is the motion-to-photon-delay (MPD). MPD is the delay from the user’s head movement to the time at which the image gets updated on the VR screen. Since the human visual system can even detect an error of a few pixels (very spatially sensitive), the MPD should be as small as possible. Popular VR vendors use the GPU-accelerated Asynchronous Time Warp (ATW) algorithm to reduce the MPD. ATW reduces the MPD if and only if the warping operation finishes just before the display refreshes. However, due to the competition between the different constituent applications for the single, shared GPU, the GPU-accelerated ATW algorithm suffers from an unpredictable ATW latency, making it challenging to find the ideal time instance for starting the time warp and ensuring that it completes with the least amount of lag relative to the screen refresh. Hence, the state-of-the-art is to use a separate hardware unit for the time-warping operation. Our approach, PredATW , uses an ML-based hardware predictor to predict the ATW latency for a VR application, and then schedule it as late as possible while running the time-warping operation on the GPU itself. As far as we know, this is the first work to do so. Our predictor achieves an error of only 0.22 ms across several popular VR applications for predicting the ATW latency. As compared to the baseline architecture, we reduce deadline misses by 80.6%. Akanksha Dixit 0002, Smruti R. Sarangi |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2023 | Perspector: Benchmarking Benchmark SuitesabstractEstimating the quality of a benchmark suite is a non-trivial task. A poorly selected or improperly configured bench-mark suite can present a distorted picture of the performance of the evaluated framework. With computing venturing into new domains, the total number of benchmark suites available is increasing by the day. Researchers must evaluate these suites quickly and decisively for their effectiveness. We present Perspector, a novel tool to quantify the performance of a benchmark suite. Perspector comprises novel metrics to characterize the quality of a benchmark suite. It provides a math-ematical framework for capturing some qualitative suggestions and observations made in prior work. The metrics are generic and domain-agnostic. Furthermore, our tool can be used to compare the efficacy of one suite vis-a-vis other benchmark suites, systematically and rigorously create a suite of workloads, and appropriately tune them for a target system. Abhisek Panda, Smruti R. Sarangi |
DATE | 3 |
| 2023 | JASS: A Tunable Checkpointing System for NVM-Based SystemsabstractCheckpointing (or snapshotting) a system's state has always been a problem of great interest and has found a lot of use in ensuring system reliability, record-replay debugging, job migration and running high-throughput transaction systems. In the last few years ultra-fast hardware-assisted NVM-based checkpointing schemes have come up that can collect incremental full-system checkpoints in milliseconds. Unfortunately, such systems have large overheads in terms of their write amplification (increased number of writes). This, in turn, seriously reduces the reliability and lifetime of NVM devices. We propose the first tunable scheme in this space, JASS, where given a checkpoint latency (CL), we near-optimally minimize the write amplification (WA). This allows us to run parallel programs in a disciplined fashion. To realize this goal, we propose many novel hardware along the way such as a rigorous method of flushing pre-checkpoint messages in the NoC, a novel DRAM scrubber and locality predictor, and a control-theoretic algorithm to guarantee a CL while minimizing the WA. We reduce WA by 35-96% as compared to the nearest state-of-the-art competing method and improve performance of PARSEC benchmarks hv 19.4%. Akshin Singh, Smruti R. Sarangi |
HiPC | 2 |
| 2023 | Securator: A Fast and Secure Neural Processing UnitabstractSecuring deep neural networks (DNNs) is a problem of significant interest since an ML model incorporates high-quality intellectual property, features of data sets painstakingly collated by mechanical turks, and novel methods of training on large cluster computers. Sadly, attacks to extract model parameters are on the rise, and thus designers are being forced to create architectures for securing such models. State-of-the-art proposals in this field take the deterministic memory access patterns of such networks into cognizance (albeit partially), group a set of memory blocks into a tile, and maintain state at the level of tiles (to reduce storage space). For providing integrity guarantees (tamper avoidance), they don’t propose any significant optimizations, and still maintain block-level state.We observe that it is possible to exploit the deterministic memory access patterns of DNNs even further, and maintain state information for only the current tile and current layer, which may comprise a large number of tiles. This reduces the storage space, reduces the number of memory accesses, increases performance, and simplifies the design without sacrificing any security guarantees. The key techniques in our proposed accelerator architecture, Securator, are to encode memory access patterns to create a small HW-based tile version number generator for a given layer, and to store layer-level MACs. We completely eliminate the need for having a MAC cache and a tile version number store (as used in related work). We show that using intelligently-designed mathematical operations, these structures are not required. By reducing such overheads, we show a speedup of 20.56% over the closest competing work. Nivedita Shrivastava, Smruti R. Sarangi |
HPCA | 2 |
| 2023 | SnapStore: A Snapshot Storage System for Serverless SystemsabstractServerless computing is getting increasingly popular because of its fine-grained billing model and autoscaling features. To speed up the process of functions' sandbox creation, cloud providers typically utilize snapshot and restore-based mechanisms for pre-warmed snapshots. This effectively trades off the startup latency with the storage requirements and the overhead of creating/restoring these snapshots. Hence, there is a need to compress the snapshots by identifying identical data chunks across snapshots and then design methods to quickly deduplicate snapshots and retrieve them. We propose SnapStore -- a novel method of finding such duplicates. As opposed to conventional work that relies on better hashing methods, we use the natural structure of the program's memory map to reduce wasted work during deduplication. Furthermore, we sequentialize and minimize disk accesses as much as possible while retrieving a snapshot into a RAM-based cache. Both of these optimizations, yield a reasonably large speedup in the deduplication process as compared to the state-of-the-art (≈ 46% in the snapshot deduplication time and ≈ 82.6% in the retrieval time on HDDs). Upon integration with FaaSnap (a state-of-the-art serverless platform), SnapStore improves the end-to-end latency of serverless functions by 25.9% along with 2.4× storage space reduction over vanilla FaaSnap on HDDs. With SSDs, our deduplication time and retrieval time reduce by 36.2% and 75.8%, respectively, with almost no degradation in the end-to-end latency. Abhisek Panda, Smruti R. Sarangi |
Middleware | 2 |
| 2022 | Poster Abstract: Polar Code-based Approximate Communication System for Multimedia Web PagesabstractPolar codes have hitherto been used in the control plane of 5G-NR systems. However, in line with other contemporary works, we propose a novel use of them in the data plane by leveraging their natural property: different bit positions suffer from different degrees of errors. The idea is to map different components of web pages to different bit positions (based on their priority). We evaluate our approach for web page transmission over a wireless link that traditionally uses TCP and demonstrate benefits for image and video-based web pages. For an image-based web page, there is a 47.96% to 81.12% gain in performance, while the received image quality score varies between 0.97 and 0.99. We observe up to a 63.56% gain in performance for web pages with embedded videos with only a 7.45% loss in the received video quality. Aman Shreshtha, Priyanka Singla 0001, Smruti R. Sarangi |
IPSN | 3 |
| 2022 | SGXGauge: A Comprehensive Benchmark Suite for Intel SGXabstractTrusted execution environments (TEEs) such as Intel SGX facilitate the secure execution of an application on untrusted machines. A plethora of work focuses on improving the performance of such environments necessitating the need for a standard, widely accepted benchmark suite. We present SGXGauge, a benchmark suite for SGX containing a diverse set of workloads from different domains. We also thoroughly characterize the behavior of the benchmark suite on a native platform and on a platform that uses a library OS-based shim layer (GrapheneSGX). Abhisek Panda, Smruti R. Sarangi |
ISPASS | 3 |
| 2022 | SecureLease: Maintaining Execution Control in The Wild using Intel SGXabstractModern software programs have dedicated license-check modules that restrict access to users, who possess valid credentials. They also have a large number of add-on pluggable modules that can be separately purchased and have their dedicated license managers. Sadly, recent work shows that regardless of their complexity, it is possible to break their security using a novel class of attacks known as control flow bending attacks (CFB attacks), where the program is run on a virtual CPU, unbeknownst to it. Abhisek Panda, Smruti R. Sarangi |
Middleware | 3 |
| 2022 | PredStereo: An Accurate Real-time Stereo Vision SystemabstractStereo vision algorithms are important building blocks of self-driving applications. The two primary requirements of a self-driving vehicle are real-time operation and nearly 100% accuracy in constructing the 3D scene regardless of the weather conditions and the degree of ambient light. Sadly, most real-time systems as of today provide a level of accuracy that is inadequate and this endangers the life of the passengers; consequently, it is necessary to supplement such systems with expensive LiDAR-based sensors. We observe that for a given scene, different stereo matching algorithms can have vastly different accuracies, and among these algorithms there is no clear winner. This makes the case for a hybrid stereo vision system where the best stereo vision algorithm for a stereo image pair is chosen by a predictor dynamically, in real-time.We implement such a system called PredStereo in ASIC1that combines two diametrically different stereo vision algorithms, CNN-based and traditional, and chooses the best one at runtime. In addition, it associates a confidence with the chosen algorithm, such that the higher-level control system can be switched on in case of a low confidence value. We show that designing a predictor that is explainable and a system that respects soft real-time constraints is non-trivial. Hence, we propose a variety of hardware optimizations that enable our system to work in real-time. Overall, PredStereo improves the disparity estimation error over a state-of-the-art CNN-based stereo vision system by up to 18% (on average 6.25%) with a negligible area overhead (0.003mm2) while respecting real-time constraints. Diksha Moolchandani, Nivedita Shrivastava, Smruti R. Sarangi |
WACV | 4 |
| 2022 | Hardware-assisted mechanisms to enforce control flow integrity: A comprehensive survey
Diksha Moolchandani, Smruti R. Sarangi |
J. Syst. Archit. | 3 |
| 2022 | A survey and experimental analysis of checkpointing techniques for energy harvesting devices
Priyanka Singla 0001, Smruti R. Sarangi |
J. Syst. Archit. | 2 |
| 2022 | Performance and Power Prediction for Concurrent Execution on GPUsabstractThe unprecedented growth of edge computing and 5G has led to an increased offloading of mobile applications to cloud servers or edge cloudlets. 1 The most prominent workloads comprise computer vision applications. Conventional wisdom suggests that computer vision workloads perform significantly well on SIMD/SIMT architectures such as GPUs owing to the dominance of linear algebra kernels in their composition. In this work, we debunk this popular belief by performing a lot of experiments with the concurrent execution of these workloads, which is the most popular pattern in which these workloads are executed on cloud servers. We show that the performance of these applications on GPUs does not scale well with an increase in the number of concurrent applications primarily because of contention at the shared resources and lack of efficient virtualization techniques for GPUs. Hence, there is a need to accurately predict the performance and power of such ensemble workloads on a GPU. Sadly, most of the prior work in the area of performance/power prediction is for only a single application. To the best of our knowledge, we propose the first machine learning-based predictor to predict the performance and power of an ensemble of applications on a GPU. In this article, we show that by using the execution statistics of stand-alone workloads and the fairness of execution when these workloads are executed with three representative microbenchmarks, we can get a reasonably accurate prediction. This is the first such work in the direction of performance and power prediction for concurrent applications that does not rely on the features extracted from concurrent executions or GPU profiling data. Our predictors achieve an accuracy of 91% and 96% in estimating the performance and power of executing two applications concurrently, respectively. We also demonstrate a method to extend our models to four or five concurrently running applications on modern GPUs. Diksha Moolchandani, Smruti R. Sarangi |
ACM Trans. Archit. Code Optim. | 3 |
| 2022 | Game Theory-Based Parameter Tuning for Energy-Efficient Path Planning on Modern UAVsabstractPresent-day path planning algorithms for UAVs rely on various parameters that need to be tuned at runtime to be able to plan the best possible route. For example, for a sampling-based algorithm, the number of samples plays a crucial role. The dimension of the space that is being searched to plan the path, the minimum distance for extending a path in a direction, and the minimum distance that the drone should maintain with respect to obstacles while traversing the planned path are all important variables. Along with this, we have a choice of vision algorithms, their parameters, and platforms. Finding a suitable configuration for all these parameters at runtime is very challenging because we need to solve a complicated optimization problem, and that too within tens of milliseconds. The area of theoretical exploration of the optimization problems that arise in such settings is dominated by traditional approaches that use regular nonlinear optimization techniques often enhanced with AI-based techniques such as genetic algorithms. These techniques are sadly rather slow, have convergence issues, and are typically not suitable for use at runtime. In this article, we leverage recent and promising research results that propose to solve complex optimization problems by converting them into approximately equivalent game-theoretic problems. The computed equilibrium strategies can then be mapped to the optimal values of the tunable parameters. With simulation studies in virtual worlds, we show that our solutions are 5-21% better than those produced by traditional methods, and our approach is 10× faster. Diksha Moolchandani, Kishore Yadav, Geesara Kulathunga, Ilya Afanasyev 0001, Manuel Mazzara, Smruti R. Sarangi |
ACM Trans. Cyber Phys. Syst. | 7 |
| 2021 | EHDSktch: A Generic Low Power Architecture for Sketching in Energy Harvesting DevicesabstractEnergy harvesting devices (EHDs) are becoming extremely prevalent in remote and hazardous environments. They sense the ambient parameters and compute some statistics on them, which are then sent to a remote server. Due to the resource-constrained nature of EHDs, it is challenging to perform exact computations on streaming data; however, if we are willing to tolerate a slight amount of inaccuracy, we can leverage the power of sketching algorithms to provide quick answers with significantly lower energy consumption. Priyanka Singla 0001, Chandran Goodchild, Smruti R. Sarangi |
ASP-DAC | 3 |
| 2021 | Radiant: efficient page table management for tiered memory systemsabstractModern enterprise servers are increasingly embracing tiered memory systems with a combination of low latency DRAMs and large capacity but high latency non-volatile main memories (NVMMs) such as Intel’s Optane DC PMM. Prior works have focused on the efficient placement and migration of data on a tiered memory system, but have not studied the optimal placement of page tables. Aravinda Prasad, Smruti R. Sarangi, Sreenivas Subramoney |
ISMM | 3 |
| 2021 | SecureFS: A Secure File System for Intel SGXabstractA trusted execution environment or a TEE facilitates the secure execution of an application on a remote untrusted server. In a TEE, the confidentiality, integrity, and freshness properties for the code and data hold throughout the execution. In a TEE setting, specifically Intel SGX, even the operating system (OS) is not trusted. This results in certain limitations of a secure application’s functionality, such as no access to the file system and network – as it requires OS support. Smruti R. Sarangi |
RAID | 2 |
| 2021 | Accelerating CNN Inference on ASICs: A Survey
Diksha Moolchandani, Smruti R. Sarangi |
J. Syst. Archit. | 3 |
| 2021 | A survey of hardware architectures for generative adversarial networks
Nivedita Shrivastava, Muhammad Abdullah Hanif, Sparsh Mittal, Smruti R. Sarangi, Muhammad Shafique 0001 |
J. Syst. Archit. | 4 |
| 2021 | A Formal Approach to Accountability in Heterogeneous Systems-on-ChipabstractSystems-on-chip (SoCs) are increasingly being composed of designs provided by different organizations. When such an SoC miscomputes or performs below expectation in-field, it is unclear which of the on-chip components caused the failure. The customer would like to use SoCs that provide the property of accountability, wherein the failure-causing component, and consequently its designing organization, can be unambiguously detected. Since it is a matter of trust, the various parties involved desire formal guarantees regarding any accountability solution. The solution must find the guilty component(s) in the event of a chip failure. Additionally, the solution must not falsely implicate any component that functioned correctly. This article formally describes the property of accountability, a formal methodology of constructing an accountability solution, and a formal game-theory based methodology to reason about and prove the viability of a proposed solution. We explore the entire space of solutions, and characterize the attack surface and methods to provide accountability for each setting. We show non-intuitive results in this article where seemingly simple solutions actually provide very powerful theoretical guarantees in terms of accountability. Rajshekar Kalayappan, Smruti R. Sarangi |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2020 | VarSim: A Fast and Accurate Variability and Leakage Aware Thermal SimulatorabstractSome of the fastest thermal estimation techniques at the architectural level are based on Green’s functions (impulse response of a unit power source). The resultant temperature profile can be easily obtained by computing a convolution of the Green’s function and the power profile. Sadly, existing approaches do not take process and temperature variation into account, which are integral aspects of today’s technologies. This problem is still open. In this paper, we provide a closed-form solution for the Green’s function after taking process, temperature, and thermal conductivity variation into account. Moreover, during the process of computing the thermal map, we reduce the amount of redundant work by identifying similar regions in the chip using an unsupervised learning-based approach. We were able to obtain a 700,000X speedup over state-of-the-art proposals with a mean absolute error limited to 0.7◦C (1.5%). Hameedah Sultan, Smruti R. Sarangi |
DAC | 2 |
| 2020 | Performance Prediction for Multi-Application Concurrency on GPUsabstractWith the advent of edge computing and 5G, multiple mobile applications are being offloaded to cloud servers to meet their computational demands. Computer vision workloads dominate this space. Since the vision workloads are composed of linear algebra kernels, they perform significantly well on SIMT/SIMD architectures such as GPUs. While an application can maximize its performance on a GPU when it is the sole consumer of the GPU's resources, it fails to maintain that performance in a multi-application scenario. The primary cause of this problem is the lack of efficient virtualization techniques for GPUs and contention among the applications for the shared resources. Sadly, most of the prior work in this area is devoted to predicting single application performance. To the best of our knowledge we propose the first machine learning based predictor to predict the performance of an ensemble of applications on a GPU. Our predictor achieves an error of 9% across a suite of representative vision workloads for predicting the execution time. Competing algorithms that primarily work for single application scenarios have significantly inferior prediction accuracy and their error rates are more than 140%. Diksha Moolchandani, Sudhanshu Gupta 0002, Smruti R. Sarangi |
ISPASS | 4 |
| 2020 | SoftMon: A Tool to Compare Similar Open-source Software from a Performance PerspectiveabstractOver the past two decades, a rich ecosystem of open-source software has evolved. For every type of application, there are a wide variety of alternatives. We observed that even if different applications that perform similar tasks and compiled with the same versions of the compiler and the libraries, they perform very differently while running on the same system. Sadly prior work in this area that compares two code bases for similarities does not help us in finding the reasons for the differences in performance. Shubhankar Suman Singh, Smruti R. Sarangi |
MSR | 2 |
| 2020 | SecONet: A Security Framework for a Photonic Network-on-ChipabstractPhotonic networks are already commercially available at the board-level, and many fabrication facilities can fabricate optical networks and integrate them with traditional silicon-based SoCs. Almost all the research in on-chip photonics has been in the areas of performance enhancement and static power reduction. However, before the large-scale adoption of such technologies, it is necessary to solve security problems. As opposed to electrical NoCs, optical NoCs are shared to a much larger extent, and are significantly more sensitive to the latencies of cryptographic operations. Hence, it is necessary to design a novel protocol for securing such networks. We propose a novel, secure, and efficient optical network in this paper (SecONet) that is immune to eavesdropping, spoofing, replay, and message-removal attacks. Using a combination of speculative execution and pre-computation, we reduce the performance overhead of 39.53% with a conventional implementation to 14.2% for a suite of Splash2 and Parsec benchmarks. The additional area overhead of our proposed hardware is modest: 1.6%. Janibul Bashir, Chandran Goodchild, Smruti R. Sarangi |
NOCS | 3 |
| 2020 | Predict, Share, and Recycle Your Way to Low-power Nanophotonic NetworksabstractHigh static power consumption is widely regarded as one of the largest bottlenecks in creating scalable optical NoCs. The standard techniques to reduce static power are based on sharing optical channels and modulating the laser. We show in this article that state-of-the-art techniques in these areas are suboptimal, and there is a significant room for further improvement. We propose two novel techniques—a neural network--based method for laser modulation by predicting optical traffic and a distributed and altruistic algorithm for channel sharing—that are significantly closer to a theoretically ideal scheme. In spite of this, a lot of laser power still gets wasted. We propose to reuse this energy to heat micro-ring resonators (achieve thermal tuning) by efficiently recirculating it. These three methods help us significantly reduce the energy requirements. Our design consumes 4.7× lower laser power as compared to other state-of-the-art proposals. In addition, it results in a 31% improvement in performance and 39% reduction in ED 2 for a suite of Splash2 and Parsec benchmarks. Janibul Bashir, Smruti R. Sarangi |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2020 | GPUOPT: Power-efficient Photonic Network-on-Chip for a Scalable GPUabstractOn-chip photonics is a disruptive technology, and such NoCs are superior to traditional electrical NoCs in terms of latency, power, and bandwidth. Hence, researchers have proposed a wide variety of optical networks for multicore processors. The high bandwidth and low latency features of photonic NoCs have led to the overall improvement in the system performance. However, there are very few proposals that discuss the usage of optical interconnects in Graphics Processor Units (GPUs). GPUs can also substantially gain from such novel technologies, because they need to provide significant computational throughput without further stressing their power budgets. The main shortcoming of optical networks is their high static power usage, because the lasers are turned on all the time by default, even when there is no traffic inside the chip, and thus sophisticated laser modulation schemes are required. Such modulation schemes base their decisions on an accurate prediction of network traffic in the future. In this article, we propose an energy-efficient and scalable optical interconnect for modern GPUs called GPUOPT that smartly creates an overlay network by dividing the symmetric multiprocessors (SMs) into clusters. It furthermore has separate sub-networks for coherence and non-coherence traffic. To further increase the throughput, we connect the off-chip memory with optical links as well. Subsequently, we show that traditional laser modulation schemes (for reducing static power consumption) that were designed for multicore processors are not that effective for GPUs. Hence, there was a need to create a bespoke scheme for predicting the laser power usage in GPUs. Using this set of techniques, we were able to improve the performance of a modern GPU by 45% as compared to a state-of-the-art electrical NoC. Moreover, as compared to competing optical NoCs for GPUs, our scheme reduces the laser power consumption by 67%, resulting in a net 65% reduction in ED 2 for a suite of Rodinia benchmarks. Janibul Bashir, Smruti R. Sarangi |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2020 | Enhancing Network-on-Chip Performance by Reusing Trace BuffersabstractEnsuring the functional correctness of networks-on-chip (NoCs) can be particularly challenging, and communication-centric debug methodologies have been widely used by engineers to validate NoC functionality during post-silicon validation. Design-for-debug structures, such as trace buffers and monitors, are usually inserted in such systems-on-chip to enhance signal visibility. However, this debug hardware becomes underutilized once the chip goes into production. While the size and organization of the router buffers directly impact network throughput, these buffers also dominate the on-chip router area. We propose a scheme augmented virtual channel (AugVC) to reuse trace buffers to augment router buffers, with the objective of improving the overall network performance. The experimental results for a 64-node mesh network show that our proposed approach can reduce latency by up to 38.25% for transpose traffic compared to a baseline design with reduced buffer sizes. We also propose an extension, output port directed virtual channel (ODVC), that uses a modified virtual channel assignment strategy, on the basis of the designated output port of a network packet. This strategy reduces the average packet latency and area of the router by 45% and 32.4%, respectively. Neetu Jindal, Shubhani Gupta, Divya Praneetha Ravipati, Preeti Ranjan Panda, Smruti R. Sarangi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | VisSched: An Auction-Based Scheduler for Vision Workloads on Heterogeneous ProcessorsabstractWith the growth of edge computing, applicationspecific workloads based on computer vision are steadily migrating to edge cloudlets. Scheduling has been identified to be a major problem in these cloudlets. In this article, we propose a generic architectural solution, VisSched, that leverages the fact that most vision workloads share similar code kernels (such as library code for linear algebra), and as a result, they tend to exhibit similar phase behavior. This allows us to create an auction theory-based scheduling mechanism, where we give each thread a replenishable virtual wallet, and threads are scheduled based on the amounts that they bid for executing on a free core. We show that in 20%-40% of the cases, our scheduling algorithm is theoretically optimal, and in the remaining cases, it reaches a global optimum obtained using Monte Carlo simulations 90%-95% of the time. Our results for the MEVBench vision workloads show a 17% higher performance and a 14% lower ED2as compared to the nearest competing algorithm in the literature. Diksha Moolchandani, José F. Martínez, Smruti R. Sarangi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | A Fast Leakage-Aware Green's-Function-Based Thermal Simulator for 3-D ChipsabstractIn this article, we propose a fast thermal modeling tool, 3DSim, using a Green's-function-based approach. Green's-function-based approaches have been shown to be faster than the traditional finite-difference-based techniques. Our proposed tool can model steady-state and transient thermal profiles for both 2-D and 3-D chips, which may contain multiple active layers and fluid-carrying microchannels for heat removal. The unique advantage of our tool is that it models leakage power analytically using a piecewise-linear leakage model, thereby eliminating the need to iterate multiple times through the leakage-temperature feedback loop. We use several algebraic techniques and transforms to compute the thermal profile analytically and thereby speedup the process of calculation. To the best of our knowledge, transform-based approaches have not been used before to model the temperature in 3-D chips with microchannels. Our approach provides a 150× speedup over state-of-the-art thermal simulators, with an error limited to 5%. Hameedah Sultan, Smruti R. Sarangi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2019 | FlexiCheck: An Adaptive Checkpointing Architecture for Energy Harvesting DevicesabstractWith the advent of 5G and M2M architectures, energy harvesting devices are expected to become far more prevalent. Such devices harvest energy from ambient sources such as solar energy or vibration energy (from machines) and use it for sensing the environmental parameters and further processing them. Given that the rate of energy consumption is more than the rate of energy production, it is necessary to frequently halt the processor and accumulate energy from the environment. During this period it is mandatory to take a checkpoint to avoid the loss of data. State of the art algorithms use software based methods that extensively rely on compiler analyses. In this paper, we provide the first formal model for such systems, and show that we can arrive at an optimal check-pointing schedule using a quadratically constrained linear program (QCLP) solver. Using this as a baseline, we show that existing algorithms for checkpointing significantly underperform. Furthermore, we prove and demonstrate that when we have a relatively constant energy source, a greedy algorithm provides an optimal solution. To model more complex situations where the energy varies, we create a novel checkpointing algorithm that adapts itself according to the ambient energy. We obtain a speedup of 2 - 5× over the nearest competing approach, and we are within 3 - 8% of the optimal solution in the general case where the ambient energy exhibits variations. Priyanka Singla 0001, Shubhankar Suman Singh, Smruti R. Sarangi |
DATE | 3 |
| 2019 | NanoTherm: An Analytical Fourier-Boltzmann Framework for Full Chip Thermal SimulationsabstractTemperature simulation is a classic problem in EDA, and researchers have been working on it for at least the last 15 years. In this paper, we focus on fast Green's function based approaches, where computing the temperature profile is as simple as computing the convolution of the power profile with the Green's function. We observe that for many problems of interest the process of computing the Green's function is the most time consuming phase, because we need to compute it with the slower finite difference or finite element based approaches. In this paper we propose a solution, NanoTherm, to compute the Green's function using a fast analytical approach that exploits the symmetry in the thermal distribution. Secondly, conventional analyses based on the Fourier's heat transfer equation fail to hold at the nanometer level. To accurately compute the temperature at the level of a standard cell, it is necessary to solve the Boltzmann transport equation (BTE) that accounts for quantum mechanical effects. This research area is very sparse. Conventional approaches ignore the quantum effects, which can result in a 25 to 60% error in temperature calculation. Hence, we propose a fast analytical approach to solve the BTE and obtain an exact solution in the Fourier transform space. Using our fast analytical models, we demonstrate a speedup of 7-668X over state of the art techniques with an error limited to 3% while computing the combined Green's function. Shashank Varshney, Hameedah Sultan, Palkesh Jain, Smruti R. Sarangi |
ICCAD | 4 |
| 2019 | Power efficient photonic network-on-chip for a scalable GPUabstractIn this paper, we propose an energy efficient and scalable optical interconnect for GPUs. We intelligently divide the components in a GPU into different types of clusters and enable these clusters to communicate optically with each other. In order to reduce the network delay, we use separate networks for coherence and non-coherence traffic. Moreover, to reduce the static power consumption in optical interconnects, we modulate the off-chip light source by proposing a novel GPU specific prediction scheme for on-chip network traffic. Using our design, we were able to increase the performance by 17% and achieve a 65% reduction in ED2 as compared to a state-of-the-art optical topology. Janibul Bashir, Khushal Sethi, Smruti R. Sarangi |
NOCS | 3 |
| 2019 | A Reference Architecture for Smart and Software-Defined BuildingsabstractThe vision encompassing Smart and Software-defined Buildings (SSDB) is becoming more popular and its implementation is now more accessible due to the widespread adoption of the Internet of Things (IoT) infrastructure. Some of the most important applications sustaining this vision are energy management, environmental comfort, safety and surveillance. This paper surveys IoT and SSB technologies and their cooperation towards the realization of smart spaces. We propose a four-layer reference architecture and we organize related concepts around it. This conceptual frame is useful to identify the current literature on the topic and to connect the dots into a coherent vision of the future of residential and commercial buildings. Manuel Mazzara, Ilya Afanasyev 0001, Smruti R. Sarangi, Salvatore Distefano, Vivek Kumar 0007, Muhammad Ahmad 0002 |
SMARTCOMP | 3 |
| 2019 | BigBus: A Scalable Optical InterconnectabstractThis article presents BigBus , a novel design of an on-chip photonic network for a 1,024-node system. For such a large on-chip network, performance and power reduction are two mutually conflicting goals. This article uses a combination of strategies to reduce static power consumption while simultaneously improving performance and the energy-delay 2 ( ED 2 ) product. The crux of the article is to segment the entire system into smaller clusters of nodes and adopt a hybrid strategy for each segment that includes conventional laser modulation, as well as a novel technique for sharing power across nodes dynamically. We represent energy internally as tokens, where one token will allow a node to send a message to any other node in its cluster. We allow optical stations to arbitrate for tokens at a global level, and then we predict the number of token equivalents of power that the off-chip laser needs to generate. Using these techniques, BigBus outperforms other competing proposals. We demonstrate a speedup of 14--34% over state of the art proposals and a 20--61% reduction in ED 2 . Janibul Bashir, Eldhose Peter, Smruti R. Sarangi |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2018 | HPXA: A highly parallel XML parserabstractState of the art XML parsing approaches read an XML file byte by byte, and use complex finite state machines to process each byte. In this paper, we propose a new parser, HPXA, which reads and processes 16 bytes at a time. We designed most of the components ab initio, to ensure that they can process multiple XML tokens and tags in parallel. We propose two basic elements - a sparse 1D array compactor, and a hardware unit called LTMAdder that takes its decisions based on adding the rows of a lower triangular matrix. We demonstrate that we are able to process 16 bytes in parallel with very few pipeline stalls for a suite of widely used XML benchmarks. Moreover, for a 28nm technology node, we can process XML data at 106 Gbps, which is roughly 6.5X faster than competing prior work. Isaar Ahmad, Sanjog Patil, Smruti R. Sarangi |
DATE | 3 |
| 2018 | Probabilistic Sequential Consistency in Social NetworksabstractResearchers have proposed numerous consistency models in distributed systems that offer higher performance than classical sequential consistency (SC). Even though these models do not guarantee sequential consistency; they either behave like an SC model under certain restrictive scenarios, or ensure SC behavior for a part of the system. We propose a different line of thinking where we try to accurately estimate the number of SC violations, and then try to adapt our system to optimally tradeoff performance, resource usage, and the number of SC violations. In this paper, we propose a generic theoretical model that can be used to analyze systems that are comprised of multiple sub-domains - each sequentially consistent. It is validated with real world measurements. Next, we use this model to propose a new form of consistency called social consistency, where socially connected users perceive an SC execution, whereas the rest of the users need not. We create a prototype social network application and implement it on the Cassandra key-value store. We show that our system has 2.4× more throughput than Cassandra and provides 37% better quality-of-experience. Priyanka Singla 0001, Shubhankar Suman Singh, K. Gopinath, Smruti R. Sarangi |
HiPC | 4 |
| 2018 | Providing Accountability in Heterogeneous Systems-on-ChipabstractWhen modern systems-on-chip (SoCs), containing designs from different organizations, miscompute or underperform in the field, discerning the responsible component is a non-trivial task. A perfectly accountable system is one in which the on-chip component at fault is always unambiguously detected. The achievement of accountability can be greatly aided by the collection of runtime information that captures the events in the system that led to the error. Such information collection must be fair and impartial to all parties. In this article, we prove that logging messages communicated between components from different organizations is sufficient to provide accountability, provided the logs are authentic. We then construct a solution based on this premise, with an on-chip trusted auditing system to authenticate the logs. We present a thorough design of the auditing system, and demonstrate that its performance overhead is a mere 0.49%, and its area overhead is a mere 0.194% (in a heterogeneous 48 core, 400 mm 2 chip). We also demonstrate the viability of this solution using three representative bugs found in popular commercial SoCs. Rajshekar Kalayappan, Smruti R. Sarangi |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | Reusing Trace Buffers as Victim Caches
Neetu Jindal, Preeti Ranjan Panda, Smruti R. Sarangi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | POSTER: BigBus: A Scalable Optical InterconnectabstractThis paper presents BigBus, a novel on-chip photonic network for a 1024 node system. The crux of the idea is to segment the entire system into smaller clusters of nodes, and adopt a hybrid strategy for each segment that includes conventional laser modulation, as well as a novel technique for sharing power across nodes dynamically. We represent energy internally as tokens, where one token will allow a node to send a message to any other node in its cluster. We allow optical stations to arbitrate for tokens and at a global level, we predict the number of token equivalents of power that the off-chip laser needs to generate. Eldhose Peter, Janibul Bashir, Smruti R. Sarangi |
PACT | 3 |
| 2017 | Reusing trace buffers to enhance cache performanceabstractWith the increasing complexity of modern Systems-on-Chip, the possibility of functional errors escaping design verification is growing. Post-silicon validation targets the discovery of these errors in early hardware prototypes. Due to limited visibility and observability, dedicated design-for-debug (DFD) hardware such as trace buffers are inserted to aid post-silicon validation. In spite of its benefit, such hardware incurs area overheads, which impose size limitations. However, the overhead could be overcome if the area dedicated to DFD could be reused in-field. In this work, we present a novel method for reusing an existing trace buffer as a victim cache of a processor to enhance performance. The trace buffer storage space is reused for the victim cache, with a small additional controller logic. Experimental results on several benchmarks and trace buffer sizes show that the proposed approach can enhance the average performance by up to 8.3% over a baseline architecture. We also propose a strategy for dynamic power management of the structure, to enable saving energy with negligible impact on performance. Neetu Jindal, Preeti Ranjan Panda, Smruti R. Sarangi |
DATE | 3 |
| 2017 | A hardware implementation of the MCAS synchronization primitiveabstractLock-based parallel programs are easy to write. However, they are inherently slow as the synchronization is blocking in nature. Non-blocking lock-free programs, which use atomic instructions such as compare-and-set (CAS), are significantly faster. However, lock-free programs are notoriously difficult to design and debug. This can be greatly eased if the primitives work on multiple memory locations instead of one. We propose MCAS, a hardware implementation of a multi-word compare-and-set primitive. Ease of programming aside, MCAS-based programs are 13.8X and 4X faster on an average than lock-based and traditional lock-free programs respectively. The area overhead, in a 32-core 400mm2 chip, is a mere 0.046%. Srishty Patel, Rajshekar Kalayappan, Ishani Mahajan, Smruti R. Sarangi |
DATE | 4 |
| 2017 | A fast leakage aware thermal simulator for 3D chipsabstractIn this paper, we propose, 3DSim, which is an ultrafast thermal simulator for 3D chips. It simulates the effects of both dynamic and leakage power. Our technique captures the steady state as well as the transient response with a high speed and good accuracy. 3DSim uses an approach based on Green's functions, where a Green's function is defined as the impulse response of a unit power source. Our approach incorporates the effects of the leakage-temperature feedback loop, exploits the radial symmetry in the thermal profile, and uses Hankel transforms to yield a closed form solution for the leakage aware Green's function. To further speed up our technique, we use fast numerical discrete Hankel transforms, and pre-compute and store certain functions in a lookup table. Our approach fundamentally converts a 3D problem to a set of 1D problems, thus leading to a 68X speedup as compared to competing simulators with an error limited to 1.5 C. Hameedah Sultan, Smruti R. Sarangi |
DATE | 2 |
| 2017 | Expander: Lock-Free Cache for a Concurrent Data StructureabstractParallel programming models and paradigms are increasingly becoming more expressive with a steady increase in the number of cores that can be placed on a single chip. Concurrent data structures for shared memory parallel pro- grams are now being used in operating systems, middle-ware, and device drivers. In such a shared memory model, processes communicate and synchronize by applying primitive operations on memory words. To implement concurrent data structures that are linearizable and possibly lock-free or wait-free, it is often necessary to add additional information to memory words in a data structure. This additional information can range from a single bit to multiple bits that typically represent thread ids, request ids, timestamps, and other application dependent fields. Since most processors can perform compare-And-Set (CAS) or load-link/store-conditional (LL/SC) operations on only 64 bits at a time, current approaches either use some bits in a memory word to pack additional information (packing), or use the bits to store a pointer to an object that contains additional information (redirection), and the original data item. The former approach restricts the number of bits for each additional field and this reduces the range of the field, and the latter approach is wasteful in terms of space. We propose a novel and universal method called a memory word expander in this paper. It caches information for a set of memory locations that need to be augmented with additional information. It supports traditional atomic get, set, and CAS operations, and tries to maintain state for a minimum number of entries. We experimentally demonstrate that it is possible to reduce the runtime memory footprint by 20-35% for algorithms that use redirection. For algorithms that use packing, the use of the EXPANDER can make them feasible. The performance overhead is within 2-13% for 32 threads. When we compare the performance of the EXPANDER based non-blocking algorithms with the version that uses locks, we have a performance gain of at least 10-100X. Pooja Aggarwal, Smruti R. Sarangi |
HiPC | 2 |
| 2017 | NUPLet: A Photonic Based Multi-Chip NUCA ArchitectureabstractArea, manufacturing yield and lack of scalable interconnects restrict single chip designs to a small number of cores (16-32). However, multi-chip designs with the help of silicon photonics can overcome area and yield constraints and make it possible to design a virtual chip, which can scale to a large number of cores. Sadly, the scalability of such designs is limited by the high percentage of inter-chip messages and relatively lower hit rate in remote cache banks. In this paper, we propose NUPLet, a multi-chip architecture that tries to remove these limitations by separating the intra and inter chip networks. It proposes to use a non-uniform cache architecture (NUCA) scheme on top of a virtual chip in order to decrease inter chip communication and increase the hit rate in the last level cache. In addition, we propose a prediction mechanism for predicting the number of inter chip messages in the network. This is used to modulate the laser accordingly, and reduce static power consumption. We simulated a four chip based NUPLet design with each chip containing 32 cores. For a suite of Splash2 and Parsec benchmarks, NUPLet increased the last level cache hit rate by 70% as compared to other state of the art proposals. Furthermore, NUPLet improved performance by 28%, reduced power consumption by 39%, and reduced ED2by 41%. Janibul Bashir, Smruti R. Sarangi |
ICCD | 2 |
| 2017 | Schedtask: a hardware-assisted task schedulerabstractThe execution of workloads such as web servers and database servers typically switches back and forth between different tasks such as user applications, system call handlers, and interrupt handlers. The combined size of the instruction footprints of such tasks typically exceeds that of the i-cache (16--32 KB). This causes a lot of i-cache misses and thereby reduces the application's performance. Hence, we propose SchedTask, a hardware-assisted task scheduler that improves the performance of such workloads by executing tasks with similar instruction footprints on the same core. We start by decomposing the combined execution of the OS and the applications into sequences of instructions called SuperFunctions. We propose a scheme to determine the amount of overlap between the instruction footprints of different SuperFunctions by using Bloom filters. We then use a hierarchical scheduler to execute SuperFunctions with similar instruction footprints on the same core. For a suite of 8 popular OS-intensive workloads, we report an increase in the application's performance of up to 29 percentage points (mean: 11.4 percentage points) over state of the art scheduling techniques. Prathmesh Kallurkar, Smruti R. Sarangi |
MICRO | 2 |
| 2017 | Optical Overlay NUCA: A High-Speed Substrate for Shared L2 CachesabstractIn this article, we propose using optical networks-on-chip (NoCs) to design cache access protocols for large shared L2 caches. We observe that the problem is unique because optical networks have very low latency, and in principle all of the cache banks are very close to each other. A naive approach is to broadcast a request to a set of banks that might possibly contain the copy of a block. However, this approach is wasteful in terms of energy and bandwidth. Hence, we propose a set of novel schemes that create a set of virtual networks ( overlays ) of cache banks over a physical optical NoC. We search for a block inside each overlay using a combination of multicast and unicast messages. We first propose two simple protocols: TSI and Broadcast . The former uses unicast messages, and the latter uses multicast messages. We subsequently propose an improved scheme, OP_BCAST , that combines the best of TSI and Broadcast , and mainly uses restricted multicast messages. Then we propose a set of novel hardware structures for creating and managing overlays, for efficiently locating blocks in the overlay, and for implementing dynamically changing overlays with OP_BCAST . The performance of the TSI scheme is within 2% to 3% of a broadcast scheme, and it is faster than traditional schemes with electrical networks by 26%. Compared to the broadcast scheme, it reduces the number of accesses, and consequently the dynamic energy of the caches by 6% to 8%. OP_BCAST is 34% faster than the best solutions with copper-based NoCs; moreover, it reduces the dynamic energy for cache access by 33% compared to the TSI scheme. Eldhose Peter, Anuj Arora, Janibul Bashir, Akriti Bagaria, Smruti R. Sarangi |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2017 | Managing Trace Summaries to Minimize Stalls During Postsilicon ValidationabstractOn-chip trace buffers are increasingly being used for at-speed debug during postsilicon validation. The limited size of these buffers results in their frequent overflowing. In scenarios when such overflowing is not desirable, the chip is stalled, and the state data recorded in these buffers are transferred off-chip. Such frequent stalling significantly impedes efficient debugging. We propose a novel scheme to minimize the number of such stalls using a portion of the trace buffer to also store summaries of trace messages. We describe an overlapped trace buffer architecture that uses a reduced number of ports to capture tapered summaries where both detailed and summary versions of traces are stored simultaneously. We propose a simple hardware structure to generate two kinds of trace summaries-spatial and temporal-as specified by the validation engineer. We introduce a storage specification language that allows the validation engineer to unambiguously specify the information to be captured in these summaries to the debug hardware. We demonstrate that our proposal significantly reduces the number of stalls for off-chip transfer of captured traces in four bug scenarios that are representative of different classes of bugs encountered during postsilicon validation. Sandeep Chandran, Preeti Ranjan Panda, Smruti R. Sarangi, Ayan Bhattacharyya, Deepak Chauhan, Sharad Kumar |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Extending trace history through tapered summaries in post-silicon validationabstractOn-chip trace buffers are increasingly being used for at-speed debug during post-silicon validation. However, the activity history captured by these buffers is small due to their limited size. We propose a novel scheme that extends the captured trace history (by upto 162%) by using a portion of the trace buffer to also store summaries of trace messages. We describe an Overlapped trace buffer architecture that uses a reduced number of ports to capture tapered summaries where both detailed and summary versions of traces are stored simultaneously. We demonstrate the usefulness of the proposed methodology for debugging various classes of bugs encountered during post-silicon validation. Sandeep Chandran, Preeti Ranjan Panda, Deepak Chauhan, Sharad Kumar, Smruti R. Sarangi |
ASP-DAC | 5 |
| 2016 | Noise Aware Scheduling in Data CentersabstractAs the demand for large scale computing is rapidly increasing to serve billions of users across the world, more powerful and densely packed server configurations are being used. Often in developing countries, and in small and medium enterprises, it is hard to place such servers in sound-proof server rooms. Hence, servers are typically placed in close proximity to employees. The noise from the cooling fans in servers adversely affects employees' health, and reduces their productivity. In this paper, we provide a framework for computer architects to measure the acoustic profile in a data center along with the temperature profile, and estimate the sound power levels at points of interest. Additionally, we studied the noise levels obtained upon using algorithms targeted at homogenizing the temperature profile. We found that these algorithms result in high noise levels, sometimes above the permissible levels. So, we propose two heuristics to redistribute workloads in a data center such that noise can be reduced at certain target locations. We obtain a noise reduction of 2-13 dB when compared with uniform workload distribution, and upto 16 dB as compared to temperature aware workload placement, with a reduction of at least 5-6 dB in 75% of the cases. The performance overhead is limited to 1%. Hameedah Sultan, Arpit Katiyar, Smruti R. Sarangi |
ICS | 3 |
| 2016 | pTask: A smart prefetching scheme for OS intensive applicationsabstractInstruction prefetching is a standard approach to improve the performance of operating system (OS) intensive workloads such as web servers, file servers and database servers. Sophisticated instruction prefetching techniques such as PIF [12] and RDIP [17] record the execution history of a program in dedicated hardware structures and use this information for prefetching if a known execution pattern is repeated. The storage overheads of the additional hardware structures are prohibitively high (64-200 KB per core). This makes it difficult for the deployment of such schemes in real systems. We propose a solution that uses minimal hardware modifications to tackle this problem. We notice that the execution of server applications keeps switching between tasks such as the application, system call handlers, and interrupt handlers. Each task has a distinct instruction footprint, and is separated by a special OS event. We propose a sophisticated technique to capture the instruction stream in the vicinity of such OS events; the captured information is then compressed significantly and is stored in a process's virtual address space. Special OS routines then use this information to prefetch instructions for the OS and the application codes. Using modest hardware support (4 registers per core), we report an increase in instruction throughput of 2-14% (mean: 7%) over state of the art instruction prefetching techniques for a suite of 8 popular OS intensive applications. Prathmesh Kallurkar, Smruti R. Sarangi |
MICRO | 2 |
| 2016 | FluidCheck: A Redundant Threading-Based Approach for Reliable Execution in Manycore ProcessorsabstractSoft errors have become a serious cause of concern with reducing feature sizes. The ability to accommodate complex, Simultaneous Multithreading (SMT) cores on a single chip presents a unique opportunity to achieve reliable execution, safe from soft errors, with low performance penalties. In this context, we present FluidCheck , a checker architecture that allows highly flexible assignment and migration of checking duties across cores. In this article, we present a mechanism to dynamically use the resources of SMT cores for checking the results of other threads, and propose a variety of heuristics for migration of such checker threads across cores. Secondly, to make the process of checking more efficient, we propose a set of architectural enhancements that reduce power consumption, decrease the length of the critical path, and reduce the load on the Network-on-Chip (NoC). Based on our observations, we design a 16 core system for running SPEC2006 based bag-of-tasks applications. Our experiments demonstrate that fully reliable execution can be attained with a mere 27% slowdown, surpassing traditional redundant threading based techniques by roughly 42%. Rajshekar Kalayappan, Smruti R. Sarangi |
ACM Trans. Archit. Code Optim. | 2 |
| 2016 | Lock-Free and Wait-Free Slot Scheduling AlgorithmsabstractIn this paper, we consider the design space of parallel non-blocking slot scheduling algorithms. Slot schedulers divide time into discrete quanta called slots, and schedule resources at the granularity of slots. They are typically used in high throughput I/O systems, data centers, video servers, and network drivers. We propose a family of parallel slot scheduling problems of increasing complexity, and then propose parallel lock-free and wait-free algorithms to solve them. In specific, we propose problems that can reserve, as well as free a set of contiguous slots in a non-blocking manner. We show that in a system with 64 threads, it is possible to get speedups of 10X by using lock-free algorithms as compared to a baseline implementation that uses locks. We additionally propose wait-free algorithms, whose mean performance is roughly the same as the version with locks. However, they suffer from significantly lower jitter and ensure a high degree of fairness among threads. Pooja Aggarwal, Smruti R. Sarangi |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | Area-Aware Cache Update Trackers for Postsilicon ValidationabstractThe internal state of the complex modern processors often needs to be dumped out frequently during postsilicon validation. Since the caches hold most of the state, the volume of data dumped and the transfer time are dominated by the large caches present in the architecture. The limited bandwidth to transfer data present in these large caches off-chip results in stalling the processor for long durations when dumping the cache contents off-chip. To alleviate this, we propose to transfer only those cache lines that were updated since the previous dump. Since maintaining a bit-vector with a separate bit to track the status of individual cache lines is expensive, we propose two methods: 1) where a bit tracks multiple cache lines and 2) an Interval Table which stores only the starting and ending addresses of continuous runs of updated cache lines. Both methods require significantly lesser space compared with a bit-vector, and allow the designer to choose the amount of space to allocate for this design-for-debug feature. The impact of reducing storage space is that some nonupdated cache lines are dumped too. We attempt to minimize such overheads. We propose a scheme to share such cache update tracking hardware (or Update Trackers) across multiple caches in case of physically distributed caches so that they are replicated fewer times, thereby limiting the area overhead. We show that the proposed Update Trackers occupy less than 1% of cache area for both the shared and distributed caches. Sandeep Chandran, Smruti R. Sarangi, Preeti Ranjan Panda |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | ColdBus: A Near-Optimal Power Efficient Optical BusabstractHigh static power dissipation is one of the largest hurdles in the widespread adoption of on-chip optical networks. There are numerous proposals in literature for the reduction of static power by predicting the activity in future time intervals, and then powering off the laser, and also by sharing optical bandwidth. We propose a novel scheme called ColdBus in this paper that introduces a new method of predicting the optical NoC traffic using program counter addresses, moreover, we then propose novel structures and predictors for handling mispredictions. These units allow optical transmitters to share bandwidth between them while simultaneously minimizing static power. We show that ColdBus consumes 4.1X less optical power than other state of the art proposals. It additionally has a speedup of 3.1%, and improves ED2 by 18% for a suite of Splash, Parsec, and Parboil benchmarks. Eldhose Peter, Arun Thomas, Anuj Dhawan, Smruti R. Sarangi |
HiPC | 4 |
| 2015 | FP-NUCA: A Fast NOC Layer for Implementing Large NUCA CachesabstractNUCA caches have traditionally been proposed as a solution for mitigating wire delays, and delays introduced due to complex networks on chip. Traditional approaches have reported significant performance gains with intelligent block placement, location, replication, and migration schemes. In this paper, we propose a novel approach in this space, called FP-NUCA. It differs from conventional approaches, and relies on a novel method of co-designing the last level cache and the network on chip. We artificially constrain the communication pattern in the NUCA cache such that all the messages travel along a few predefined paths (fast paths) for each set of banks. We leverage this communication pattern by designing a new type of NOC router called the Freeze router, which augments a regular router by adding a layer of circuitry that gates the clock of the regular router when there is a fast path message waiting to be transmitted. Messages along the fast path do not require buffering, switching, or routing. We incorporate a bank predictor with our novel NOC for reducing the number of messages, and resultant energy consumption. We compare our performance with state of the art protocols, and report speedups of up to 31 percent (mean: 6.3 percent), and ED2 reduction up to 46 percent (mean: 10.4 percent) for a suite of Splash and Parsec benchmarks. We implement the Freeze router in VHDL and show that the additional fast path logic has minimal area and timing overheads. Anuj Arora, Mayur Harne, Hameedah Sultan, Akriti Bagaria, Smruti R. Sarangi |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2014 | LightSim: A leakage aware ultrafast temperature simulatorabstractIn this paper, we propose the design of an ultra-fast temperature simulator (LightSim) that can perform both steady state and transient thermal analysis, and also take the effect of leakage power into account. We use a novel Hankel transform based technique to derive a transient version of the Green's function for a chip, which takes into account the feedback loop between temperature and leakage. Subsequently, we calculate the temperature map of a chip by convolving the derived Green's function with the power map. Our simulator is at least 3500 times faster than HotSpot, and at least 2.3 times faster than competing research prototypes [4, 12]. The total error is limited to 0.18 °C. Smruti R. Sarangi, Gayathri Ananthanarayanan, M. Balakrishnan |
ASP-DAC | 1 |
| 2014 | RADIR: Lock-free and wait-free bandwidth allocation models for solid state drivesabstractNovel applications such as micro-blogging and algorithmic trading typically place a very high load on the underlying storage system. They are characterized by a stream of very short requests, and thus they require a very high I/O throughput. The traditional solution for supporting such applications is to use an array of hard disks. With the advent of solid state drives (SSDs), storage vendors are increasingly preferring them because their I/O throughput can scale up to a million IOPS (I/O operations per second). In this paper, we design a family of algorithms, RADIR, to schedule requests for such systems. Our algorithms are lock-free/wait-free, lineariz-able, and take the characteristics of requests into account such as the deadlines, request sizes, dependences, and the amount of available redundancy in RAID configurations. We perform simulations with workloads derived from traces provided by Microsoft and demonstrate a scheduling throughput of 900K IOPS on a 64 thread Intel server. Our algorithms are 2-3 orders of magnitude faster than the versions that use locks. We show detailed results for the effect of deadlines, request sizes, and the effect of RAID levels on the quality of the schedule. Pooja Aggarwal, Giridhar Yasa, Smruti R. Sarangi |
HiPC | 3 |
| 2014 | TriKon: A hypervisor aware manycore processorabstractVirtualization is increasingly being deployed to run applications in a cloud computing environment. Sadly, there are overheads associated with hypervisors that can prohibitively reduce application performance. A major source of the overheads is the destructive interference between the application, OS, and hypervisor in the memory system. We characterize such overheads in this paper, and propose the design of a novel Triangle cache that can effectively mitigate destructive interference across these three classes of workloads. We subsequently, proceed to design the TriKon manycore processor that consists of a set of heterogeneous cores with caches of different sizes, and Triangle caches. To maximize the throughput of the system as a whole, we propose a dynamic scheduling algorithm for scheduling a class of system and CPU intensive applications on the set of heterogeneous cores. The area of the TriKon processor is within 2% of a baseline processor, and with such a system, we could achieve a performance gain of 12% for a suite of benchmarks. Within this suite, the system intensive benchmarks show a performance gain of 20% while the performance of the compute intensive ones remains unaffected. Also, by allocating extra area for cores with sophisticated cache designs, we further improved the performance of the system intensive benchmarks to 30%. Rohan Bhalla, Prathmesh Kallurkar, Smruti R. Sarangi |
HiPC | 4 |
| 2014 | GpuTejas: A parallel simulator for GPU architecturesabstractIn this paper, we introduce a new Java-based parallel GPGPU simulator, GpuTejas. GpuTejas is a fast trace driven simulator, which uses relaxed synchronization, and non-blocking data structures to derive its speedups. Secondly, it introduces a novel scheduling and partitioning scheme for parallelizing a GPU simulator. We evaluate the performance of our simulator with a set of Rodinia benchmarks. We demonstrate a mean speedup of 17.33x with 64 threads over sequential execution, and a speedup of 429X over the widely used simulator GPGPU-Sim. We validated our timing and simulation model by comparing our results with a native system (NVIDIA Tesla M2070). As compared to the sequential version of GpuTejas, the parallel version has an error limited to <;7.67% for our suite of benchmarks, which is similar to the numbers reported by competing parallel simulators. Geetika Malhotra, Seep Goel, Smruti R. Sarangi |
HiPC | 3 |
| 2014 | Optical overlay NUCA: A high speed substrate for shared L2 cachesabstractIn this paper, we propose to use optical NOCs to design cache access protocols for large shared L2 caches. We observe that the problem is unique because optical networks have very low latency, and in principle all the cache banks are very close to each other. A naive approach is to broadcast a request to a set of banks that might possibly contain the copy of a block. However, this approach is wasteful in terms of energy and bandwidth. Hence, we propose a novel scheme in this paper, TSI, which proposes to create a set of virtual networks (overlays) of cache banks over a physical optical NOC. We search for a block inside each overlay using a combination of multicast and unicast messages. We additionally create support for our overlay networks by proposing optimizations to the previously proposed R-SWMR network. We also propose a set of novel hardware structures for creating and managing overlays, and for efficiently locating blocks in the overlay. The performance of the TSI scheme is within 2-3% of a broadcast scheme, and it is faster than traditional static NUCA schemes by 50%. As compared to the broadcast scheme it reduces the number of accesses, and consequently the dynamic energy by 20-30%. Eldhose Peter, Anuj Arora, Akriti Bagaria, Smruti R. Sarangi |
HiPC | 4 |
| 2014 | ParTejas: A parallel simulator for multicore processorsabstractIn this paper, we present the design of a novel multicore simulator called ParTejas. It is a fast shared memory based parallel simulator written in Java. Unlike recently released parallel simulators that mainly rely on sampling, high level models, and highly relaxed synchronization, we primarily rely on novel concurrent data structures. In specific, we use a lock free parallel slot scheduler for synchronizing the accesses of multiple threads at a shared resource, and we use flexible barriers known as phasers to relax synchronization within bounds. We leverage additional language specific features of Java, and demonstrate a mean speedup of 11.8X (simulation speed of 4-8 MIPS) with 64 threads for a suite of Splash2 and Parsec benchmarks. Geetika Malhotra, Pooja Aggarwal, Abhishek Sagar, Smruti R. Sarangi |
ISPASS | 4 |
| 2014 | Architectural Support for Handling Jitterin Shared Memory Based Parallel ApplicationsabstractWith an increasing number of cores per chip, it is becoming harder to guarantee optimal performance for parallel shared memory applications due to interference caused by kernel threads, interrupts, bus contention, and temperature management schemes (referred to as jitter). We demonstrate that the performance of parallel programs gets reduced (up to 35.22 percent) in large CMP based systems. In this paper, we characterize the jitter for large multi-core processors, and evaluate the loss in performance. We propose a novel jitter measurement unit that uses a distributed protocol to keep track of the number of wasted cycles. Subsequently, we try to compensate for jitter by using DVFS across a region of timing critical instructions called a frame. Additionally, we propose an OS cache that intelligently manages the OS cache lines to reduce memory interference. By performing detailed cycle accurate simulations, we show that we are able to execute a suite of Splash2 and Parsec benchmarks with a deterministic timing overhead limited to 2 percent for 14 out of 17 benchmarks with modest DVFS factors. We reduce the overall jitter by an average 13.5 percent for Splash2 and 6.4 percent for Parsec. The area overhead of our scheme is limited to 1 percent. Sandeep Chandran, Prathmesh Kallurkar, Smruti R. Sarangi |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2013 | Space sensitive cache dumping for post-silicon validationabstractThe internal state of complex modern processors often needs to be dumped out frequently during post-silicon validation. Since the last level cache (considered L2 in this paper) holds most of the state, the volume of data dumped and the transfer time are dominated by the L2 cache. The limited bandwidth to transfer data off-chip coupled with the large size of L2 cache results in stalling the processor for long durations when dumping the cache contents off-chip. To alleviate this, we propose to transfer only those cache lines that were updated since the previous dump. Since maintaining a bit-vector with a separate bit to track the status of individual cache lines is expensive, we propose 2 methods: (i) where a bit tracks multiple cache lines and (ii) an Interval Table which stores only the starting and ending addresses of continuous runs of updated cache lines. Both methods require significantly lesser space compared to a bit-vector, and allow the designer to choose the amount of space to allocate for this design-for-debug (DFD) feature. The impact of reducing storage space is that some non-updated cache lines are dumped too. We attempt to minimize such overheads. Further, the Interval Table is independent of the cache size which makes it ideal for large caches. Through experimentation, we also determine the break-even point below which a t-lines/bit bit-vector is beneficial compared to an Interval Table. Sandeep Chandran, Smruti R. Sarangi, Preeti Ranjan Panda |
DATE | 2 |
| 2013 | Lock-Free and Wait-Free Slot Scheduling AlgorithmsabstractScalable scheduling is being increasingly regarded as an important requirement in high performance systems. There is a demand for high throughput schedulers in servers, data-centers, networking hardware, large storage systems, and in multi-cores of the future. In this paper, we consider an important subset of schedulers namely slot schedulers that discretize time into quanta called slots. Slot schedulers are commonly used for scheduling jobs in a large number of applications. Current implementations of slot schedulers are either sequential, or use locks. Sadly, lock based synchronization can lead to blocking, and deadlocks, and effectively reduces concurrency. To mitigate these problems, we propose a set of parallel lock-free and wait-free slot scheduling algorithms. Our algorithms are immune to operating system jitter, and guarantee forward progress. Additionally, all our algorithms are linearizable and expose the scheduler's interface as a shared data structure with standard semantics. We empirically demonstrate the scalability of our algorithms for a setup with thousands of requests per second on a 24 thread server. The wait free algorithms are most of the time as fast as the lock-free versions (3X-8X slower in the worst case). Pooja Aggarwal, Smruti R. Sarangi |
IPDPS | 2 |
| 2012 | Efficient on-line algorithm for maintaining k-cover of sparse bit-stringsabstractWe consider the on-line problem of representing a sparse bit string by a set of k intervals, where k is much smaller than the length of the string. The goal is to minimize the total length of these intervals under the condition that each 1-bit must be in one of these intervals. We give an efficient greedy algorithm which takes time O(log k) per update (an update involves converting a 0-bit to a 1-bit), which is independent of the size of the entire string. We prove that this greedy algorithm is 2-competitive. We use a natural linear programming relaxation for this problem, and analyze the algorithm by finding a dual feasible solution whose value matches the cost of the greedy algorithm. Amit Kumar 0001, Preeti Ranjan Panda, Smruti R. Sarangi |
FSTTCS | 3 |
| 2010 | DUST: a generalized notion of similarity between uncertain time seriesabstractLarge-scale sensor deployments and an increased use of privacy-preserving transformations have led to an increasing interest in mining uncertain time series data. Traditional distance measures such as Euclidean distance or dynamic time warping are not always effective for analyzing uncertain time series data. Recently, some measures have been proposed to account for uncertainty in time series data. However, we show in this paper that their applicability is limited. In specific, these approaches do not provide an intuitive way to compare two uncertain time series and do not easily accommodate multiple error functions. Smruti R. Sarangi, Karin Murthy |
KDD | 1 |
| 2008 | EVAL: Utilizing processors with variation-induced timing errorsabstractParameter variation in integrated circuits causes sections of a chip to be slower than others. If, to prevent any resulting timing errors, we design processors for worst-case parameter values, we may lose substantial performance. An alternate approach explored in this paper is to design for closer to nominal values, and provide some transistor budget to tolerate unavoidable variation-induced errors. To assess this approach, this paper first presents a novel framework that shows how microarchitecture techniques can trade off variation-induced errors for power and processor frequency. Then, the paper introduces an effective technique to maximize performance and minimize power in the presence of variation-induced errors, namely High-Dimensional dynamic adaptation. For efficiency, the technique is implemented using a machine-learning algorithm. The results show that our best configuration increases processor frequency by 56% on average, allowing the processor to cycle 21% faster than without variation. Processor performance increases by 40% on average, resulting in a performance that is 14% higher than without variation - at only a 10.6% area cost. Smruti R. Sarangi, Brian Greskamp, Abhishek Tiwari 0002, Josep Torrellas |
MICRO | 1 |
| 2007 | ReCycle: : pipeline adaptation to tolerate process variationabstractProcess variation affects processor pipelines by making some stages slower and others faster, therefore exacerbating pipeline unbalance. This reduces the frequency attainable by the pipeline. To improve performance, this paper proposes ReCycle, an architectural framework that comprehensively applies cycle time stealing to the pipeline - transferring the time slack of the faster stages to the slow ones by skewing clock arrival times to latching elements after fabrication. As a result, the pipeline can be clocked with a period equal to the average stage delay rather than the longest one. In addition, ReCycle's frequency gains are enhanced with Donor stages, which are empty stages added to "donate" slack to the slow stages. Finally, ReCycle can also convert slack into power reductions. Abhishek Tiwari 0002, Smruti R. Sarangi, Josep Torrellas |
ISCA | 2 |
| 2007 | Threshold Voltage Variation Effects on Aging-Related Hard Failure RatesabstractThis paper quantifies the impact of threshold voltage variation on aging-related hard failure rates in a high-performance 65nm processor. Simulations show that threshold voltage variations can accelerate aging substantially, depending on the thermal resistance of the heatsink and the total leakage power of the processor before variation. For unfavorable values of these parameters, our models suggest that the time at which 1% of the processors have failed can decrease by about 60%. Brian Greskamp, Smruti R. Sarangi, Josep Torrellas |
ISCAS | 2 |
| 2006 | CADRE: Cycle-Accurate Deterministic Replay for Hardware DebuggingabstractOne of the main reasons for the difficulty of hardware verification is that hardware platforms are typically nondeterministic at clock-cycle granularity. Uninitialized state elements, I/O, and timing variations on high-speed buses all introduce nondeterminism that causes different behavior on different runs starting from the same initial state. To improve our ability to debug hardware, we would like to completely eliminate nondeterminism. This paper introduces the cycle-accurate deterministic replay (CADRE) architecture, which cost-effectively makes a board-level computer cycle-accurate deterministic. We characterize the sources of nondeterminism in computers and show how to address them. In particular, we introduce a novel scheme to ensure deterministic communication on source-synchronous buses that cross clock-domain boundaries. Experiments show that CADRE on a 4-way multiprocessor server enables cycle-accurate deterministic execution of one-second intervals with modest buffering requirements (around 200MB) and minimal performance loss (around 1%). Moreover, CADRE has modest hardware requirements Smruti R. Sarangi, Brian Greskamp, Josep Torrellas |
DSN | 1 |
| 2006 | Phoenix: Detecting and Recovering from Permanent Processor Design Bugs with Programmable HardwareabstractAlthough processor design verification consumes ever-increasing resources, many design defects still slip into production silicon. In a few cases, such bugs have caused expensive chip recalls. To truly improve productivity, hardware bugs should be handled like system software ones, with vendors periodically releasing patches to fix hardware in the field. Based on an analysis of serious design defects in current AMD, Intel, IBM, and Motorola processors, this paper proposes and evaluates Phoenix - novel field-programmable on-chip hardware that detects and recovers from design defects. Phoenix taps key logic signals and, based on downloaded defect signatures, combines the signals into conditions that flag defects. On defect detection, Phoenix flushes the pipeline and either retries or invokes a customized recovery handler. Phoenix induces negligible slowdown, while adding only 0.05% area and 0.48% wire overheads. Phoenix detects all the serious defects that are triggered by concurrent control signals. Moreover, it recovers from most of them, and simplifies recovery for the rest. Finally, we present an algorithm to automatically size Phoenix for new processors Smruti R. Sarangi, Abhishek Tiwari 0002, Josep Torrellas |
MICRO | 1 |
| 2005 | Thread-Level Speculation on a CMP can be energy efficientabstractChip Multiprocessors (CMP) with Thread-Level Speculation (TLS) have become the subject of intense research. However, TLS is suspected of being too energy inefficient to compete against conventional processors. In this paper, we refute this claim. To do so, we first identify the main sources of dynamic energy consumption in TLS. Then, we present simple energy-saving optimizations that cut the energy cost of TLS by over 60% on average with minimal performance impact. The resulting TLS CMP, populated with four 3-issue cores, speeds-up full SPECint 2000 codes by 1.27 on average, while keeping the fraction of the chip's energy consumption due to TLS to only 20%. Compared to a 6-issue superscalar at the same frequency, the TLS CMP is on average faster, while consuming only 85% of its total on-chip power. Jose Renau, Karin Strauss, Luis Ceze, Wei Liu 0014, Smruti R. Sarangi, James Tuck 0001, Josep Torrellas |
ICS | 5 |
| 2005 | ReSlice: Selective Re-Execution of Long-Retired Misspeculated Instructions Using Forward SlicingabstractAs more data value speculation mechanisms are being proposed to speed-up processors, there is growing pressure on the critical processor structures that must buffer the state of the speculative instructions. A scalable solution is to checkpoint the processor and retire speculative instructions. However, in this environment, misprediction recovery becomes very wasteful, as it involves discarding and re-executing all the instructions executed since the checkpoint. To speed-up execution in this environment, this paper presents a novel architecture (ReSlice) that selectively re-executes only the speculatively-retired instructions that directly depended on the mispredicted value, namely its Forward Slice. ReSlice buffers the (typically very few) instructions in the forward slice of the predicted value as such instructions initially execute. Then, potentially thousands of instructions later, ReSlice can quickly re-execute the slice if a misprediction is declared, and merge its state with the program state. In addition, this paper develops a sufficient condition for correct slice re-execution and merge. As one possible use of ReSlice, we apply it to recover from cross-task dependence violations in a chip multiprocessor with thread-level speculation (TLS). ReSlice speeds up SpecInt applications over aggressive TLS by up to 33%, with a geometric mean of 12%. Moreover, E /spl times/ D/sup 2/ decreases by 20%. All this is obtained by saving on average 61% of the task squashes through slice re-execution. On average, a slice re-executes only 6.6 instructions, compared to the 210 that would be re-executed on a squash. Smruti R. Sarangi, Wei Liu 0014, Yuanyuan Zhou 0001 |
MICRO | 1 |