Sparsh Mittal

dblp:97/9523 · DBLP profile ↗
← Back
79ranked-venue papers
35as first author
41since 2021 · last 2026
0000-0002-2908-993XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 58 · 32 first-author · 22 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Efficient On-chip Adaptation Technique for MAML models on Memristive Hardware
Aabid Amin Fida, Sparsh Mittal
ISCAS2
2026 SilentBite: A Novel LLM-based framework for Automated Hardware Trojan Insertion
Shresth Shankhdhar, Ansh Bharadwaj, Vishesh Mishra, Sparsh Mittal
ISCAS4
2026 Dressing the Imagination: A Dataset for AI-Powered Translation of Text into Fashion Outfits and A Novel NeRA Adapter for Enhanced Feature Adaptation
abstract
Specialized datasets that capture the fashion industry’s rich language and styling elements can boost progress in AI-driven fashion design. We present FLORA (Fashion Language Outfit Representation for Apparel Generation), the first comprehensive dataset containing 4,330 curated pairs of fashion outfits and corresponding textual descriptions. Each description utilizes industry-specific terminology and jargon commonly used by professional fashion designers, providing precise and detailed insights into the outfits. Hence, the dataset captures the delicate features and subtle stylistic elements necessary to create high-fidelity fashion designs. We demonstrate that fine-tuning generative models on the FLORA dataset significantly enhances their capability to generate accurate and stylistically rich images from textual descriptions of fashion sketches. FLORA will catalyze the creation of advanced AI models capable of comprehending and producing subtle, stylistically rich fashion designs. It will also help fashion designers and end-users to bring their ideas to life. As a second orthogonal contribution, we introduce NeRA (Nonlinear low-rank Expressive Representation Adapter), a novel adapter architecture based on Kolmogorov-Arnold Networks (KAN). Unlike traditional PEFT techniques such as LoRA, LoKR, DoRA, and LoHA that use MLP adapters, NeRA uses learnable spline-based nonlinear transformations, enabling superior modeling of complex semantic relationships, achieving strong fidelity, faster convergence and semantic alignment. Extensive experiments on our proposed FLORA and LAION-5B datasets validate the superiority of NeRA over existing adapters. See Project page for data and code.
Gayatri Deshmukh, Somsubhra De, Chirag Sehgal, Jishu Sen Gupta, Sparsh Mittal
WACV5
2026 Pyramidal Spectrum: Frequency-based Hierarchically Vector Quantized VAE for Videos
abstract
Variational Autoencoders (VAEs) form the foundation of modern video generation models. In particular, discrete latent VAEs with vector quantization have gained prominence for their superior perceptual sharpness, ability to model long-range dynamics, and efficient adaptation to downstream tasks. However, existing discrete VAEs face two key limitations: (i) a lack of frequency-domain modeling to enhance global spatiotemporal understanding, and (ii) fixed-resolution quantization schemes, preventing effective modeling of coarse-to-fine spatiotemporal hierarchies essential for video generation. To address these limitations, we propose a Pyramidal Vector Quantized Variational Autoencoder (PVQ-VAE) for videos. PVQ-VAE’s encoder–decoder leverages Fast Fourier Transform and Discrete Wavelet Transform to capture global semantics and multi-scale local details jointly. We introduce Pyramidal Vector Quantization (PVQ), a hierarchical quantization scheme that discretizes features at multiple resolutions to better capture multi-scale information. To further boost fidelity, we introduce a cross-modal contrastive loss guided by a pretrained high-resolution image VAE. PVQ-VAE achieves state-of-the-art performance on WebVid-val, COCO-val, and MCL-JCV, reconstructing videos with high perceptual quality at up to 32× spatial and 16× temporal compression. Project page of PVQ-VAE.
Tushar Prakash, Onkar Susladkar, Inderjit S. Dhillon, Sparsh Mittal
WACV4
2026 Confidence Through Parallel Attention for Depth and Uncertainty Estimation in Dynamic Environments
abstract
Monocular depth estimation is crucial for robotics, offering a lightweight and scalable alternative to stereo or LiDAR-based systems. While recent methods have achieved high accuracy, their efficacy degrades under real-world conditions such as occlusion, and domain shifts. We introduce ConFiDeNet, a unified framework that jointly predicts metric depth and associated aleatoric uncertainty, enabling risk-aware robotic perception. ConFiDeNet employs a lightweight parallel attention module that efficiently fuses semantic cues from DINOv2 dense descriptors and SAM2-based segmentation for densely occluded objects, enhancing structural understanding without sacrificing real-time performance. Further, we explicitly condition the model on environment type, improving generalization across diverse indoor and outdoor scenes without retraining. Our method achieves state-of-the-art results across six datasets under both supervised and zero-shot settings, outperforming nine prior techniques, including Marigold, ZoeDepth, PatchFusion, and MonoProb. With significantly faster inference and high prediction confidence, ConFiDeNet is readily deployable for embodied AI, self-driving applications, and robotic manipulation tasks. Keywords: Monocular depth estimation, Feature fusion, Robotic Vision, point cloud estimation. For code and weights Check the project page here
Onkar Susladkar, Rohit Pawar, Chirag Sehgal, Samaksh Ujjawal, Sparsh Mittal
WACV5
2026 Dual-source attention-guided frequency-spatial ensemble attacks for cross-architecture adversarial transferability
Rohit Singh Nitwal, Sparsh Mittal, Manav Aggarwal
J. Syst. Archit.2
2026 A 252.8-GOPS and 10.1-TOPS/W Split DP-8T SRAM-Based Analog CIM Macro for 9-bit Signed MAC Operations With High Signal Margin
abstract
We present a reconfigurable and scalable current-based analog compute-in-memory (CIM) macro based on a split dual-port 8T (DP-8T) SRAM bitcell. This bitcell can perform MAC operations between signed 9-bit inputs and signed 9-bit weights and logical operations. Throughput is enhanced by sensing MAC and logical operations from both true and complementary bitcell sides via dual bitline sensing, without requiring external circuitry to decode the MAC output from the complementary side. To improve the signal margin and reduce the number of ADCs, we split the 8-bit magnitudes of inputs and weights into 2-bit groups and perform efficient weight encoding with a 2C–1C network combined with analog shift-and-add using 4C and 1C at each bitline. The proposed$64\times 128$macro performs 4096 signed MAC operations in 4 cycles with a latency of 16.2 ns. The design is implemented in TSMC 65-nm technology at 1.2 V. The architecture performs linear MAC operations across different inputs, weights, and process corners. It delivers a throughput of 252.83 GOPS, an MAC energy efficiency of 10.11 TOPS/W, and a signal margin of 26 mV for signed MAC operations. In the logical compute mode, the architecture performsnor,and,nand, andorBoolean operations in a single cycle, as well asexnorandexorusing additional logical gates, achieving a throughput of 3276.8 GOPS and a latency of 1.25 ns at 0.8 V. The work achieves inference accuracies of 99.1%, 91.65%, and 71.8% on MNIST, CIFAR-10, and CIFAR-100, respectively.
Abhishek Goel, Cheena Singhal, Sparsh Mittal, Sudeb Dasgupta
IEEE Trans. Very Large Scale Integr. Syst.3
2025 A 101 TOPS/W and 1.73 TOPS/mm2 6T SRAM-Based Digital Compute-in-Memory Macro Featuring a Novel 2T Multiplier
abstract
In this paper, we propose a 6T SRAM-based all-digital Compute-in-memory (CIM) macro for multi-bit multiply-and-accumulate (MAC) operations. We propose a novel 2T bitwise multiplier, which is a direct improvement over the previously proposed 4T NOR gate-based multiplier. The 2T multiplier also eliminates the need to invert the input bits, which is required when using NOR gates for multipliers. We propose an efficient digital MAC computation flow based on a barrel shifter, which significantly reduces the latency of shift operation. This brings down the overall latency incurred while performing MAC operations to 13ns/25ns (in 65nm CMOS)for 4b/8b operands (in 65nm CMOS @ 0.6V), compared to 10ns/18ns (in 22nm CMOS @ 0.72V) of the previous work. The proposed CIM macro is fully re-configurable in weight bits (4/8/12/16) and input (4/8) bits. It can perform concurrent MAC and weight update operations. Moreover, its fully complete digital implementation circumvents the challenges associated with analog CIM macros. For MAC operation with 4b weight and input, the macro achieves 24 TOPS/W at 1.2 V and 81 TOPS/W at 0.7 V. When using low-threshold-voltage transistors in the 2T multiplier, the macro works reliably even at 0.6V while achieving 101 TOPS/W.
Priyanshu Tyagi, Sparsh Mittal
DATE2
2025 Historic Scripts to Modern Vision: A Novel Dataset and A VLM Framework for Transliteration of Modi Script to Devanagari
Harshal Kausadikar, Tanvi Kale, Onkar Susladkar, Sparsh Mittal
ICDAR (5)4
2025 MotionAura: Generating High-Quality and Motion Consistent Videos using Discrete Diffusion
abstract
The spatio-temporal complexity of video data presents significant challenges in tasks such as compression, generation, and inpainting. We present four key contributions to address the challenges of spatiotemporal video processing. First, we introduce the 3D Mobile Inverted Vector-Quantization Variational Autoencoder (3D-MBQ-VAE), which combines Variational Autoencoders (VAEs) with masked modeling to enhance spatiotemporal video compression. The model achieves superior temporal consistency and state-of-the-art (SOTA) reconstruction quality by employing a novel training strategy with full frame masking. Second, we present MotionAura, a text-to-video generation framework that utilizes vector-quantized diffusion models to discretize the latent space and capture complex motion dynamics, producing temporally coherent videos aligned with text prompts. Third, we propose a spectral transformer-based denoising network that processes video data in the frequency domain using the Fourier Transform. This method effectively captures global context and long-range dependencies for high-quality video generation and denoising. Lastly, we introduce a downstream task of Sketch Guided Video Inpainting. This task leverages Low-Rank Adaptation (LoRA) for parameter-efficient fine-tuning. Our models achieve SOTA performance on a range of benchmarks. Our work offers robust frameworks for spatiotemporal modeling and user-driven video content manipulation.
Onkar Susladkar, Jishu Sen Gupta, Chirag Sehgal, Sparsh Mittal, Rekha Singhal
ICLR4
2025 CADAEC: Content-Aware Deployment of AI Workloads in Edge-Cloud Ecosystem
abstract
The rapid growth of edge devices has revolutionized industrial AI applications, including robotics, autonomous systems, and IoT, where real-time processing is essential. These systems face the challenge of managing concurrent, high-volume workloads across resource-constrained edge devices and cloud infrastructure. A major hurdle is optimizing deep learning model deployment across edge-cloud environments in dynamic conditions, particularly where input quality and noise fluctuate under concurrent demands. This paper introduces a novel optimization framework, that addressed these challenges and dynamically selects the most suitable models from a diverse model zoo and determines optimal deployment locations (edge or cloud). The proposed framework leverages content-aware approach to minimize both communication and computation latency while considering hardware limitations and environmental factors. Using a binary linear programming (BILP) approach, our method efficiently balances model distribution of an AI pipeline, maximizing end-to-end performance. We validate this framework on a robotic AI pipeline in real-world, noise-variant environments, comparing content-aware and content-agnostic deployment strategies. Our results demonstrate significant optimization in deployment latency and system performance under high-concurrency conditions, using both content-agnostic and content-aware approaches, highlighting the framework's robustness and scalability. Additionally, we showed the effectiveness of the content-aware approach over the content-agnostic method in optimizing deployment choices and reducing latency, while maintaining the desired qualitative outcomes of the AI pipeline with different communication set up. This makes the content-aware strategy more suitable for complex, real-world environments where input quality and noise vary significantly. Overall, The proposed method presents a compelling solution for optimizing AI pipelines in edge-cloud ecosystems, offering potential for broader applications domains.
Ratul Kishore Saha, Sparsh Mittal, Rekha Singhal, Manoj Nambiar 0001
ICPE2
2025 Novel hybrid probabilistic-statistical error metrics for approximate adders
Vishesh Mishra, Sparsh Mittal, Urbi Chatterjee
J. Syst. Archit.2
2025 SATGuard: SAT-driven Countermeasures for Protecting Approximate Circuits from Hardware Trojan
abstract
Approximate arithmetic circuits have gained prominence in modern computing systems due to their ability to trade accuracy for improved performance and energy efficiency. However, their susceptibility to stealthy Trojan attacks poses a significant security concern. This work analyzes Trojan attacks on approximate circuits, focusing specifically on approximate adders and multipliers. We propose SATGuard, a boolean satisfiability (SAT)-based methodology to identify Trojan activating inputs (TAIs) for all approximate adder and multiplier families. We also claim that TAIs for approximate circuits are analogous to test input patterns for accurate circuits. Subsequently, we propose design-specific countermeasures to safeguard approximate circuits. The proposed countermeasures nullify the Hardware Trojan Horse (HTH)-based accuracy degradation, thus upholding the application-level accuracy requirements. We conduct experiments where potential Trojans are implanted into various approximate adders and multipliers. We evaluate their impact on the error metrics and the quality of results in real-world applications such as image processing and deep neural networks (DNNs). Our findings demonstrate that the proposed methodology successfully reverses the HTH-based accuracy degradation by 99.4%, and 99.8% in approximate adders and multipliers, respectively. This improvement is achieved with an average area overhead of 5.3% and a power-delay-product overhead of 7.6% in approximate adders and 1.7% and 1.9% in multipliers, respectively.
Vishesh Mishra, Dipesh, Sparsh Mittal, Urbi Chatterjee
ACM Trans. Embed. Comput. Syst.3
2024 GRIZAL: Generative Prior-guided Zero-Shot Temporal Action Localization
abstract
Zero-shot temporal action localization (TAL) aims to temporally localize actions in videos without prior training examples.To address the challenges of TAL, we offer GRIZAL, a model that uses multimodal embeddings and dynamic motion cues to localize actions effectively.GRIZAL achieves sample diversity by using large-scale generative models such as GPT-4 for generating textual augmentations and DALL-E for generating image augmentations.Our model integrates vision-language embeddings with optical flow insights, optimized through a blend of supervised and self-supervised loss functions.On Activi-tyNet, Thumos14 and Charades-STA datasets, GRIZAL vastly outperforms state-of-the-art zero-shot TAL models, demonstrating its robustness and adaptability across a wide range of video content.The code and models are available on https://github.com/CandleLabAI/ GRIZAL-EMNLP2024.
Onkar Susladkar, Gayatri Deshmukh, Vandan Gorade, Sparsh Mittal
EMNLP4
2024 Harmonized Spatial and Spectral Learning for Generalized Medical Image Segmentation
Vandan Gorade, Sparsh Mittal, Debesh Jha, Rekha Singhal, Ulas Bagci
ICPR (13)2
2024 D2Styler: Advancing Arbitrary Style Transfer with Discrete Diffusion Methods
Onkar Susladkar, Gayatri Deshmukh, Sparsh Mittal, Parth Shastri
ICPR (6)3
2024 Textual Alchemy: CoFormer for Scene Text Understanding
abstract
The paper presents CoFormer (Convolutional Fourier Transformer), a robust and adaptable transformer architecture designed for a range of scene text tasks. CoFormer integrates convolution and Fourier operations into the transformer architecture. Thus, it leverages convolution properties such as shared weights, local receptive fields, and spatial subsampling, while the Fourier operation emphasizes composite characteristics from the frequency domain. The research further proposes two new pretraining datasets, named Textverse10M-E and Textverse10M-H. Using these datasets, we demonstrate the efficacy of pretraining for scene text understanding. CoFormer achieves state-of-the-art results with and without pretraining on two downstream tasks: scene text recognition (STR) and scene text editing (STE). The paper further proposes LISTNet (Language Invariant Style Transfer), a novel framework for bi-lingual STE. It also introduces three datasets, viz., TST500K for STE, CSTR2.5M and Akshara550 for STR. The source-code of CoFormer is available at https://github.com/CandleLabAI/CoFormer-WACV-2024.
Gayatri Deshmukh, Onkar Susladkar, Dhruv Makwana, Sparsh Mittal, R. Sai Chandra Teja
WACV4
2024 SynergyNet: Bridging the Gap between Discrete and Continuous Representations for Precise Medical Image Segmentation
abstract
In recent years, continuous latent space (CLS) and discrete latent space (DLS) deep learning models have been proposed for medical image analysis for improved performance. However, these models encounter distinct challenges. CLS models capture intricate details but often lack interpretability in terms of structural representation and robustness due to their emphasis on low-level features. Conversely, DLS models offer interpretability, robustness, and the ability to capture coarse-grained information thanks to their structured latent space. However, DLS models have limited efficacy in capturing fine-grained details. To address the limitations of both DLS and CLS models, we propose SynergyNet, a novel bottleneck architecture designed to enhance existing encoder-decoder segmentation frameworks. SynergyNet seamlessly integrates discrete and continuous representations to harness complementary information and successfully preserves both fine and coarse-grained details in the learned representations. Our extensive experiment on multi-organ segmentation and cardiac datasets demonstrates that SynergyNet outperforms other state of the art methods including TransUNet: dice scores improving by 2.16%, and Hausdorff scores improving by 11.13%, respectively. When evaluating skin lesion and brain tumor segmentation datasets, we observe a remarkable improvement of 1.71% in Intersection-over-Union scores for skin lesion segmentation and of 8.58% for brain tumor segmentation. Our innovative approach paves the way for enhancing the overall performance and capabilities of deep learning models in the critical domain of medical image analysis.
Vandan Gorade, Sparsh Mittal, Debesh Jha, Ulas Bagci
WACV2
2024 LIVENet: A novel network for real-world low-light image denoising and enhancement
abstract
Low-light image enhancement (LLIE) is the process of improving the quality of images taken in low-light conditions while striking a balance between enhancing image illumination and maintaining their natural appearance. This involves reducing noise, enhancing details, and correcting colors, all while avoiding artifacts such as halo effects or color distortions. We propose LIVENet, a novel deep neural network that jointly performs noise reduction on lowlight images and enhances illumination and texture details. LIVENet has two stages: the image enhancement stage and the refinement stage. For the image enhancement stage, we propose a Latent Subspace Denoising Block (LSDB) that uses a low-rank representation of low-light features to suppress the noise and predict a noise-free grayscale image. We propose enhancing an RGB image by eliminating noise. This is done by converting it into YCbCr color space and replacing the noisy luminance (Y) channel with the predicted noise-free grayscale image. LIVENet also predicts the transmission map and atmospheric light in the image enhancement stage. LIVENet produces an enhanced image with rich color and illumination by feeding them to an atmospheric scattering model. In the refinement stage, the texture information from the grayscale image is incorporated into the improved image using a Spatial Feature Transform (SFT) layer. Experiments on different datasets demonstrate that LIVENet’s enhanced images consistently outperform previous techniques across various quality metrics. The source code can be obtained from https://github.com/CandleLabAI/LiveNet.
Dhruv Makwana, Gayatri Deshmukh, Onkar Susladkar, Sparsh Mittal, R. Sai Chandra Teja
WACV4
2023 RAxC: Reflexivity-based Approximate Computing techniques for efficient remote sensing
abstract
Hyperspectral images (HSI) have a huge size, which makes their processing through neural networks cumbersome. We propose novel approximate computing techniques that leverage physical properties of the reflectance spectra to accelerate the processing of HSI images. This makes the images interpretable across various applications. We propose three spectral dimensionality reduction techniques. These techniques use spectral clustering methods that rely on reflectance values to capture inherent characteristics from hyperspectral images across diverse domains. We also evaluate existing spatial dimension reduction techniques and a combination of spatial + spectral dimension reduction techniques. We conduct extensive experiments on three real-world open-source datasets, encompassing urban and rural landscapes. Our techniques reduce the training time by up to 8x and inference time by up to 5x, while reducing the model size by up to 3x. Our techniques have a negligible impact on accuracy. By contrast, PCA and MNF techniques incur 3X higher pre-processing latency overheads than our techniques and also degrade the accuracy. Our techniques are promising for addressing the computational challenges of HSI processing.
Aaditi Kapre, Shruti Kunde, Sparsh Mittal, Rekha Singhal
IEEE Big Data3
2023 TPFNet: A Novel Text In-painting Transformer for Text Removal
Onkar Susladkar, Dhruv Makwana, Gayatri Deshmukh, Sparsh Mittal, R. Sai Chandra Teja, Rekha Singhal
ICDAR (6)4
2023 GAFNet: A Global Fourier Self Attention Based Novel Network for multi-modal downstream tasks
abstract
In "vision and language" problems, multimodal inputs are simultaneously processed for combined visual and textual understanding for image-text embedding. In this paper, we discuss the necessity of considering the difference between the feature space and the distribution when performing multimodal learning. We deal with this problem through deep learning and a generative model approach. We introduce a novel network, GAFNet (Global Attention Fourier Net), which learns through large-scale pre-training over three image-text datasets (COCO, SBU, and CC-3M), for achieving high performance on downstream vision and language tasks. We propose a GAF (Global Attention Fourier) module, which integrates multiple modalities into one latent space. GAF module is independent of the type of modality, and it allows combining shared representations at each stage. Various ways of thinking about the relationships between different modalities directly affect the model’s design. In contrast to previous research, our work considers visual grounding as a pretrainable and transferable quality instead of something that must be trained from scratch. We show that GAFNet is a versatile network that can be used for a wide range of downstream tasks. Experimental results demonstrate that our technique achieves state-of-the-art performance on multimodal classification on the CrisisMD dataset and image generation on the COCO dataset. For image-text retrieval, our technique achieves competitive performance.
Onkar Susladkar, Gayatri Deshmukh, Dhruv Makwana, Sparsh Mittal, R. Sai Chandra Teja, Rekha Singhal
WACV4
2023 PCBSegClassNet - A light-weight network for segmentation and classification of PCB component
Dhruv Makwana, R. Sai Chandra Teja, Sparsh Mittal
Expert Syst. Appl.3
2023 An active memristor based rate-coded spiking neural network
Aabid Amin Fida, Farooq Ahmad Khanday, Sparsh Mittal
Neurocomputing3
2023 A survey of techniques for optimizing transformer inference
Krishna Teja Chitty-Venkata, Sparsh Mittal, Murali Emani, Venkatram Vishwanath, Arun K. Somani
J. Syst. Archit.2
2023 At the Locus of Performance: Quantifying the Effects of Copious 3D-Stacked Cache on HPC Workloads
abstract
Over the last three decades, innovations in the memory subsystem were primarily targeted at overcoming the data movement bottleneck. In this paper, we focus on a specific market trend in memory technology: 3D-stacked memory and caches. We investigate the impact of extending the on-chip memory capabilities in future HPC-focused processors, particularly by 3D-stacked SRAM. First, we propose a method oblivious to the memory subsystem to gauge the upper-bound in performance improvements when data movement costs are eliminated. Then, using the gem5 simulator, we model two variants of a hypothetical LARge Cache processor (LARC), fabricated in 1.5 nm and enriched with high-capacity 3D-stacked cache. With a volume of experiments involving a broad set of proxy-applications and benchmarks, we aim to reveal how HPC CPU performance will evolve, and conclude an average boost of 9.56× for cache-sensitive HPC applications, on a per-chip basis. Additionally, we exhaustively document our methodological exploration to motivate HPC centers to drive their own technological agenda through enhanced co-design.
Jens Domke, Emil Vatai, Balazs Gerofi, Yuetsu Kodama, Mohamed Wahib, Artur Podobas, Sparsh Mittal, Miquel Pericàs, Lingqi Zhang 0001, Peng Chen 0035, Aleksandr Drozd, Satoshi Matsuoka
ACM Trans. Archit. Code Optim.7
2023 VADF: Versatile Approximate Data Formats for Energy-Efficient Computing
abstract
Approximate computing (AC) techniques provide overall performance gains in terms of power and energy savings at the cost of minor loss in application accuracy. For this reason, AC has emerged as a viable method for efficiently supporting several compute-intensive applications, e.g., machine learning, deep learning, and image processing, that can tolerate bounded errors in computations. However, most prior techniques do not consider the possibility of soft errors or malicious bit-flips in AC systems. These errors may interact with approximation-introduced errors in unforeseen ways, leading to disastrous consequences, such as the failure of computing systems. A recent research effort, FTApprox (DATE’21) proposes an error-resilient approximate data format. FTApprox stores two blocks, starting from the one containing the most significant valid (MSV) bit. It also stores location of the MSV block and protects them using error-correcting bits (ECBs). However, FTApprox has crucial limitations such as lack of flexibility, redundantly storing zeros in the MSV, etc. In this paper, we propose a novel storage format named Versatile Approximate Data Format (VADF) for storing approximate integer numbers while providing resilience to soft errors. VADF prescribes rules for storing, for example, a 32-bit number in either 8-bit, 12-bit or 16-bit numbers. VADF identifies the MSV bit and stores a certain number of bits following the MSV bit. It also stores the location of the MSV bit and protects it by ECBs. VADF does not explicitly store the MSB bit itself and this prevents VADF from accruing significant errors. VADF incurs lower error than both truncation methodologies and FTApprox. We further evaluate five image-processing and machine-learning applications and confirm that VADF provides higher application quality than FTApprox in the presence and absence of soft errors. Finally, VADF allows the use of narrow arithmetic units. For example, instead of using a 32-bit multiplier/adder, one can first use VADF (or FTApprox) to compress the data and then use a 8-bit multiplier/adder. Through this approach, VADF facilitates 95.97% and 79.3% energy savings in multiplication and addition, respectively. However, the subsequent re-conversion of the 8-bit output data to 32-bit data using Inv-VADF(16,3,32) diminishes the energy savings by 9.6% for addition and 0.56% for multiplication operation, respectively. The code is available at https://github.com/CandleLabAI/VADF-ApproximateDataFormat-TECS .
Vishesh Mishra, Sparsh Mittal, Neelofar Hassan, Rekha Singhal, Urbi Chatterjee
ACM Trans. Embed. Comput. Syst.2
2023 A Survey of Deep Learning Techniques for Underwater Image Classification
abstract
In recent years, there has been an enormous interest in using deep learning to classify underwater images to identify various objects, such as fishes, plankton, coral reefs, seagrass, submarines, and gestures of sea divers. This classification is essential for measuring the water bodies' health and quality and protecting the endangered species. Furthermore, it has applications in oceanography, marine economy and defense, environment protection, underwater exploration, and human-robot collaborative tasks. This article presents a survey of deep learning techniques for performing underwater image classification. We underscore the similarities and differences of several methods. We believe that underwater image classification is one of the killer application that would test the ultimate success of deep learning techniques. Toward realizing that goal, this survey seeks to inform researchers about state-of-the-art on deep learning on underwater images and also motivate them to push its frontiers forward.
Sparsh Mittal, Srishti Srivastava 0001, J. Phani Jayanth
IEEE Trans. Neural Networks Learn. Syst.1
2022 MEGA-MAC: A Merged Accumulation based Approximate MAC Unit for Error Resilient Applications
abstract
This paper proposes a novel merged-accumulation-based approximate MAC (multiply-accumulate) unit, MEGA-MAC, for accelerating error-resilient applications. MEGA-MAC utilizes a novel rearrangement and compression strategy in the multiplication stage and a novel approximate "carry predicting adder" (CPA) in the accumulation stage. Addition and multiplication operations are merged, which reduces the delay. MEGA-MAC provides knobs to exercise a tradeoff between accuracy and resource overhead. Compared to the accurate MAC unit, MEGA-MAC(8,6) (i.e., a MEGA-MAC unit with a chunk size of 6 bits, operating on 8-bit input operands) reduces the power-delay-product (PDP) by 49.4%, while incurring a mean error percentage of only 4.2%. Compared to state-of-art approximate MAC units, MEGA-MAC achieves a better balance between resource-saving and accuracy-loss. The source code is available at https://sites.google.com/view/mega-mac-approximate-mac-unit/.
Vishesh Mishra, Sparsh Mittal, Divy Pandey, Rekha Singhal
ACM Great Lakes Symposium on VLSI2
2022 ClarifyNet: A high-pass and low-pass filtering based CNN for single image dehazing
Onkar Susladkar, Gayatri Deshmukh, Subhrajit Nag, Ananya Mantravadi, Dhruv Makwana, Sujitha Ravichandran, R. Sai Chandra Teja, Gajanan H. Chavhan, C. Krishna Mohan, Sparsh Mittal
J. Syst. Archit.10
2022 CORIDOR: Using COherence and TempoRal LocalIty to Mitigate Read Disurbance ErrOR in STT-RAM Caches
abstract
In the deep sub-micron region, “spin-transfer torque RAM” (STT-RAM ) suffers from “read-disturbance error” (RDE) , whereby a read operation disturbs the stored data. Mitigation of RDE requires restore operations, which imposes latency and energy penalties. Hence, RDE presents a crucial threat to the scaling of STT-RAM. In this paper, we offer three techniques to reduce the restore overhead. First, we avoid the restore operations for those reads, where the block will get updated at a higher level cache in the near future. Second, we identify read-intensive blocks using a lightweight mechanism and then migrate these blocks to a small SRAM buffer. On a future read to these blocks, the restore operation is avoided. Third, for data blocks having zero value, a write operation is avoided, and only a flag is set. Based on this flag, both read and restore operations to this block are avoided. We combine these three techniques to design our final policy, named CORIDOR. Compared to a baseline policy, which performs restore operation after each read, CORIDOR achieves a 31.6% reduction in total energy and brings the relative CPI (cycle-per-instruction) to 0.64×. By contrast, an ideal RDE-free STT-RAM saves 42.7% energy and brings the relative CPI to 0.62×. Thus, our CORIDOR policy achieves nearly the same performance as an ideal RDE-free STT-RAM cache. Also, it reaches three-fourths of the energy-saving achieved by the ideal RDE-free cache. We also compare CORIDOR with four previous techniques and show that CORIDOR provides higher restore energy savings than these techniques.
Sheel Sindhu Manohar, Sparsh Mittal, Hemangee K. Kapoor
ACM Trans. Embed. Comput. Syst.2
2022 A Survey of Deep Learning on CPUs: Opportunities and Co-Optimizations
abstract
CPU is a powerful, pervasive, and indispensable platform for running deep learning (DL) workloads in systems ranging from mobile to extreme-end servers. In this article, we present a survey of techniques for optimizing DL applications on CPUs. We include the methods proposed for both inference and training and those offered in the context of mobile, desktop/server, and distributed systems. We identify the areas of strength and weaknesses of CPUs in the field of DL. This article will interest practitioners and researchers in the area of artificial intelligence, computer architecture, mobile systems, and parallel computing.
Sparsh Mittal, Poonam Rajput, Sreenivas Subramoney
IEEE Trans. Neural Networks Learn. Syst.1
2021 A survey on hardware security of DNN models and accelerators
Sparsh Mittal, Himanshi Gupta, Srishti Srivastava 0001
J. Syst. Archit.1
2021 A survey On hardware accelerators and optimization techniques for RNNs
Sparsh Mittal, Sumanth Umesh
J. Syst. Archit.1
2021 A survey of accelerator architectures for 3D convolution neural networks
Sparsh Mittal, Vibhu
J. Syst. Archit.1
2021 A survey of SRAM-based in-memory computing techniques and applications
Sparsh Mittal, Gaurav Verma 0002, Brajesh Kumar Kaushik, Farooq Ahmad Khanday
J. Syst. Archit.1
2021 CURATING: A multi-objective based pruning technique for CNNs
Santanu Pattanayak, Subhrajit Nag, Sparsh Mittal
J. Syst. Archit.3
2021 A survey of hardware architectures for generative adversarial networks
Nivedita Shrivastava, Muhammad Abdullah Hanif, Sparsh Mittal, Smruti R. Sarangi, Muhammad Shafique 0001
J. Syst. Archit.3
2021 A survey of deep learning techniques for vehicle detection from UAV images
Srishti Srivastava 0001, Sarthak Narayan, Sparsh Mittal
J. Syst. Archit.3
2021 A survey of techniques for intermittent computing
Sumanth Umesh, Sparsh Mittal
J. Syst. Archit.2
2021 Modeling Data Reuse in Deep Neural Networks by Taking Data-Types into Cognizance
abstract
In recent years, researchers have focused on reducing the model size and number of computations (measured as “multiply-accumulate” or MAC operations) of DNNs. The energy consumption of a DNN depends on both the number of MAC operations and the energy efficiency of each MAC operation. The former can be estimated at design time; however, the latter depends on the intricate data reuse patterns and underlying hardware architecture. Hence, estimating it at design time is challenging. This article shows that the conventional approach to estimate the data reuse, viz. arithmetic intensity, does not always correctly estimate the degree of data reuse in DNNs since it gives equal importance to all the data types. We propose a novel model, termed “data type aware weighted arithmetic intensity” (DI), which accounts for the unequal importance of different data types in DNNs. We evaluate our model on 25 state-of-the-art DNNs on two GPUs. We show that our model accurately models data-reuse for all possible data reuse patterns for different types of convolution and different types of layers. We show that our model is a better indicator of the energy efficiency of DNNs. We also show its generality using the central limit theorem.
Nandan Kumar Jha, Sparsh Mittal
IEEE Trans. Computers2
2020 ULSAM: Ultra-Lightweight Subspace Attention Module for Compact Convolutional Neural Networks
abstract
The capability of the self-attention mechanism to model the long-range dependencies has catapulted its deployment in vision models. Unlike convolution operators, self-attention offers infinite receptive field and enables compute- efficient modeling of global dependencies. However, the existing state-of-the-art attention mechanisms incur high compute and/or parameter overheads, and hence unfit for compact convolutional neural networks (CNNs). In this work, we propose a simple yet effective "Ultra-Lightweight Subspace Attention Mechanism" (ULSAM), which infers different attention maps for each feature map subspace. We argue that leaning separate attention maps for each feature subspace enables multi-scale and multi-frequency feature representation, which is more desirable for fine-grained image classification. Our method of subspace attention is orthogonal and complementary to the existing state-of-the- arts attention mechanisms used in vision models. ULSAM is end-to-end trainable and can be deployed as a plug-and- play module in the pre-existing compact CNNs. Notably, our work is the first attempt that uses a subspace attention mechanism to increase the efficiency of compact CNNs. To show the efficacy of ULSAM, we perform experiments with MobileNet-V1 and MobileNet-V2 as backbone architectures on ImageNet-1K and three fine-grained image classification datasets. We achieve ≈13% and ≈25% reduction in both the FLOPs and parameter counts of MobileNet-V2 with a 0.27% and more than 1% improvement in top-1 accuracy on the ImageNet-1K and fine-grained image classification datasets (respectively). Code and trained models are available at https://github.com/Nandan91/ULSAM.
Rajat Saini, Nandan Kumar Jha, Bedanta Das, Sparsh Mittal, C. Krishna Mohan
WACV4
2020 A survey on evaluating and optimizing performance of Intel Xeon Phi
abstract
Summary Intel's Xeon Phi combines the parallel processing power of a many‐core accelerator with the programming ease of CPUs. In this paper, we present a survey of works that study the architecture of Phi and use it as an accelerator for a broad range of applications. We review performance optimization strategies as well as the factors that bottleneck the performance of Phi. We also review works that perform comparison or collaborative execution of Phi with CPUs and GPUs. This paper will be useful for researchers and developers in the area of computer‐architecture and high‐performance computing.
Sparsh Mittal
Concurr. Comput. Pract. Exp.1
2020 DeepPeep: Exploiting Design Ramifications to Decipher the Architecture of Compact DNNs
abstract
The remarkable predictive performance of deep neural networks (DNNs) has led to their adoption in service domains of unprecedented scale and scope. However, the widespread adoption and growing commercialization of DNNs have underscored the importance of intellectual property (IP) protection. Devising techniques to ensure IP protection has become necessary due to the increasing trend of outsourcing the DNN computations on the untrusted accelerators in cloud-based services. The design methodologies and hyper-parameters of DNNs are crucial information, and leaking them may cause massive economic loss to the organization. Furthermore, the knowledge of DNN’s architecture can increase the success probability of an adversarial attack where an adversary perturbs the inputs and alters the prediction. In this work, we devise a two-stage attack methodology “DeepPeep,” which exploits the distinctive characteristics of design methodologies to reverse-engineer the architecture of building blocks in compact DNNs. We show the efficacy of “DeepPeep” on P100 and P4000 GPUs. Additionally, we propose intelligent design maneuvering strategies for thwarting IP theft through the DeepPeep attack and proposed “Secure MobileNet-V1.” Interestingly , compared to vanilla MobileNet-V1, secure MobileNet-V1 provides a significant reduction in inference latency (≈60%) and improvement in predictive performance (≈2%) with very low memory and computation overheads.
Nandan Kumar Jha, Sparsh Mittal, Binod Kumar 0001, Govardhan Mattela
ACM J. Emerg. Technol. Comput. Syst.2
2020 A survey on modeling and improving reliability of DNN algorithms and accelerators
Sparsh Mittal
J. Syst. Archit.1
2020 A survey of FPGA-based accelerators for convolutional neural networks
Sparsh Mittal
Neural Comput. Appl.1
2019 Address-stride assisted approximate load value prediction in GPUs
abstract
Value prediction holds the promise of significantly improving the performance and energy efficiency. However, if the values are predicted incorrectly, significant performance overheads are observed due to execution rollbacks. To address these overheads, value approximation is introduced, which leverages the observation that the rollbacks are not necessary as long as the application-level loss in quality due to value misprediction is acceptable to the user. However, in the context of Graphics Processing Units (GPUs), our evaluations show that the existing approximate value predictors are not optimal in improving the prediction accuracy as they do not consider memory request order, a key characteristic in determining the accuracy of value prediction. As a result, the overall data movement reduction benefits are capped as it is necessary to limit the percentage of predicted values (i.e., prediction coverage) for an acceptable value of application-level error.
Mohamed Assem Ibrahim, Sparsh Mittal, Adwait Jog
ICS3
2019 A survey of techniques for dynamic branch prediction
abstract
Summary Branch predictor (BP) is an essential component in modern processors since high BP accuracy can improve performance and reduce energy by decreasing the number of instructions executed on wrong‐path. However, reducing the latency and storage overhead of BP while maintaining high accuracy presents significant challenges. In this paper, we present a survey of dynamic branch prediction techniques. We classify the works based on key features to underscore their differences and similarities. We believe this paper will spark further research in this area and will be useful for computer architects, processor designers, and researchers.
Sparsh Mittal
Concurr. Comput. Pract. Exp.1
2019 A survey of techniques for improving efficiency of mobile web browsing
abstract
Summary Mobile web traffic has now surpassed the desktop web traffic and has become the primary means for service providers to reach out to the billions of end users. Due to this trend, optimization of mobile web browsing (MWB) has gained significant attention. In this paper, we present a survey of techniques for improving the efficiency of web browsing on mobile systems, proposed in the last 6‐7 years. We review the techniques from both the networking domain (eg, proxy and browser enhancements) and the processor architecture domain (eg, hardware customization and thread‐to‐core scheduling). We organize the research works based on key parameters to highlight their similarities and differences. Beyond summarizing the recent works, this survey aims to emphasize the need of architecting for MWB as the first principle, instead of retrofitting for it.
Sparsh Mittal, Venkat Mattela
Concurr. Comput. Pract. Exp.1
2019 A Survey on optimized implementation of deep learning models on the NVIDIA Jetson platform
Sparsh Mittal
J. Syst. Archit.1
2019 A survey on applications and architectural-optimizations of Micron's Automata Processor
Sparsh Mittal
J. Syst. Archit.1
2019 A survey of encoding techniques for reducing data-movement energy
Sparsh Mittal, Subhrajit Nag
J. Syst. Archit.1
2019 A survey of techniques for optimizing deep learning on GPUs
Sparsh Mittal, Shraiysh Vaishay
J. Syst. Archit.1
2019 A survey of spintronic architectures for processing-in-memory and neural networks
Sumanth Umesh, Sparsh Mittal
J. Syst. Archit.2
2018 A survey of techniques for architecting SLC/MLC/TLC hybrid Flash memory-based SSDs
abstract
Summary Flash memory–based solid‐state drives (SSDs) offer several attractive features and benefits compared to hard disk drive (HDD), such as shock resistance and better performance especially for random data access. Depending on the number of bits in each cell, Flash memory can be designed as single/multi/triple level cell (SLC/MLC/TLC), which have different performance, density, cost and write endurance characteristics. To bring the best of these together, several researchers have proposed designing SSD using hybrid SLC/MLC/TLC Flash memory. However, these SSDs also present several challenges such as buffer management, placement of hot/cold data in suitable portion, and intelligent garbage collection. Several recent techniques aim to address these challenges. In this paper, we present a survey of techniques for managing SSDs designed with SLC/MLC/TLC Flash memory. We classify the works on several axes to bring out their similarities and differences. We aim to synthesize the state‐of‐art progress in hybrid SSD management and also spark further research in this area.
Ahmed Izzat Alsalibi, Sparsh Mittal, Mohammed Azmi Al-Betar, Putra Sumari
Concurr. Comput. Pract. Exp.2
2018 A survey of techniques for improving error-resilience of DRAM
Sparsh Mittal, Maruthi Seshidhar Inukonda
J. Syst. Archit.1
2017 Building a Fast and Power Efficient Inductive Charge Pump System for 3D Stacked Phase Change Memories
abstract
Phase change memory emerges as one of the most promising alternatives to traditional DRAM in terabyte main memory constructions. 3D stacking technology further enhances the scalability and capacity of PCM. However, write bandwidth of emerging 3D stacked PCMs is seriously limited by capacitive change pump widely adopted in 2D PCM chips. In this paper, we propose an inductive charge pump system design for 3D stacked PCMs. We further present a novel ICP aware page management technique to reduce average charging latency and improve memory bandwidth in a 3D stacked PCM-based main memory. Compared to a capacitive charge pump system, our inductive charge pump and page management technique improve the CPU performance by 54%and reduce the system energy by 29%.
Lei Jiang 0001, Sparsh Mittal, Wujie Wen
ACM Great Lakes Symposium on VLSI2
2017 A survey of techniques for designing and managing CPU register file
abstract
Summary Processor register file (RF) is an important microarchitectural component used for storing operands and results of instructions. The design and operation of RF have crucial impact on the performance, energy efficiency, and reliability of the processor, and hence, several techniques have been recently proposed to manage RF in modern processors. In this paper, we present a survey of techniques for architecting and managing CPU register file. We classify the techniques across several parameters to underscore their similarities and differences. We hope that this paper will provide insights to researchers into working of RF and inspire even more efforts towards optimization of RF in next‐generation computing systems. Copyright © 2016 John Wiley & Sons, Ltd.
Sparsh Mittal
Concurr. Comput. Pract. Exp.1
2017 A survey of techniques for architecting TLBs
abstract
Summary Translation lookaside buffer (TLB) caches virtual to physical address translation information and is used in systems ranging from embedded devices to high‐end servers. Because TLB is accessed very frequently and a TLB miss is extremely costly, prudent management of TLB is important for improving performance and energy efficiency of processors. In this paper, we present a survey of techniques for architecting and managing TLBs. We characterize the techniques across several dimensions to highlight their similarities and distinctions. We believe that this paper will be useful for chip designers, computer architects, and system engineers.
Sparsh Mittal
Concurr. Comput. Pract. Exp.1
2017 A survey of value prediction techniques for leveraging value locality
abstract
Summary Value locality (VL) refers to recurrence of values in a memory structure, and value prediction (VP) refers to predicting VL and leveraging it for diverse optimizations. VP holds the promise of exceeding true‐data dependencies and provide performance and bandwidth advantages in both single‐ and multi‐threaded applications. Fully exploiting the potential of VL, however, requires addressing several challenges, such as achieving high accuracy and coverage, reducing hardware and latency overheads, etc. In this paper, we present a survey of techniques for leveraging value locality. We categorize the research works based on key parameters to provide insights and highlight similarities and differences. This paper is expected to be useful for researchers, processor architects, and chip‐designers.
Sparsh Mittal
Concurr. Comput. Pract. Exp.1
2017 A Survey of Techniques for Architecting and Managing GPU Register File
abstract
To support their massively-multithreaded architecture, GPUs use very large register file (RF) which has a capacity higher than even L1 and L2 caches. In total contrast, traditional CPUs use tiny RF and much larger caches to optimize latency. Due to these differences, along with the crucial impact of RF in determining GPU performance, novel and intelligent techniques are required for managing GPU RF. In this paper, we survey the techniques for designing and managing GPU RF. We discuss techniques related to performance, energy and reliability aspects of RF. To emphasize the similarities and differences between the techniques, we classify them along several parameters. The aim of this paper is to synthesize the state-of-art developments in RF management and also stimulate further research in this area.
Sparsh Mittal
IEEE Trans. Parallel Distributed Syst.1
2016 Reducing Soft-error Vulnerability of Caches using Data Compression
abstract
With ongoing chip miniaturization and voltage scaling, particle strike-induced soft errors present increasingly severe threat to the reliability of on-chip caches. In this paper, we present a technique to reduce the vulnerability of caches to soft-errors. Our technique uses data compression to reduce the number of vulnerable data bits in the cache and performs selective duplication of more critical data-bits to provide extra protection to them. Microarchitectural simulations have shown that our technique is effective in reducing cache vulnerability and outperforms another technique. For single and dual-core system configuration, the average reduction in cache vulnerability is 5.59X and 8.44X, respectively. Also, the implementation and performance overheads of our technique are minimal and it is useful for a broad range of workloads.
Sparsh Mittal, Jeffrey S. Vetter
ACM Great Lakes Symposium on VLSI1
2016 Algorithm-Directed Data Placement in Explicitly Managed Non-Volatile Memory
abstract
The emergence of many non-volatile memory (NVM) techniques is poised to revolutionize main memory systems because of the relatively high capacity and low lifetime power consumption of NVM. However, to avoid the typical limitation of NVM as the main memory, NVM is usually combined with DRAM to form a hybrid NVM/DRAM system to gain the benefits of each. However, this integrated memory system raises a question on how to manage data placement and movement across NVM and DRAM, which is critical for maximizing the benefits of this integration. The existing solutions have several limitations, which obstruct adoption of these solutions in the high performance computing (HPC) domain. In particular, they cannot take advantage of application semantics, thus losing critical optimization opportunities and demanding extensive hardware extensions; they implement persistent semantics for resilience purpose while suffering large performance and energy overhead. In this paper, we re-examine the current hybrid memory designs from the HPC perspective, and aim to leverage the knowledge of numerical algorithms to direct data placement. With explicit algorithm management and limited hardware support, we optimize data movement between NVM and DRAM, improve data locality, and implement a relaxed memory persistency scheme in NVM. Our work demonstrates significant benefits of integrating algorithm knowledge into the hybrid memory design to achieve multi-dimensional optimization (performance, energy, and resilience) in HPC.
Panruo Wu, Dong Li 0001, Zizhong Chen, Jeffrey S. Vetter, Sparsh Mittal
HPDC5
2016 A Survey of Architectural Techniques for Near-Threshold Computing
abstract
Energy efficiency has now become the primary obstacle in scaling the performance of all classes of computing systems. Low-voltage computing, specifically, near-threshold voltage computing (NTC), which involves operating the transistor very close to and yet above its threshold voltage, holds the promise of providing many-fold improvement in energy efficiency. However, use of NTC also presents several challenges such as increased parametric variation, failure rate, and performance loss. This article surveys several recent techniques that aim to offset these challenges for fully leveraging the potential of NTC. By classifying these techniques along several dimensions, we also highlight their similarities and differences. It is hoped that this article will provide insights into state-of-the-art NTC techniques to researchers and system designers and inspire further research in this field.
Sparsh Mittal
ACM J. Emerg. Technol. Comput. Syst.1
2016 A Survey of Techniques for Architecting Processor Components Using Domain-Wall Memory
abstract
Recent trends of increasing core-count and bandwidth/memory wall have motivated researchers to explore novel memory technologies for designing processor components such as cache, register file, shared memory, and so on. Domain-wall memory (DWM), also known as racetrack memory, is a promising emerging technology due to its non-volatility and very high density. However, use of DWM presents challenges due to characteristics of both DWM itself (e.g., requirement of shift operations, variable latency) and processor components. Recently, several techniques have been proposed to address these challenges. This article presents a survey of architectural techniques for using DWM for designing components in both CPU and GPU. We discuss techniques related to performance, energy, and reliability and also discuss works that compare DWM with other memory technologies. We also highlight the opportunities and obstacles in using DWM for designing processor components. This survey is expected to spark further research in this area and be useful for researchers, chip designers, and computer architects.
Sparsh Mittal
ACM J. Emerg. Technol. Comput. Syst.1
2016 A Survey of Techniques for Cache Locking
abstract
Cache memory, although important for boosting application performance, is also a source of execution time variability, and this makes its use difficult in systems requiring worst-case execution time (WCET) guarantees. Cache locking is a promising approach for simplifying WCET estimation and providing predictability, and hence, several commercial processors provide ability for locking cache. However, cache locking also has several disadvantages (e.g., extra misses for unlocked blocks, complex algorithms required for selection of locking contents) and hence, a careful management is required to realize the full potential of cache locking. In this article, we present a survey of techniques proposed for cache locking. We categorize the techniques into several groups to underscore their similarities and differences. We also discuss the opportunities and obstacles in using cache locking. We hope that this article will help researchers gain insight into cache locking schemes and will also stimulate further work in this area.
Sparsh Mittal
ACM Trans. Design Autom. Electr. Syst.1
2016 A Survey of Techniques for Modeling and Improving Reliability of Computing Systems
abstract
Recent trends of aggressive technology scaling have greatly exacerbated the occurrences and impact of faults in computing systems. This has made `reliability' a first-order design constraint. To address the challenges of reliability, several techniques have been proposed. This paper provides a survey of architectural techniques for improving resilience of computing systems. We especially focus on techniques proposed for microarchitectural components, such as processor registers, functional units, cache and main memory etc. In addition, we discuss techniques proposed for non-volatile memory, GPUs and 3D-stacked processors. To underscore the similarities and differences of the techniques, we classify them based on their key characteristics. We also review the metrics proposed to quantify vulnerability of processor structures. We believe that this survey will help researchers, system-architects and processor designers in gaining insights into the techniques for improving reliability of computing systems.
Sparsh Mittal, Jeffrey S. Vetter
IEEE Trans. Parallel Distributed Syst.1
2016 A Survey Of Architectural Approaches for Data Compression in Cache and Main Memory Systems
abstract
As the number of cores on a chip increases and key applications become even more data-intensive, memory systems in modern processors have to deal with increasingly large amount of data. In face of such challenges, data compression presents as a promising approach to increase effective memory system capacity and also provide performance and energy advantages. This paper presents a survey of techniques for using compression in cache and main memory systems. It also classifies the techniques based on key parameters to highlight their similarities and differences. It discusses compression in CPUs and GPUs, conventional and non-volatile memory (NVM) systems, and 2D and 3D memory systems. We hope that this survey will help the researchers in gaining insight into the potential role of compression approach in memory components of future extreme-scale systems.
Sparsh Mittal, Jeffrey S. Vetter
IEEE Trans. Parallel Distributed Syst.1
2016 A Survey of Software Techniques for Using Non-Volatile Memories for Storage and Main Memory Systems
abstract
Non-volatile memory (NVM) devices, such as Flash, phase change RAM, spin transfer torque RAM, and resistive RAM, offer several advantages and challenges when compared to conventional memory technologies, such as DRAM and magnetic hard disk drives (HDDs). In this paper, we present a survey of software techniques that have been proposed to exploit the advantages and mitigate the disadvantages of NVMs when used for designing memory systems, and, in particular, secondary storage (e.g., solid state drive) and main memory. We classify these software techniques along several dimensions to highlight their similarities and differences. Given that NVMs are growing in popularity, we believe that this survey will motivate further research in the field of software technology for NVMs.
Sparsh Mittal, Jeffrey S. Vetter
IEEE Trans. Parallel Distributed Syst.1
2016 A Survey Of Techniques for Architecting DRAM Caches
abstract
Recent trends of increasing core-count and memory/bandwidth-wall have led to major overhauls in chip architecture. In face of increasing cache capacity demands, researchers have now explored DRAM, which was conventionally considered synonymous to main memory, for designing large last level caches. Efficient integration of DRAM caches in mainstream computing systems, however, also presents several challenges and several recent techniques have been proposed to address them. In this paper, we present a survey of techniques for architecting DRAM caches. Also, by classifying these techniques across several dimensions, we underscore their similarities and differences. We believe that this paper will be very helpful to researchers for gaining insights into the potential, tradeoffs and challenges of DRAM caches.
Sparsh Mittal, Jeffrey S. Vetter
IEEE Trans. Parallel Distributed Syst.1
2016 EqualWrites: Reducing Intra-Set Write Variations for Enhancing Lifetime of Non-Volatile Caches
abstract
Driven by the trends of increasing core-count and bandwidth-wall problem, the size of last level caches has greatly increased, and hence the researchers have explored non-volatile memories (NVMs) that provide high density and consume low-leakage power. Since NVMs have low write endurance and the existing cache management policies are write variation (WV) unaware, effective wear-leveling techniques (WLTs) are required for achieving reasonable cache lifetimes using NVMs. We present EqualWrites, a technique for mitigating intra-set WV. Our technique works by recording the number of writes on a block and changing the cache-block location of a hot data item to redirect the future writes to a cold block to achieve wear leveling. Simulation experiments have been performed using an x86-64 simulator and benchmarks from SPEC06 and high-performance computing field. The results show that for single-, dual-, and quad-core system configurations, EqualWrites improves cache lifetime by 6.31× , 8.74×, and 10.54×, respectively. In addition, its implementation overhead is very small and it provides larger improvement in lifetime than three other intra-set WLTs and a cache replacement policy.
Sparsh Mittal, Jeffrey S. Vetter
IEEE Trans. Very Large Scale Integr. Syst.1
2015 DESTINY: a tool for modeling emerging 3D NVM and eDRAM caches
Matthew Poremba, Sparsh Mittal, Dong Li 0001, Jeffrey S. Vetter, Yuan Xie 0001
DATE2
2015 AYUSH: Extending Lifetime of SRAM-NVM Way-Based Hybrid Caches Using Wear-Leveling
abstract
The features and limitations of both SRAM and NVM (non-volatile memory) technologies have led the researchers to study SRAM-NVM way-based hybrid last level caches (LLCs). Since large leakage power consumption of SRAM allows including only few SRAM ways, the small write-endurance of NVM may still lead to small lifetime of these hybrid caches. We propose AYUSH, a technique for improving lifetime of SRAM-NVM hybrid caches. AYUSH uses data-migration approach to preferentially utilize SRAM for storing write-intensive data. Microarchitectural simulations have shown that AYUSH provides larger improvement in lifetime than three previous techniques. For single, dual and quad-core system configurations, the average increase in cache lifetime with AYUSH is 6.90×, 24.06× and 47.62×, respectively. Also, it does not harm performance or energy efficiency and works well for a range of system and algorithm parameters.
Sparsh Mittal, Jeffrey S. Vetter
MASCOTS1
2015 A Survey Of Architectural Approaches for Managing Embedded DRAM and Non-Volatile On-Chip Caches
abstract
Recent trends of CMOS scaling and increasing number of on-chip cores have led to a large increase in the size of on-chip caches. Since SRAM has low density and consumes large amount of leakage power, its use in designing on-chip caches has become more challenging. To address this issue, researchers are exploring the use of several emerging memory technologies, such as embedded DRAM, spin transfer torque RAM, resistive RAM, phase change RAM and domain wall memory. In this paper, we survey the architectural approaches proposed for designing memory systems and, specifically, caches with these emerging memory technologies. To highlight their similarities and differences, we present a classification of these technologies and architectural approaches based on their key characteristics. We also briefly summarize the challenges in using these technologies for architecting caches. We believe that this survey will help the readers gain insights into the emerging memory device technologies, and their potential use in designing future computing systems.
Sparsh Mittal, Jeffrey S. Vetter, Dong Li 0001
IEEE Trans. Parallel Distributed Syst.1
2014 WriteSmoothing: improving lifetime of non-volatile caches using intra-set wear-leveling
abstract
Driven by the trends of increasing core-count and bandwidth-wall problem, the size of last level caches (LLCs) has greatly increased. Since SRAM consumes high leakage power, researchers have explored use of non-volatile memories (NVMs) for designing caches as they provide high density and consume low leakage power. However, since NVMs have low write-endurance and the existing cache management policies are write variation-unaware, effective wear-leveling techniques are required for achieving reasonable cache lifetimes using NVMs. We present WriteSmoothing, a technique for mitigating intra-set write variation in NVM caches. WriteSmoothing logically divides the cache-sets into multiple modules. For each module, WriteSmoothing collectively records number of writes in each way for any of the sets. It then periodically makes most frequently written ways in a module unavailable to shift the write-pressure to other ways in the sets of the module. Extensive simulation results have shown that on average, for single and dual-core system configurations, WriteSmoothing improves cache lifetime by 2.17X and 2.75X, respectively. Also, its implementation overhead is small and it works well for a wide range of algorithm and system parameters.
Sparsh Mittal, Jeffrey S. Vetter, Dong Li 0001
ACM Great Lakes Symposium on VLSI1
2014 Improving energy efficiency of embedded DRAM caches for high-end computing systems
abstract
The number of cores in a single chip in the nodes of high-end computing systems is on rise, due, in part, to a number of constraints, such as power consumption. With this, the size of the last level cache (LLC) has also increased significantly. Since LLCs built with SRAM consume high leakage power, power consumption of LLCs is becoming a significant fraction of processor power consumption. To address this issue, researchers have used embedded DRAM (eDRAM) LLCs which consume low leakage power. However, eDRAM caches consume a significant amount of energy in the form of refresh energy. In this paper, we propose ESTEEM, an energy saving technique for embedded DRAM caches. ESTEEM uses dynamic cache reconfiguration to turn off a portion of the cache to save both leakage and refresh energy. It logically divides the cache sets into multiple modules and turns off possibly different number of ways in each module. Microarchitectural simulations confirm that ESTEEM is effective in improving performance and energy efficiency and provides better results compared to a recently-proposed eDRAM cache energy saving technique, namely Refrint. For single and dual-core simulations, the average energy saving in memory subsystem (LLC+main memory) with ESTEEM is 25.8% and 32.6% respectively, and the average weighted speedup is 1.09x and 1.22x respectively. Additional experiments confirm that ESTEEM works well for a wide-range of system and algorithm parameters.
Sparsh Mittal, Jeffrey S. Vetter, Dong Li 0001
HPDC1
2014 Quantitatively Modeling Application Resilience with the Data Vulnerability Factor
abstract
Recent strategies to improve the observable resilience of applications require the ability to classify vulnerabilities of individual components (e.g., Data structures, instructions) of an application, and then, selectively apply protection mechanisms to its critical components. To facilitate this vulnerability classification, it is important to have accurate, quantitative techniques that can be applied uniformly and automatically across real-world applications. Traditional methods cannot effectively quantify vulnerability, because they lack a holistic view to examine system resilience, and come with prohibitive evaluation costs. In this paper, we introduce a data-driven, practical methodology to analyze these application vulnerabilities using a novel resilience metric: the data vulnerability factor (DVF). DVF integrates knowledge from both the application and target hardware into the calculation. To calculate DVF, we extend a performance modeling language to provide a structured, fast modeling solution. We evaluate our methodology on six representative computational kernels, we demonstrate the significance of DVF by quantifying the impact of algorithm optimization on vulnerability, and by quantifying the effectiveness of specific hardware protection mechanisms.
Li Yu 0006, Dong Li 0001, Sparsh Mittal, Jeffrey S. Vetter
SC3
2014 MASTER: A Multicore Cache Energy-Saving Technique Using Dynamic Cache Reconfiguration
abstract
With increasing number of on-chip cores and CMOS scaling, the size of last-level caches (LLCs) is on the rise and hence, managing their leakage energy consumption has become vital for continuing to scale performance. In multicore systems, the locality of memory access stream is significantly reduced because of multiplexing of access streams from different running programs and hence, leakage energy-saving techniques such as decay cache, which rely on memory access locality, do not save a large amount of energy. The techniques based on way level allocation provide very coarse granularity and the techniques based on offline profiling become infeasible to use for large number of cores. We present a multicore cache energy saving technique using dynamic cache reconfiguration (MASTER) that uses online profiling to predict energy consumption of running programs at multiple LLC sizes. Using these estimates, suitable cache quotas are allocated to different programs using cache coloring scheme and the unused LLC space is turned off to save energy. Even for four core systems, the implementation overhead of MASTER is only 0.8% of L2 size. We evaluate MASTER using out-of-order simulations with multiprogrammed workloads from SPEC2006 and compare it with conventional cache leakage energy-saving techniques. The results show that MASTER gives the highest saving in energy and does not harm performance or cause unfairness. For twoand four-core simulations, the average savings in memory subsystem (which includes LLC and main memory) energy over shared baseline LLC are 15% and 11%, respectively. Also, the average values of weighted speedup and fair speedup are close to one (≥0.98).
Sparsh Mittal, Yanan Cao 0002, Zhao Zhang 0010
IEEE Trans. Very Large Scale Integr. Syst.1
2013 FlexiWay: A cache energy saving technique using fine-grained cache reconfiguration
abstract
Recent trends of CMOS scaling and use of large last level caches (LLCs) have led to significant increase in the leakage energy consumption of LLCs and hence, managing their energy consumption has become extremely important in modern processor design. The conventional cache energy saving techniques require offline profiling or provide only coarse granularity of cache allocation. We present FlexiWay, a cache energy saving technique which uses dynamic cache reconfiguration. FlexiWay logically divides the cache sets into multiple (e.g. 16) modules and dynamically turns off suitable and possibly different number of cache ways in each module. FlexiWay has very small implementation overhead and it provides fine-grain cache allocation even with caches of typical associativity, e.g. an 8-way cache. Microarchitectural simulations have been performed using an x86-64 simulator and workloads from SPEC2006 suite. Also, FlexiWay has been compared with two conventional energy saving techniques. The results show that FlexiWay provides largest energy saving and incurs only small loss in performance. For single, dual and quad core systems, the average energy saving using FlexiWay are 26.2%, 25.7% and 22.4%, respectively.
Sparsh Mittal, Zhao Zhang 0010, Jeffrey S. Vetter
ICCD1