Virendra Singh

dblp:31/4795 · DBLP profile ↗
← Back
101ranked-venue papers
4as first author
35since 2021 · last 2026
0000-0002-7035-7844ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 85 · 4 first-author · 21 since 2021Software engineering, systems software and programming languages · 17 · 4 since 2021Security and privacy · 9 · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FlIP: Flow-Based Instruction Processing for Out-of-Order Scheduling in GPGPUs
abstract
Graphics Processing Units (GPUs) achieve high throughput by exploiting abundant thread-level parallelism (TLP) to hide long-latency operations. However, many emerging general-purpose and scientific workloads exhibit irregular control flow and limited data-level parallelism, which reduce effective TLP and lead to frequent stalls and underutilized execution resources. While out-of-order execution can mitigate such stalls by exploiting instruction-level parallelism (ILP), conventional CPU-style out-of-order approaches are prohibitively complex and costly to scale to GPUs.
Munawira Kotyad, Yashvardhan Rathore, Ganesh Sai Shanmukhi, Virendra Singh
CF4
2026 POSTER: Bypassing Blocking Instructions to enable Out-Of-Order Execution in GPGPUs
Munawira Kotyad, Ganesh Sai Shanmukhi, Yashvardhan Rathore, Virendra Singh
CF4
2026 A SAT-Hard Compound Logic Locking Scheme with Empirical Resistance to Known Structural Attacks
abstract
Logic Locking aims to hide the original functionality of the design using a secret key. It protects hardware intellectual properties (IPs) against IP piracy or IC overproduction. However, an attacker analyzes the structural traces and/or uses Boolean satisfiability based technique called SAT attack to break such logic locking schemes. This motivates us to find a logic locking technique that can work against both SAT and structural analysis attacks. Therefore, this paper introduces a novel multiplier-based logic locking scheme. Leveraging the inherent complexity of multiplier circuits, the proposed scheme exponentially increases the time required for each iteration of a SAT attack. Moreover, a heuristic is also proposed to identify appropriate locations for inserting multiplier instances to increase the number of iterations. The multiplier-based logic locking scheme is further combined with the Anti-SAT scheme to create a robust and effective defense mechanism against SAT attacks and the various other attacks exploiting structural traces. The proposed technique is resilient to the state-of-the-art attack dedicated to the existing compound logic locking schemes. Moreover, the proposed compound logic locking scheme, requires half the number of key inputs than the state-of-the-art logic locking scheme while providing the similar level of security.
Sonali Shukla, Govind Rajhans Jadhav, Durgesh Sardan, Suryakant Toraskar, Jaynarayan T. Tudu, Masahiro Fujita 0004, Virendra Singh
DDECS8
2026 DREAM: Dynamic, Reinforced, and Evasive Attack Model on Spatio-Temporal GNNs
Mandru Suma Sri, Uddhav Narayan Gilda, Tarun Bisht, Virendra Singh
ICAART (3)4
2026 Sparse Rewards as Preferences: Investigating Reward Shaping with Preference-Based RL Methods in Sparse Reward Domains
Veerendrababu Vakkapatla, Shariq Faraz, Virendra Singh
ICAART (1)3
2026 Continuation-Preserving Tiling for Pointer-Chasing Optimization in Structured Mutual Recursion
Virendra Singh, Supratim Biswas
ICS2
2026 SPARTON: Secure Dynamic Partition Scheme for Last-Level Cache
Tejeshwar Bhagatsing Thorawade, Rishab Ravi, Varun Venkitaraman, Keerthisagar Kokkiligadda, Nirmal Kumar Boran, Virendra Singh
SECRYPT (1)6
2025 Weight-Aware Scan Chain Stitching for Shift Power Minimization under Routing Constraints
abstract
Scan chain architecture is a widely adopted design-for-testability (DFT) technique in modern VLSI circuits. However, with the increasing complexity of integrated circuits, excessive shift power during scan testing has become a critical concern, especially for power-constrained designs, as it can lead to increased thermal stress and reduced reliability. In this paper, we propose a weight-aware scan chain stitching methodology that effectively reduces both shift-in and shift-out power while maintaining optimized routing length. Experimental evaluations on ISCAS’89 and ITC’99 benchmark circuits demonstrate average reductions of 14.5% in scan shift power and a 68.5% improvement in routing efficiency compared to state-of-the-art techniques. Index Terms-scan-based testing, scan chain stitching, design for testability (DFT), routing constraint, low power testing.
Rohit Badjatya, Masahiro Fujita 0004, Virendra Singh
ATS5
2025 GPD: Predictive Control Flow Error Detection Leveraging Data Flow Error Detection Methods
abstract
Soft errors in General Purpose Graphics Processing Units (GPGPUs) result in data or control flow errors. Error detection and correction methods for data and control flow errors are orthogonal, and these methods incur separate area, power, and performance overheads. This paper proposes a low-overhead predictive control flow error detection method called GPGPU Predictive Detector (GPD), which leverages data flow error detection and correction methods to detect control flow errors. GPD is non-intrusive to application software and transparent to users. GPD is built on earlier work on data flow error detection and correction methods DDSR and TREFU. In GPD, DDSR and TREFU combined architecture protects all non-control flow instructions. The control flow error is detected by calculating the address of the instruction that succeeds the control flow instruction in advance and comparing it with the actual address it accesses. The effectiveness of GPD has been shown through a set of ISPASS-2009 and RODINIA benchmarks. Relative to a non-fault-tolerant GPGPU architecture, GPD has a performance overhead of 5% and average and peak power overheads of 4% and 3%, respectively. We prove through induction that the GPD provides fault coverage against GPGPU control flow and data flow errors.
Raghunandana K. K, Yogesh Prasad K. R, Matteo Sonza Reorda, Virendra Singh
IOLTS4
2025 A New Hardware Trojan Attack on Scan-obfuscated Logic-locked Circuits
abstract
Logic locking has emerged as a crucial defense mechanism for securing ICs against threats such as IP theft, counterfeiting, and Hardware Trojans (HTs). However, advanced attacks like Boolean Satisfiability (SAT) attack have exposed vulnerabilities by exploiting scan-unlocked oracle to retrieve secret keys. To mitigate this risk, scan obfuscation techniques were introduced to secure scan access and enhance protection against SAT attack. However, ScanSAT attack has been shown to bypass these defenses, successfully retrieving secret keys even from scan-obfuscated circuits. Recently, a test authentication scheme combined with scan obfuscation has been proposed as a countermeasure against ScanSAT attack.This work presents an attack that employs a stealthy HT to subvert the security provided by the test authentication scheme. Our analysis demonstrates that the inserted Trojan not only facilitates the execution of the ScanSAT attack but also eludes detection by the state-of-the-art Hardware Trojan detection techniques. These results highlight critical vulnerabilities in current IC security measures, emphasizing the need for resilient defenses against increasingly sophisticated attacks.
Anjum Riaz, Gaurav Kumar 0001, Yamuna Prasad, Satyadev Ahlawat, Virendra Singh
ISCAS5
2025 RRR: Robust runtime reconfigurable shared cache management scheme for GPGPUs
abstract
General-Purpose GPUs (GPGPUs) are ideal for parallel processing and high throughput. However, efficient on-chip memory use is challenging due to resource conflicts between threads. This leads to sub-optimal throughput, highlighting the need for better memory designs. As compute demand grows, GPUs with more Streaming Multiprocessors (SMs) emerge, increasing bandwidth needs and intensifying Network-on-Chip (NoC) traffic. Current GPUs partition the shared Last-Level Cache (LLC) into uniform slices shared by all SMs. While this reduces miss rates, workloads with high inter-SM data sharing benefit more from private LLCs, which reduce contention and improve bandwidth. We propose a dynamic profiling method to evaluate data sharing across SMs during execution. Using this, we present a logistic regression-based decision mechanism to switch between shared and private LLC configurations at runtime, with a twenty five thousand cycle reconfiguration epoch (compared to 1 million cycles in the state-of-the-art). Our approach boosts performance by up to 56% over the baseline and 29% over state-of-the-art methods. It also reduces stalls by 71% over the baseline and 59% over state-of-the-art approach.
Varun Venkitaraman, Shrasti Bhargava, Tejeshwar Bhagatsing Thorawade, Keerthisagar Kokkiligadda, Virendra Singh
ISCAS6
2025 Test Data Compaction Techniques with Improved Diagnostic Capabilities and Reduced Tester Time
abstract
The IC manufacturing has become very challenging in today’s world due to miniaturization of transistor sizes, high logic density and complex fabrication processes. At the same time, with the shrinking time-to-market schedule, it becomes very important to be able to detect these manufacturing defect and rectify the fabrication processes faster in order to have better quality products during mass production. Compression techniques with decompressors at input and compactors at output have been adopted to shorten the scan chain lengths but, the data loss during compaction makes the diagnosis harder and the difficulty increases with high data compaction techniques like Multiple-Input-Signature-Registers(MISRs), predominantly used in Logic-BIST(LBIST). These type of compactors are also plagued by late failure detection issue and poor diagnostics, since it is very difficult to identify the Fault Capturing Scan Flops(FCSFs) which is the pre-requisite for locating the exact fault site. In this thesis, a novel method of comparing the MISR signature intermittently inside the chip (on-chip) by using a simple circuitry of shift registers, a counter and XOR comparator, along with a dynamic resettable MISR has been proposed and explained. This method detects the failures much earlier than the conventional MISR approach and also enables testing of the ICs with lesser number of scan outputs , thus improving the parallel testing capabilities (multi-site testing) and reducing test time by 30% to 50%, eventually leading into test cost savings. In addition, the diagnostic time of LBIST at the tester was reduced by > 95% since the intermittent MISR reset ensures that the effect of the error or X-sources in the MISR does not corrupt the entire signature and is limited to few cycles. Apart from the improvements for MISR compaction, a unique, first-of-a-kind compaction method has been invented which uses prime numbers to exactly identify the location of one or two FCSFs directly from the signature without requiring any regeneration of patterns. The mechanism and the mathematical proofs for the proposed idea have been explained in detail. Even though the area overhead is higher compared to the existing methods, the flexibility in terms of implementation have also been discussed. For all the simulation runs, 100% accuracy of identifying one or 2 FCSF locations has been achieved.
Jaidev Shenoy, Virendra Singh, Kelly Ockunzzi
ITC2
2025 TRAP: Time-Aware Probabilistic In-Dram RowHammer Solution
abstract
DRAM continues to serve as the backbone of main memory in modern computing systems. However, it is increasingly vulnerable to RowHammer, a circuit-level vulnerability in which repeated activation of a row can induce bit flips in adjacent rows. Among the range of mitigation strategies, low cost inDRAM solutions are particularly attractive because they operate entirely within the DRAM chip and require no changes to the broader system architecture. Despite their appeal, existing inDRAM defenses often face two critical limitations: non-uniform mitigation probability across vulnerable rows and high non-selection rates, both of which reduce their effectiveness and reliability under worst-case access patterns.We propose TRAP: Time-Aware Probabilistic in-DRAM RowHammer Protection. TRAP is a lightweight, hardware-efficient scheme that introduces time-slot-based sampling probabilities for each row activation within a DRAM refresh interval. TRAP tracks the timing of activations using a per-bank counter and applies analytically derived probabilities from a precomputed lookup table, ensuring uniform mitigation coverage and reducing the probability of no selection to near zero. Evaluations on SPEC2017, PARSEC, and LIGRA benchmarks demonstrate that TRAP delivers strong and consistent protection with zero performance overhead and only 2.5% DRAM energy increase. When integrated with the DDR5 Refresh Management (RFM) feature, TRAP incurs just 0.2% slowdown, underscoring its practicality for future DRAM systems.
Samiksha Verma, Virendra Singh
SBAC-PAD2
2025 Stegoslayer: A Robust Browser-Integrated Approach for Thwarting Stegomalware
Rushikesh Kawale, Sarath Babu 0003, Virendra Singh
SECRYPT3
2025 EDQKD: Enhanced-Dynamic Quantum Key Distributions with Improved Security and Key Rate
Nikhil Kumar Parida, Sarath Babu 0003, Neeraj Panwar, Virendra Singh
SECRYPT4
2025 SCAM: Secure Shared Cache Partitioning Scheme to Enhance Throughput of CMPs
Varun Venkitaraman, Rishab Ravi, Tejeshwar Bhagatsing Thorawade, Nirmal Kumar Boran, Virendra Singh
SECRYPT5
2025 LiC: Low-Cost Cache Replacement Algorithm for All Cache Levels
abstract
Modern processors use caches to reduce memory access time. However, their limited size leads to frequent misses, requiring an efficient replacement policy. The Least Recently Used (LRU) policy is widely adopted for its effectiveness but becomes impractical in highly associative caches due to its high area and power costs. To address inefficiencies in the last-level cache (LLC), researchers have proposed sophisticated replacement policies. However, their complexity and hardware overhead make them unsuitable for level-one (L1) and leveltwo (L2) caches, which require fast and lightweight decisionmaking. Additionally, low-cost microcontrollers demand simple and efficient replacement mechanisms. This paper introduces the Lightweight Cache Replacement Policy (LiC) as a low-cost, power-efficient alternative to LRU. Unlike conventional policies that focus on eviction decisions, LiC prioritizes protecting the last accessed block. This approach significantly reduces hardware complexity and power consumption while maintaining performance. We evaluate LiC through simulations in both single-core and multi-core environments. Results show that LiC matches LRU’s performance while drastically reducing storage overhead. Compared to sophisticated LLC replacement policies, it reduces storage costs by up to $28 \times$. Against low-cost policies like NRU and PLRU, it achieves $4 \times$ and $3.75 \times$ lower storage demands. Additionally, LiC reduces area overhead by $16 \times$ compared to LRU. With its low hardware overhead and strong performance, LiC emerges as an efficient and scalable solution across all cache levels.
Varun Venkitaraman, Tejeshwar Bhagatsing Thorawade, Mitul Tandon, Keerthisagar Kokkiligadda, Virendra Singh, Janak Patel
VLSI-SoC5
2025 B-CAVE: A Robust Online Time Series Change Point Detection Algorithm Based on the Between-Class Average and Variance Evaluation Approach
abstract
Change point detection (CPD) is a valuable technique in time series (TS) analysis, which allows for the automatic detection of abrupt variations within the TS. It is often useful in applications such as fault, anomaly, and intrusion detection systems. However, the inherent unpredictability and fluctuations in many real-time data sources pose a challenge for existing contemporary CPD techniques, leading to inconsistent performance across diverse real-time TS with varying characteristics. To address this challenge, we have developed a novel and robust online CPD algorithm constructed from the principle of discriminant analysis and based upon a newly proposed between-class average and variance evaluation approach, termed B-CAVE. Our B-CAVE algorithm features a unique change point measure, which has only one tunable parameter (i.e. the window size) in its computational process. We have also proposed a new evaluation metric that integrates time delay and the false alarm error towards effectively comparing the performance of different CPD methods in the literature. To validate the effectiveness of our method, we conducted experiments using both synthetic and real datasets, demonstrating the superior performance of the B-CAVE algorithm over other prominent existing techniques.
Adeiza Onumanyi, Satyadev Ahlawat, Yamuna Prasad, Virendra Singh
IEEE Trans. Knowl. Data Eng.5
2024 TCC: GPGPU Architecture for Instruction Decoder and Control Flow Error Detection
abstract
The devices fabricated with the latest sub-nanometer technology node have a higher probability of parametric and wear-out failures, operational faults, and manufacturing defects, and these devices are more susceptible to intrinsic and extrinsic noise, resulting in soft errors. The parts with manufacturing defects are generally screened out during end-of-manufacturing tests. Thus, the soft errors during normal operations are of great concern. The soft errors in GPGPUs result into silent data corruption and control flow divergence errors. In order to deal with this, dual and triple modular redundancy architectures are used for soft error detection and correction, which result in large areas and power overheads. To overcome this, we propose a low overhead fault-tolerant microarchitecture called Trace Consistency Check (TCC) to detect the decoder and control flow divergence errors. The TCC is transparent to the application software. For error detection, we exploit the execution model of GPGPUs, where the warps of kernel executing in the streaming multiprocessor have temporal execution repetition. Hence, the instruction execution trace and control divergence paths across the warps are consistent. Inconsistency across warps for the same code region is attributed to decoder or control divergence errors. For error detection, new microarchitecture structures called Execution Trace buffer and Control Divergence Trace buffer were introduced to store and check the trace consistency across warps. The performance of TCC is evaluated through the ISPASS 2009 and RODINIA benchmarks. TCC's error detection capability and power overheads are evaluated. The simulation results show that TCC detects greater than 99% decoder and control flow errors with low power and no performance overheads.
Raghunandana K. K, Yogesh Prasad K. R, Matteo Sonza Reorda, Virendra Singh
DDECS4
2024 HIDC: Heterogeneous-ISA Dynamic Core
abstract
Heterogeneous-ISA architectures are emerging as an alternative for better performance and energy efficiency. These systems exploit ISA affinity within and across programs by executing them dynamically on the most suited ISA core. However, heterogeneous-ISA chip multiprocessor incurs significant power and area overheads compared to single-ISA systems. Also, due to switching overhead between different ISA cores, the migration opportunities are constrained to coarse-grained program phases (order of 100M instructions), limiting the potential of harnessing ISA diversity. We propose Heterogeneous-ISA Dynamic Core (HIDC), an architecture which reduces the migration overhead by supporting different ISAs within a single core and an improved migration strategy called Simultaneous Transformation. HIDC improves single-threaded performance and energy efficiency by exploiting ISA diversity at a finer granularity and reducing power consumption compared to heterogeneous-ISA CMP. To optimally schedule program phases as per their ISA affinity, we present a low-overhead perceptron-based scheduling mechanism along with its implementation in hardware. Overall, HIDC improves by 52.89% and 33.93% in performance per joule metric over heterogeneous-ISA CMP and single x86 ISA core, respectively. It achieves 4.62% and 19.06% performance improvement relative to heterogeneous-ISA CMP and single x86 ISA core. Also, it leads to ∼20× reduction in cross-ISA migration latency.
Nirmal Kumar Boran, Prakhar Diwan, Meet Udeshi, Shubhankit Rathore, Virendra Singh
HPCC5
2024 GNNDLD: Graph Neural Network with Directional Label Distribution
Chandramani Chaudhary, Nirmal Kumar Boran, N. Sangeeth, Virendra Singh
ICAART (2)4
2024 S-Clflush: Securing Against Flush-based Cache Timing Side-Channel Attacks
abstract
Micro-architectural attacks exploit intrinsic vulnerabilities within computing systems, circumventing advanced security techniques such as cryptographic algorithms, access control policies, and secure enclaves. These attacks encompass a range of methodologies, including cache timing side-channel attacks like Flush+Reload, Flush+Flush, and Prime+Probe, as well as speculative execution attacks such as Spectre and Meltdown. These exploits leverage specific characteristics of micro-architecture to infer sensitive data, posing a significant threat to system security. Cache timing side-channel attacks exploit the inclusive nature of the last-level cache (LLC) to deduce the memory access patterns of victim processes. By observing the timing variations associated with cache hits and misses, attackers can extract confidential information, such as cryptographic keys. Although existing mitigation strategies provide a level of security, they typically do so at the expense of system performance and increased hardware. These trade-offs limit the practical applicability of such defences in performance-critical environments. This paper proposes S-Clflush: Secure Clflush, an innovative defence mechanism specifically designed to counter flush-based cache timing side-channel attacks. S-Clflush achieves this by modifying the existing clflush instruction to prevent attackers from inferring memory access patterns based on cache access latency. Unlike traditional mitigation techniques, S-Clflush enhances security without incurring performance degradation or additional area overhead. The proposed mechanism is formally verified to ensure its security guarantees. Our evaluation against the state-of-the-art mitigation technique TimeCache shows a 0.5% improvement in performance and a 58% reduction in MPKI on average without adding area overhead.
Tejeshwar Bhagatsing Thorawade, Prajakta Yeola, Varun Venkitaraman, Virendra Singh
SBAC-PAD4
2024 Critical Behavior Sequence Monitoring for Early Malware Detection
abstract
The widespread use and the immense user base make Windows systems a prime target for attackers seeking to exploit vulnerabilities and maximize impact. Modern obfuscation techniques enable malware to evade detection tools, allowing it to intrude on systems and carry out malicious activities. Thus, to prevent potential harm to the victim's system, it is crucial to detect malware at the early execution stage and initiate adequate action. However, there is often a trade-off between accuracy and earliness in malware detection, as detecting threats at earlier stages may sometimes come at the cost of reduced detection accuracy. We introduce an early malware detection framework that balances this trade-off. Our proposed framework iteratively constructs API call prefix subsequence and applies security-sensitive embedding using API call parameters. The causal sequence encoder transforms these sequences into contextual vectors, which are then classified by a multi-layer perceptron. The proposed framework not only outperforms the state-of-the-art early malware detection method EarlyMalDetect but also demonstrates comparable performance to post-sequence malware detection methods like CTIMD and BD-MDLC, achieving accurate maliciousness prediction on or before analyzing just 3 % of the API sequence.
Tarun Bisht, Sarath Babu 0003, Virendra Singh
SIN3
2024 BD-MDLC: Behavior description-based enhanced malware detection for windows environment using longformer classifier
Sarath Babu 0003, Virendra Singh
Comput. Secur.2
2024 DAT: A robust Discriminant Analysis-based Test of unimodality for unknown input distributions
Adeiza Onumanyi, Satyadev Ahlawat, Yamuna Prasad, Virendra Singh, Adnan M. Abu-Mahfouz
Pattern Recognit. Lett.5
2023 SMASh: A State Encoding Methodology Against Attacks on Finite State Machines
abstract
Finite State Machines (FSMs) are critical components that control the functionality of chips and their individual components. They can be compromised by introducing unwanted transitions to the security-critical states. This makes them vulnerable to several attacks that aim to identify the FSM transitions. This work handles the Differential Power Analysis (DPA) attacks and Fault Injection Attacks (FIA) on FSMs. Therefore, this paper proposes a state encoding technique that ensures the maintenance of hamming distance across all transitions, providing protection against DPA. This property of hamming distance preservation is further utilized to detect and mitigate potential faults injected into the system. Hence, addresses the FIA attacks. The proposed scheme is independent of the number of FSM transitions. It also introduces confusion to attackers by dynamically varying the state encoding based on the additional bits required to maintain the Hamming Distance (HD). The effectiveness of the proposed scheme is demonstrated in terms of resource utilization compared to traditional encoding schemes through extensive experimentation on MCNC benchmark circuits. Moreover, the technique proposed in this paper is quite flexible for a wide range of FSM designs. Thus, provide generic framework to enhance the security of digital systems.
Gowthami Konganapalle, Sonali Shukla, Virendra Singh
ATS3
2023 ERrOR: Improving Performance and Fault Tolerance Using Early Execution
abstract
Contemporary integrated circuits are becoming increasingly susceptible to soft errors due to single-event upsets, effectively decreasing the reliability of operation. In this paper, we propose the ERrOR microarchitecture, that detects soft errors in processor operation using temporal redundancy with minimal hardware overhead. Previous proposals have explored the idea of introducing an Early Execution Unit (EXU) at the processor frontend in order to expeditiously execute dynamic instructions with short dependency chains for performance improvement. However, we observe that the functional units in the EXU are idle for a significant fraction of the program execution duration. ERrOR leverages these inactive frontend functional units to re-execute dynamic instructions for the purpose of error detection. A lightweight verifier introduced at the backend makes use of idle resources for redundant execution by interleaving program execution with re-execution for error detection. ERrOR provides exhaustive transient fault coverage while improving performance by 7.5% over an existing restricted OoO microarchitecture, Freeflow Core.
Raj Kumar Choudhary, Janeel Patel, Virendra Singh
IOLTS3
2023 On-Chip SRAM Disclosure Attack Prevention Technique for SoC
abstract
Recently a prominent attack vector called Volt Boot attack was disclosed, which exploits low power mode (i.e., power gating) transitions and power distribution networks of SoCs, to compromise the content of on-chip SRAMs with higher data accuracy than traditional cold boot attacks. This attack can compromise on-chip SRAMs that store plaintext data for subsequent processing. The present mitigation techniques either incur longer latency or increase the area/power of the SoC. This paper proposes address and data mapping/de-mapping with different true random number generators(TRNG). These TRNGs are used to map/de-map the address and data to newer values in the memory array when the boot/reset/power or tampering event is detected. Any memory read/write request is served after mapping/de-mapping the address/data with appropriate boolean functions using stored TRNGs. This technique changes the memory contents or traces after each boot/power cycle of SoC or tampering event with respect to the contents of previous computations. Our results show that the proposed technique makes current memory data uncorrelated to previously stored memory contents. Also, our technique shows low implementation area overhead and reduces the latency of data corruption to one clock cycle.
Prokash Ghosh, Yogesh Gholap, Virendra Singh
IOLTS3
2023 TREFU: An Online Error Detecting and Correcting Fault Tolerant GPGPU Architecture
abstract
General Purpose Graphics Processing Units (GPGPUs) are extensively used in high-performance applications/systems, whose execution times may vary from a few days to months. Many times, these systems are expected to provide high reliability and availability. On the other hand, the high-throughput GPGPUs are fabricated with the latest cutting-edge technology. The shrinking transistor feature size and aggressive voltage scaling resulted in increased susceptibility to soft errors. Hence, GPGPU execution results cannot be trusted. This necessitates the employment of error detection and correction methods for reliable results. To mitigate soft error effects in the GPGPU execution pipeline, we propose a fault-tolerant microarchitecture called Triple modular Redundant Execution with idle Functional Units (TREFU) to detect and correct errors online. The proposed method is transparent to the application software. A new microarchitecture structure, replay buffer, is introduced to store temporary operands and results and used as a checkpoint. On error detection, the data in the duplicate copy of the replay buffers are used for Triple Modular Redundant (TMR) execution and error correction. The effectiveness of TREFU is demonstrated through the ISPASS 2009 and RODINIA benchmarks. TREFU's performance and power overheads are evaluated for an error-free run and at various error rates of executed instructions ranging from 1 to 50K. The simulation results show that complete error detection and correction across all threads can be achieved with a mean performance overhead of 4%, an average power overhead of 4%, and a peak power overhead of 5%.
Raghunandana K. K, B. K. S. V. L. Varaprasad, Matteo Sonza Reorda, Virendra Singh
IOLTS4
2022 PASS-P: Performance and Security Sensitive Dynamic Cache Partitioning
Nirmal Kumar Boran, Pranil Joshi, Virendra Singh
SECRYPT3
2022 Exploiting post-silicon debug hardware to improve the fault coverage of Software Test Libraries
abstract
Functional test using a Software Test Library (STL) is becoming a standard solution for the in-field test of safety-critical systems, in compliance with functional safety standards, such as the ISO26262 for the automotive domain. However, developing high-quality test programs is considerably more challenging than generating scan test patterns through commercial tools, mainly due to the lack of mature EDA tools. As a result, in many cases, the effort needed to reach the target fault coverage is not affordable. In this paper, we propose a methodology to improve the fault coverage of an STL using already available hardware resources. The proposed approach identifies the set of sequential cells that capture fault effects before being masked during their propagation towards observable points. Using a heuristic set covering approach, we select the subset of flip-flops needed to reach the target fault coverage, and exploit post-silicon debug hardware to make fault effects observable. Experimental results gathered on an open-source RISC-V core show significant improvements in the stuck-at and delay fault coverage values.
Riccardo Cantoro, Francesco Garau, Riccardo Masante, Sandro Sartoni, Virendra Singh, Matteo Sonza Reorda
VTS5
2021 Predictive Warp Scheduling for Efficient Execution in GPGPU
abstract
Today's general-purpose graphics processing units (GPGPUs) offer phenomenal performance to applications from a variety of fields. Despite its memory-latency tolerant parallel architecture, GPGPU cores do not attain optimal performance with most of the memory-intensive applications which add a significant load on memory-resources. The existing thread scheduling mechanism is not optimized to address this challenge for latency-sensitive applications. In this paper, we propose a warp-scheduling policy that defers the execution of all those warps which will potentially cause long-latency memory-stalls, to prevent the resulting congestion in interconnect network and DRAM bandwidth. The proposed policy uses a predictor in each core of the GPU to predict whether or not the data can be retrieved from the L1-cache at the time of increased congestion to effectively reduce the number of warps that can be paused and keep the cores active. This translates to an average performance improvement of 11.6% and up to 41.2% over the state-of-the-art scheduling policies across a diverse selection of applications with a 4.2% increase in average power consumption.
Abhinish Anand, Winnie Thomas, Suryakant Toraskar, Virendra Singh
ACM Great Lakes Symposium on VLSI4
2021 Dynamic Optimizations in GPU Using Roofline Model
abstract
Massively parallel processors such as graphics processing units (GPUs) often face the challenge of resource underutilization due to varying resource proclivity of workloads. Running multiple applications on a GPU has been an efficient and known alternative to mitigate underutilization. This paper proposes a multi-application oriented framework that carries out dynamic optimizations based on the operational intensities of various applications. Our framework analyzes applications based on operational intensities to identify their bottleneck resources using Roofline model. We demonstrate that the proposed optimizations improve the utilization and system-wide throughput of the GPU co-running applications with irregular resource demands. The dynamic optimizations improve the performance by 14.8% on average and up to 72.4% over a state-of-the-art spatial multitasking technique.
Winnie Thomas, Suryakant Toraskar, Virendra Singh
ISCAS3
2021 A Framework for Configurable Joint-Scan Design-for-Test Architecture
Jaynarayan T. Tudu, Satyadev Ahlawat, Sonali Shukla, Virendra Singh
J. Electron. Test.4
2021 Enhanced Design Debugging With Assistance From Guidance-Based Model Checking
abstract
Design debugging is one of the most important steps in the modern integrated circuits (ICs) development cycle. Simulation-based verification is never sufficient for ensuring design correctness because of its incomplete nature. Formal techniques such as model checking promise to solve this issue through a complete state-space traversal approach. However, because of increasing design complexity, such methods suffer from scalability issues. Guidance-based state-space traversal techniques have been proposed in the past to assist the model checkers in overcoming the complexity bottleneck. Automatically identifying these guidance hints is relatively difficult and requires heuristic-based reasoning procedures. Additionally, to come up with quick fixes during the debug stage, an effective bug localization strategy is needed. In this article, we revisit the paradigm of guidance-based model checking and propose a methodology to improve these guidance generation mechanisms for achieving fine-grained bug localization. In particular, this work proposes a systematic methodology to localize the buggy RTL lines from the erroneous RTL simulation trace. The proposed technique involves the mining of invariant-like assertions from simulation traces. The mined assertions act as probable guidance candidates for the model checking exercise. To identify useful guidance hints from possible ones, we use the Bayesian networks that explore conditional dependence between the various hints at different levels and the target property. These guidance hints are utilized for obtaining possible buggy subregions, which are analyzed via an iterative model checking methodology for fine-grained bug localization. By using the proposed framework, bugs can be localized to within a few lines of RTL description.
Vineesh V. S., Binod Kumar 0001, Rushikesh Shinde, Masahiro Fujita 0004, Virendra Singh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2020 LUT-based Circuit Approximation with Targeted Error Guarantees
abstract
Approximate circuits are widely gaining popularity in various fields where error tolerance is applicable. However, striking the right balance between error tolerance and the output quality is a challenging step in the overall design of approximate systems. We propose a systematic approach utilizing Look-Up Table (LUT)-based netlist transformations to achieve approximation while targeting specific error guarantees. Specifically, we employ a SAT-based property checking technique to accommodate worst-case error constraints acting as error guarantees. The proposed methodology involves the formulation of templates to enable the reusability of the technique for different design choices. The analysis comprises of fitness function evaluation based on layout area or the considered error guarantees. We analyze the impact of different parameters on the quality of the output of the resulting approximation and the time taken to obtain them.
Vinod G. U, Vineesh V. S., Jaynarayan T. Tudu, Masahiro Fujita 0004, Virendra Singh
ATS5
2020 Characterization of Data Generating Neural Network Applications on x86 CPU Architecture
abstract
This paper analyzes the performance of two contemporary data-generating neural network-based workloads, Neural Style Transfer and Super Resolution GAN run on x86 hardware architecture. In understanding the impact of data-readiness, we find how certain layers benefit from forced data warming-up. In examining bandwidth utilization of these layers, we identify several memory-bound layers as not necessarily being bandwidthbound hinting at the feasibility of prefetch-based solutions for improved performance. We also observe layers with specific kernel sizes performing poorly because of their unoptimized library kernel implementation. Based on our findings, we suggest directions for removing these performance bottlenecks by utilizing available bandwidth margins ≥ 90% and realizing convolution operations through vector-based functional units with a scope of at least 20x more such software-to-hardware mappings than existing implementation.
Antara Ganguly, Shankar Balachandran, Anant Nori, Virendra Singh, Sreenivas Subramoney
ISPASS4
2020 Post-Silicon Gate-Level Error Localization With Effective and Combined Trace Signal Selection
abstract
Incorporating on-chip trace buffers (TBs) helps to overcome the limited observability by tracing selected signals during post-silicon validation. The effectiveness of TB-based techniques largely relies on selection of appropriate trace signals. For processor-based systems, the selection becomes relatively easier because important signals can be identified. However, for a general digital block in a complex system-on-chip, recognizing necessary trace signals becomes extremely challenging and requires a systematic approach. Previous research on trace signal selection has mainly focused on improving reconstruction of unknown signal values with the help of traced signals. Even though it serves as a good selection principle, an effective signal selection must consider other important factors such as error detection (ED) with the traced signals, which in turn assist in localization and root-cause discovery. Additionally, from practical point of view, the signal selection algorithm needs to cater to factors like routing congestion and minimizing routing wire length. The proposed methodology of signal selection attempts to combine these three crucial factors of signal selection: restoration of untraced signal states, ED with traced signals and routing considerations. The concurrent maximization of all these three parameters is difficult as they have conflicting preference of the candidate trace signals. Hence, the proposed signal selection approach presents a methodology of judiciously mixing the choices of these three objectives. Furthermore, the restored and traced signal states are analyzed for the purpose of error localization at the gate level for several design error models.
Binod Kumar 0001, Kanad Basu, Masahiro Fujita 0004, Virendra Singh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 A Methodology to Capture Fine-Grained Internal Visibility During Multisession Silicon Debug
abstract
Silicon debugging is carried out in multiple sessions which are characterized by run-and-halt intervals. One of the important criteria for the success of this method is that the debugging infrastructure should capture only the erroneous data which can add important insights to the debugging process. However, identification of such suspect clock cycles is not a trivial exercise and requires an systematic approach. We propose a debugging architecture for enhancing the multisession procedure using the technique of on-chip debug data compression. The first session assists in identifying those erroneous clock cycles, and the useful debug data are collected in the second session with the help of markers called tag bits. At the cost of a minimal increase in area overhead, the proposed architecture achieves finer temporal visibility expansion because of the debug data collection in a segregated manner. During the offline analysis of the collected debug data, error localization can be achieved to a finer resolution. We evaluate our methodology on several designs for different kinds of error configurations. Experimental results show that the proposed methodology can achieve better on-chip storage utilization and the expansion in the temporal observation window compared to similar techniques in the literature.
Binod Kumar 0001, Jay Adhaduk, Kanad Basu, Masahiro Fujita 0004, Virendra Singh
IEEE Trans. Very Large Scale Integr. Syst.5
2019 Validating Multi-Processor Cache Coherence Mechanisms under Diminished Observability
abstract
Modern chip multi-processors (CMP) inevitably require cache coherence mechanisms for their correct operation. However, exhaustive functional verification of a complex cache coherence mechanism is a challenging task. This leads to bugs escaping to the first silicon and necessitates validation at the post- silicon stage. In this work, an on-chip signal logging method is proposed which helps in bug detection in case of design errors and soft-errors arising out of reliability issues. The logged contents can then be further dumped off-line for fine-grained bug localization. The proposed methodology utilizes cache coherence protocol specifications to obtain the signal states of coherence transactions and the detector module flags an error once a mismatch is found between observed signal states and correct signal states. The proposed logging mechanism decreases the error detection latency at minimal area and power overheads. Experiments on a four core multiprocessor having a 7-stage MIPS pipeline implementing the widely utilized directory-based MESI protocol indicate that the proposed methodology succeeds in detecting design errors. Analysis of soft errors have also been performed and shorter error detection latency is achieved compared to a previously proposed technique in the literature.
Binod Kumar 0001, Atul Kumar Bhosale, Masahiro Fujita 0004, Virendra Singh
ATS4
2019 Orion: A Technique to Prune State Space Search Directions for Guidance-Based Formal Verification
abstract
Model checking of large designs is a challenging task because of different scalability issues. In this paper, we aim to utilize guided state space traversal to address this issue. However, providing guidance for state space traversal of complex designs is also an equally challenging problem. We adopt a simulation-based strategy combined with Bayesian modelling approach for finding effective guidance hints for state space traversal. A heuristic-based structural dependency of the design yields ineffective guidance hints which need further of filtering. To prune out the ineffective guidance hints, we first generate module-level sub-properties from static analysis of the design. These sub-properties and structural dependency-based guidance hints are analyzed in simulation traces generated from the constrained-random test benches. These conditional occurrence of sub-properties and guidance hints are inputs to a Bayesian model which can then provide us the guidance hints with the highest profitability. With the proposed methodology, we succeed in pruning out the set of unprofitable guidance hints and obtain effective search directions which are then used to assist the model checking procedure. Experiments on two complex designs for different properties show the effectiveness of the proposed methodology in reducing CPU time during model checking.
Vineesh V. S., Binod Kumar 0001, Rushikesh Shinde, Akshay Jaiswal, Harsh Bhargava, Virendra Singh
ATS6
2019 Freeflow Core: Enhancing Performance of In-Order Cores with Energy Efficiency
abstract
The Out-of-Order (OoO) superscalar core design has been widely adopted for high performance computing. It exploits both instruction level parallelism (ILP) and memory level parallelism (MLP) to speedup the program's execution. However, due to unresolved data dependencies among instructions, exploiting ILP becomes at times difficult and make the OoO core idle, thus reducing its energy-efficiency. Our study focuses on selective exploitation of inherent ILP present in the program. Our proposed architecture, Freeflow Core (FC), focuses on discovering the selective opportunities for conversion of instruction guided execution to data guided execution. Such selective mechanism improves the performance without incurring substantial energy budget overheads. Memory-related instructions are handled in FC by memory-access pipeline and compute instructions are handled by compute pipeline. Giving priority to load/store instructions has been one of the known techniques to improve performance. However, less importance has been given to non-ready instructions which may block the head of the compute pipeline. We observe that such instructions are mainly those whose producers are unresolved in the memory-access pipeline. Hence, FC detects instructions that are dependent on unresolved memory instructions and guides them to a dedicated independent execution path. This segregation enables the younger ready instructions to free flow through the compute pipeline. Our evaluations show that FC outperforms InO and state-of-the-art Load Slice Core (LSC) both in performance and energy efficiency metrics.
Raj Kumar Choudhary, Newton, Harideep Nair, Rishabh Rawat, Virendra Singh
ICCD5
2019 Securing Scan through Plain-text Restriction
abstract
Scan design-for-test (DfT) feature can be exploited as a side channel to break a cryptographic chip. The stringent test and diagnosis requirements of present-day complex system-on-chip (SoC) make use of the scan DfT feature unavoidable. However, being a threat to cryptographic chips, it needs to be secured against the scan-based side-channel attacks. In this paper, we propose a simple yet effective technique to prevent scan attack on Advanced Encryption Standard (AES) cryptographic chip. The proposed technique restricts the user from applying any random inputs at the plain-text inputs. To use the scan feature the plain-text inputs must be forced to a constant all-0 or all-1 value throughout the test session. Because of this feature, there is no possibility of mounting any differential scan attack. The proposed technique is simple to implement and does not have any impact on test coverage.
Satyadev Ahlawat, Kailash Ahirwar, Jaynarayan T. Tudu, Masahiro Fujita 0004, Virendra Singh
IOLTS5
2019 SAT-based Silicon Debug of Electrical Errors under Restricted Observability Enhancement
Binod Kumar 0001, Masahiro Fujita 0004, Virendra Singh
J. Electron. Test.3
2018 ATPG power guards: On limiting the test power below threshold
abstract
Modern circuits with high performance and low power requirements impose strict constraints on manufacturing test generation, particularly on timing test. Delay test is used for performance grading of the circuit. During the application of the test, power consumption has to be less than the functional threshold value, in order to avoid yield loss. This work proposes a new direction to generate power safe test without any changes in DFT (design for testability) structure or existing CAD (computer-aided design) tools. We propose a virtual wrapper circuitry around the circuit under test (CUT), for test generation purpose, which acts as a shield to obtain power safe vectors. The wrapper prohibits the generation of test vector if power consumption exceeds the threshold limits. We consider analytical power models for power analysis of candidate test vector patterns. Experiments performed on benchmark circuits show power safe test generation without coverage loss.
Rohini Gulve, Virendra Singh
DATE2
2018 AB-Aware: Application Behavior Aware Management of Shared Last Level Caches
abstract
In modern multicore systems, Last-Level Cache (LLC) is usually shared among multiple cores. Though it benefits applications by sharing and utilizing cache resources efficiently; the benefits come at the cost of increased conflict misses due to interference among applications. In shared LLC, conventionally used LRU-based cache replacement policies logically partition the cache on-demand basis. Thus, cache friendly applications sharing LLC with streaming applications, suffer due to high data demands and low reuse of streaming applications. Apart from different data locality behavior, applications also show different memory access behavior while accessing the LLC. Some applications inherently have parallel memory accesses while others have more isolated long-latency accesses. The cost of idle cycles processor spends waiting for off-chip memory accesses is shared by parallel misses. However, misses which occur in isolation hurt the performance most. This adds another dimension to application's behavior. We propose an application behavior aware cache replacement policy to manage shared LLC. The proposed policy simultaneously reduces the negative interference among applications sharing the LLC and the miss-penalty associated with each LLC miss. Evaluation on SPEC CPU2006 benchmarks shows that our replacement policy improves performance on dual-core systems and quad-core system by up to 15.9% and 23.8% respectively over SRRIP for shared LLC. It is worth to note that effectiveness of our policy improves with the increase in the number of cores.
Suhit Pai, Newton, Virendra Singh
ACM Great Lakes Symposium on VLSI3
2018 A PAM-4 10S/12S line coding scheme with equi-probable levels
abstract
We propose an equi-probable line coding scheme for pulse amplitude modulation (PAM)-4. This encoding scheme can be used for improving the spectral efficiency of high speed serial links. Equi-probable PAM symbols also ensure atleast one zero-crossing transition in every encoded word, which is sufficient enough to recover the symbol clock from the received data. The equi-probable encoded words make design of the comparators at the receiver easy by aiding an automatic threshold tracking mechanism. This proposed encoding technique has a maximum contiguous symbol run length of 8, and ensures a DC balancing of encoded signals. The proposed 10S/12S encoding scheme has an overhead of 20%, as compared to the 25% overhead in the commonly used 8B/10B schemes.
Sandeep Goyal, Ron Joseph, Virendra Singh, Shalabh Gupta
ISCAS3
2018 In-situ Monitoring for Slack Time Violation Without Performance Penalty
abstract
High-performance requirements are forcing circuits to have lower combinational depth. At lower depth, degradation of performance can be significant due to multiplexer present in scan flip flop during functional mode. Therefore, exclusion of multiplexer from functional path has been a crucial factor for performance. Furthermore, newer technology nodes are continuously introduced to meet the demand of high-performance system on chips. However, variations, such as environmental variation (temperature) and aging variation (negative bias temperature instability) etc., affect temporal reliability significantly in recent technology nodes. Adding pessimistic timing margin to ensure the reliable working of a circuit under worst case can cause performance as well as power loss. In this paper, we propose an in-situ error prediction monitor with scan-in capability. The proposed cell uses extra master latch. During test mode, slave latch gets scan-in data from extra master latch whereas during functional mode this extra latch helps in error prediction. The proposed cell design adheres to traditional test generation and application. The detailed simulation demonstrates the effectiveness of the design.
Nihar Hage, Satyadev Ahlawat, Virendra Singh
ISCAS3
2018 On Securing Scan Design Through Test Vector Encryption
abstract
Scan-based side-channel attacks have been gaining a prominence among the malicious attackers. The unprotected scan chains are extremely vulnerable and could be exploited to extract the secret information from a security chip such as an Advanced Encryption Standard (AES) cryptochip. To protect the secret information from being hacked it is utmost necessary to redesign the scan chain with security features. In this paper, we propose a secure scan architecture aiming at the protection of AES cryptochips against scan-based attacks. The proposed idea is based on the principle of test pattern encryption. The major contribution of our architecture is area efficiency and its security features without hampering test, diagnose, and debug capability of the original scan chain. The experimental results and security analysis shows the efficacy of proposed design.
Darshit Vaghani, Satyadev Ahlawat, Jaynarayan T. Tudu, Masahiro Fujita 0004, Virendra Singh
ISCAS5
2018 Multiple Stuck-at Fault Testability Analysis of ROBDD Based Combinational Circuit Design
Toral Shah, Anzhela Yu. Matrosova, Masahiro Fujita 0004, Virendra Singh
J. Electron. Test.4
2017 On Securing Scan Design from Scan-Based Side-Channel Attacks
abstract
Test and diagnosis requirements has made the use of scan design unavoidable for present day highly complex circuits. However, scan design can be exploited to retrieve the secret information stored on a crypto chip by mounting scan based side channel attack. The scan design poses a threat to the security of crypto chips as it gives the user the capability to control/observe the circuit state. In this paper, we propose a technique to secure the scan design that can effectively defend the crypto chips against scan based side-channel attacks. To use the scan architecture the user first needs to supply the test authorization key. Once the user is authorized, the conventional test sequence can be started. Furthermore, the proposed technique allows using the original test set without any test time and test data overhead. In addition to that, the proposed technique leaves the debug capability intact and has a marginal area overhead.
Satyadev Ahlawat, Darshit Vaghani, Jaynarayan T. Tudu, Virendra Singh
ATS4
2017 An efficient test technique to prevent scan-based side-channel attacks
abstract
Almost every complex circuit today employs scan-based Design-for-Testability (DFT) architecture to enhance controllability and observability for every flip-flop in the design, thereby improve the testability. However, the DFT structure can also be exploited to mount side-channel attacks to retrieve the secret key stored in a cryptographic chip, thus compromising its security. In this paper, we propose a new secure scan test architecture which isolates the encryption key whenever the cryptographic chip is switched to test mode. The encryption key remains isolated during the whole scan test process. It also clears the last functional states of the security sensitive scan cells as soon as the scan test mode is activated. The proposed secure scan test approach is capable of exercising all kinds of conventional stuck-at and timing test. The proposed approach has minimal hardware overhead and it is capable of preventing existing scan-based side channel attacks.
Satyadev Ahlawat, Darshit Vaghani, Virendra Singh
ETS3
2017 Combining Restorability and Error Detection Ability for Effective Trace Signal Selection
abstract
Persistent growth in design complexity has led to increased chances of bugs appearing during post-silicon validation. Debugging errors at this stage requires some arrangement for expanding the observability of internal signals of the design. Limited number of trace buffers help in increasing visibility of these internal states. However, appropriate trace signal selection is very difficult. Restorability of untraced states with the help of traced ones is a popular approach for signal selection although it fails to address the main issue of error detection. We propose a signal selection methodology which combines the restoration capability and error detection ability of these signals. Experimental evaluation of the proposed signal selection approach on benchmark circuits indicates improved error detection for various kind of design errors. Practical consideration like minimizing routing overhead has also been analyzed.
Binod Kumar 0001, Ankit Jindal, Masahiro Fujita 0004, Virendra Singh
ACM Great Lakes Symposium on VLSI4
2017 EEAL: Processors' Performance Enhancement Through Early Execution of Aliased Loads
abstract
Due to increasing speed gap between DRAM and CPU, improving the performance of memory instructions has become very critical. Load and Store instructions account for almost 20-30%. It indicates that memories have become a bottleneck in CPU's performance. Conventionally load instructions, being more critical than stores, are to be executed as early as possible. Load-Store Queues (LSQs) are used for early execution of aliased loads by forwarding them results of the executed stores. Generally memory instructions, ready to be issued from Reservation Station (RS), have their source registers available however their memory addresses are not yet known. Hence, RS cannot detect aliasing loads/stores. We propose to insert a 2$^{nd}$ RS after the address computation stage in the memory pipeline. The 2$^{nd}$ RS helps in taking data forwarding decisions based on the computed addresses of the memory instructions. If the address of a load instruction aliases with preceding store, data can be directly forwarded to it. Such forwarded loads can bypass address translation and cache access stages, leading to their early execution. Our evaluation on a 4-wide Out-of-Order (OoO) core shows that the proposed architecture outperforms the baseline architecture by up to 7.3% (avg: 2.2%) with just 0.1% area overhead and 0.05% power overhead.
Abhishek Rajgadia, Newton, Virendra Singh
ACM Great Lakes Symposium on VLSI3
2017 DAAIP: Deadblock Aware Adaptive Insertion Policy for High Performance Caching
abstract
The commonly used LRU replacement policy for management of shared last-level cache (LLC) is not efficient as the policy is sharing-oblivious. LRU policy is suitable for applications which show a high-degree of data locality i.e. applications which are cache-friendly. However, applications with working data set greater than the available cache size or poor temporal locality perform poorly with LRU as most of the cache lines inserted by them simply traverse from MRU to LRU position without being re-referenced. Such applications are streaming in their cache behaviour and have very less data reuse. LRU policy makes inefficient use of shared caches for application mixes which are a combination of cache-friendly and streaming applications as the policy treats each cache line independently and doesn't learn from application's past cache reuse behaviour. We show that simple adaptive changes to the insertion policy can significantly improve system's performance. We propose Deadblock Aware Adaptive Insertion Policy (DAAIP) which dynamically adapts to the changing cache behaviour of applications sharing the LLC. DAAIP protects the data of application having high temporal locality from high access rate thrashing/streaming applications. Our proposed mechanism monitors each application at-runtime using cost-effective hardware circuits. The information collected is used to dynamically modify the insertion policy and implicitly partition the cache in favour of application showing more data locality. Our evaluation, with 39 multiprogrammed workloads, shows that DAAIP improves performance of dual-core system by up to 21% and on an average 5.8% over LLC caches managed by SRRIP replacement policy. We show that DAAIP also outperforms state-of-the-art cache replacement policy ABRip by 4.6% on system throughput metric.
Newton, Sujit Kr Mahto, Suhit Pai, Virendra Singh
ICCD4
2017 Revisiting random access scan for effective enhancement of post-silicon observability
abstract
Due to tremendous growth in complexity of modern designs, bugs inevitably escape the pre-silicon verification stage. This has led to considerable increase in the time and effort dedicated to post-silicon validation. Debugging designs at postsilicon stage faces a severe bottleneck of limited observability of the internal states. This paper presents a methodology for post-silicon debug utilizing the special features of progressive random access scan (PRAS). The PRAS offers a read-out of nondestructive scan values which is the bottleneck in the process of debugging. The proposed methodology avoids the large overhead of additional resources for debugging as the DfT architecture is reused. PRAS provides a simultaneous solution to the problems of power, data volume and application time during testing at the cost of routing overhead. The PRAS based proposed architecture offers visibility of internal states in fewer clock cycles than traditional serial scan chain based debug methods. The proposed debug scheme offers reconfigurability which enables selective visibility of internal states of a certain portion of the design. Experimental results indicate the better performance of the proposed methodology as compared to the state restoration based observability enhancement techniques.
Binod Kumar 0001, Ankit Jindal, Jaynarayan T. Tudu, Brajesh Pandey, Virendra Singh
IOLTS5
2017 Instruction-based self-test for delay faults maximizing operating temperature
abstract
In today's technology, reliability is one of the major challenges. Process variations and increasing power density in advanced technology nodes make the condition even worse. Process variations during manufacturing induce delay variations and such variations may manifest into faults at the higher temperature during operation. Delay faults at higher functional temperature is a critical factor for the reliability of a system. Current DFT-based methods do not handle this properly. In this paper, we propose an instruction based self-test (IBST) technique to elevate the temperature of a chip near to functional temperature and test it for delay faults at that temperature. In the first step, integer linear programming (ILP) is used to find out power hungry instructions. These instructions cause maximum toggling in a unit. In the second step, delay test instructions are combined to create a program for testing functional unit of an out-of-order superscalar processor. Experimentation of the proposed technique is carried out on portable ISA (PISA) based superscalar processor. Higher operating temperature 95°C ± 2°C is achieved and it is maintained while applying delay test.
Nihar Hage, Rohini Gulve, Masahiro Fujita 0004, Virendra Singh
IOLTS4
2017 Test pattern generation to detect multiple faults in ROBDD based combinational circuits
abstract
Deduced Ordered Binary Decision Diagram (ROBDD) based circuit syntheses is known to have complete testability under single stuck-at, multiple stuck-at and path delay fault models. In this paper we propose a test generation methodology for ROBDD based implementations. The test vectors required for all multiple stuck-at faults and delay faults are derived from the disjoint sum-of-products (DSOPs) representation thereby allowing test generation at design time with minimum effort.
Toral Shah, Anzhela Yu. Matrosova, Virendra Singh
IOLTS3
2017 A low cost technique for scan chain diagnosis
abstract
Scan based diagnosis plays a critical role in yield enhancement of sub-nanometer technology based chips. However, the scan chain itself can be subject to defects due to the large logic circuitry associated with it which constitute a significant fraction of total chip area. In some cases, it has been observed that scan chain failures may account for 30% to 50% of chip failures. Hence, scan chain testing and diagnosis has become very crucial in recent years. In this paper, we propose a hardware-assisted low cost and low complexity scan chain diagnosis technique. The proposed technique is very simple in operation and provides maximum resolution for stuck-at fault diagnosis.
Satyadev Ahlawat, Darshit Vaghani, Rohini Gulve, Virendra Singh
ISCAS4
2017 Exploiting path delay test generation to develop better TDF tests for small delay defects
abstract
Localized small delay defects, for example due to degraded transistor drive strength caused by a broken fin, are a growing concern in current FinFET and emerging gate all around (GAA) technologies. Such defects are currently targeted by timing-aware Transition Delay Fault (TDF) tests that aim to test the target nodes along the longest path. The resulting tests often require considerable test generation time, have high test data volume, and at times do not provide the desired coverage. In this paper, we show that Path Delay Fault (PDF) test generation can be exploited to not only generate the timing tests more efficiently, but the resulting TDF test sets are also more compact and perform better on commonly used delay test coverage metrics. This is because all TDF faults along a PDF targeted timing-critical path can be detected efficiently by generating a single PDF test. This efficiency is not explicitly exploited by node oriented TDF test generation even when the TDFs are targeted along the longest paths. We demonstrate the effectiveness of our methodology for a range of benchmark circuits by comparing the results from a commercial timing-aware ATPG (TA-ATPG) with our new approach that efficiently exploit PDF tests wherever possible. The proposed new approach results in approximately 12.5% reduction in pattern volume, 35% reduction in ATPG runtime and also a 5% improvement in delay test coverage (DTC) when compared to existing TA-ATPG approaches.
Ankush Srivastava, Adit D. Singh, Virendra Singh, Kewal K. Saluja
ITC3
2017 Improving post-silicon error detection with topological selection of trace signals
abstract
Drastic growth in design complexity of VLSI circuits has increased the chances of bugs escaping to first released silicon. This has resulted in an increased emphasis on post-silicon validation and debug which is typically hindered by limited observability of internal signals. Trace buffers assist in curbing this bottleneck by storing selected signal states for limited clock cycles. For efficient use of these on-chip buffers, devising a proper selection criterion is of utmost importance. Maximization of restoration of untraced signals is a widely utilized signal selection metric. However, this approach has been seen to not be very effective for error detection. This paper proposes a trace signal selection technique based on error transmission, taking into account the topology of the design. The proposed signal selection methodology can be effectively applied to trace as well as a combination of trace and scan based observability techniques. Experimental evaluation of the proposed methodology on different design errors indicates improvement in error detection as compared to restorability based selection techniques.
Binod Kumar 0001, Kanad Basu, Ankit Jindal, Masahiro Fujita 0004, Virendra Singh
VLSI-SoC5
2017 A Reliability-Aware Methodology to Isolate Timing-Critical Paths under Aging
Ankush Srivastava, Virendra Singh, Adit D. Singh, Kewal K. Saluja
J. Electron. Test.2
2016 A high performance scan flip-flop design for serial and mixed mode scan test
abstract
Over the years, serial scan design has became the defacto Design for Testability (DFT) technique. The ease of testing and high test coverage has made it to gain wide spread industrial acceptance. However, there are associated penalties with serial scan. These penalties include performance degradation, test data volume, test application time, and test power dissipation. The performance overhead of scan design is due to the scan multiplexers added to the inputs of every flip-flop. In today's very high speed designs with minimum possible combinational depth, the performance degradation caused by scan multiplexer has became magnified. Hence to maintain the circuit performance the timing overhead of scan design must be addressed. In this paper we propose a new scan flip-flop design that eliminates the performance overhead of serial scan. The proposed design removes the scan multiplexer off the functional path. The proposed design can help in improving the functional frequency of performance critical designs. Furthermore, the proposed design can be used as a common scan flip-flop in mixed mode scan test wherein it can be used as a serial scan cell as well as random access scan RAS) cell.
Satyadev Ahlawat, Jaynarayan T. Tudu, Anzhela Yu. Matrosova, Virendra Singh
IOLTS4
2016 REMO: Redundant execution with minimum area, power, performance overhead fault tolerant architecture
abstract
Relentless scaling in CMOS fabrication technology has made contemporary integrated circuits continue to evolve and grow in functionality with high clock frequencies and exponentially increasing transistor counts. However, it also makes them more susceptible to transient faults effectively decreasing their reliability. Therefore, ensuring correct and reliable operation of these microprocessors at low cost has become a challenging task. This paper proposes a light weight error detection method called REMO which aims to incorporate simple fault tolerance mechanisms as part of the basic architecture. It dynamically verifies the execution results of the instructions by exploiting spatial and temporal redundancy and detects soft errors. REMO shows that with minimal area, power and performance overhead, and a very low detection latency, a very high degree of fault coverage can be achieved. Our simulation results shows an increase in area is about 0.4%, power overhead near to 9% and a negligible performance penalty during fault free run.
Shoba Gopalakrishnan, Virendra Singh
IOLTS2
2015 A New Scan Flip Flop Design to Eliminate Performance Penalty of Scan
abstract
The demand for high performance system-on-chips (SoC) in communication and computing has been growing continuously. To meet the performance goals, very aggressive circuit design techniques such as the use of smallest possible logic depth are being practiced. Replacement of normal flip-flops with scan flip-flops adds an additional multiplexer delay to critical path. Furthermore as the combinational depth decreases, the performance degradation caused by scan multiplexer delay become more critical. Elimination of the scan multiplexer delay off the functional path has become crucial in maintaining the circuit performance. In this work we propose a new transistor level scan cell design to eliminate the scan multiplexer off the functional path. The proposed scan cell uses separate master latch for functional and test mode where as the slave latch is same in both the modes. Our proposed scan flip-flop fully comply with the conventional test flow. Post layout experimental results justify the effectiveness of the proposed scan cell design in eliminating the performance penalty of scan, and thus in improving the timing performance of integrated circuits.
Satyadev Ahlawat, Jaynarayan T. Tudu, Anzhela Yu. Matrosova, Virendra Singh
ATS4
2015 A Soft Error Resilient Low Leakage SRAM Cell Design
abstract
Semiconductor industry has been aggressively following the Moore's Law ever since its was proposed in the late sixties in its pursuit for smaller device sizes and higher performance metrics. However, this vigorous scaling has brought in several scaling induced side effects like single event upsets into the technology regime. SRAMs are highly susceptible to such upsets since they are designed at minimum device sizes to keep the on-chip memory density high. This paper presents a novel SEU-hardened SRAM cell employing single bitline. The proposed cell is 4 times more immune than a standard 6T-SRAM cell and also achieves 68% reduction in bitline leakage.
Adithyalal P. M, Shankar Balachandran, Virendra Singh
ATS3
2015 A Methodology for Identifying High Timing Variability Paths in Complex Designs
abstract
In some complex deep sub-micron designs, the variations in interconnect delay has a significant impact on the production yield of the product. In this paper, we develop a theoretical explanation for the unexpectedly higher process related timing variability shown by long interconnects that are driven by high drive strength gates. This gets even worse due to conventional gate delay variability and other random process effects. Our analysis is supported by actual silicon data and further validated by detailed Monte-Carlo (MC) simulations. Unfortunately, traditional scan based transition delay fault (TDF) timing tests can miss these variability induced delay faults on long interconnects which lies on the critical paths. We propose a methodology to identify high variability paths dominated by such long interconnects, with the aim of developing high quality delay timing tests. Specifically, we develop a heuristic based path selection algorithm to identify potentially slow paths that can contribute to test escapes in production. We further extend our approach to generate high quality delay timing tests for the target paths using the proposed "three pass" method.
Virendra Singh, Adit D. Singh, Kewal K. Saluja
ATS1
2015 Application behavior aware re-reference interval prediction for shared LLC
abstract
In modern CMPs, Last Level Cache (LLC) is shared among cores for better utilization. Interference among data, mapped from multiple cores, increases conflict misses in shared LLCs. Such interference is highly dependent on cache behavior of applications and access rate difference among them. We observe that interference among applications is not eliminated completely even using existing state-of-the-art mechanism for applications having high cache access rate difference and different memory characteristic. Applications with highly diverse cache behavior can be observed in homogeneous as well as heterogeneous multicore processors. Streaming applications, having high access rate, can still interfere with cache friendly applications having low access rate. We propose Application behavior aware replacement policy that predicts re-reference interval of the block based on block locality as well as application behavior. By providing more priority to application behavior over cache block locality, we reduce interference between streaming applications and cache friendly applications. Our evaluation on set of SPEC CPU2006 workloads running on CMP with shared LLC shows that proposed replacement policy outperforms the state-of-the-art replacement policy, on throughput metric. We achieve performance gain up to 16.2% over SRRIP for application mixes of cache-friendly and streaming applications. Our replacement policy achieves maximum of 59.9% reduction in number of misses as compared to SRRIP with average misses per kilo instructions (mpki) reduction of 5.9% over SRRIP.
Parth Lathigara, Shankar Balachandran, Virendra Singh
ICCD3
2012 SEU tolerant robust memory cell design
abstract
The implementation of semiconductor circuits and systems in nano-technology makes it possible to achieve high speed, lower voltage level and smaller area. The unintended and undesirable result of this scaling is that it makes integrated circuits susceptible to soft errors normally caused by alpha particle or neutron hits. These events of radiation strike resulting into bit upsets referred to as single event upsets(SEU), become increasingly of concern for the reliable circuit operation in the field. Storage elements are worst hit by this phenomenon. As we further scale down, there is greater interest in reliability of the circuits and systems, apart from the performance, power and area aspects. In this paper we propose an improved 12T SEU tolerant SRAM cell design. The proposed SRAM cell is economical in terms of area overhead. It is easy to fabricate as compared to earlier designs. Simulation results show that the proposed cell is highly robust, as it does not flip even for a transient pulse with 62 times the Qcritof a standard 6T SRAM cell.
Mohammed Shayan, Virendra Singh, Adit D. Singh, Masahiro Fujita 0004
IOLTS2
2012 Impact of process variations on computers used for image processing
abstract
Manufacturing process variations (PV) of transistors in the deep-submicron regime present the single biggest design challenge for large die size VLSI circuits such as processor arrays, GPUs, and FPGAs. However, there are a few applications in signal processing, such as image processing, and speech processing, where errors in computation by the underlying hardware could be tolerated or corrected off-line with readily available image restoration algorithms. In this paper, we qualitatively and quantitatively evaluate the effect of process variation in the underlying hardware (for different technology nodes) on a high level application program such as image processing. We rely on gate level simulation, of the data-path of an image processor comprising of a dedicated multiply-accumulate (MAC) array of size 256 × 256, with individual gate delays of the processor sampled from a delay distribution as appropriate for each technology node. Our results show that processing images with PV degraded hardware in technologies beyond 65nm is discernible to the human eye; image quality degrades further at 45nm, and is of unacceptable quality at 32nm and beyond. We also use image restoration algorithms to restore the images corrupted due to processing on PV degraded hardware. Our results show that with standard restoration algorithms, even images processed with high levels of PV (as in 32nm) can be restored to almost the same quality as the image processed on fault-free hardware.
Suraj Sindia, Foster F. Dai, Vishwani D. Agrawal, Virendra Singh
ISCAS4
2012 Efficient regular expression pattern matching for network intrusion detection systems using modified word-based automata
abstract
Network Intrusion Detection Systems (NIDS) intercept the traffic at an organization's network periphery to thwart intrusion attempts. Signature-based NIDS compares the intercepted packets against its database of known vulnerabilities and malware signatures to detect such cyber attacks. These signatures are represented using Regular Expressions (REs) and strings. Regular Expressions, because of their higher expressive power, are preferred over simple strings to write these signatures. We present Cascaded Automata Architecture to perform memory efficient Regular Expression pattern matching using existing string matching solutions. The proposed architecture performs two stage Regular Expression pattern matching. We replace the substring and character class components of the Regular Expression with new symbols. We address the challenges involved in this approach. We augment the Word-based Automata, obtained from the re-written Regular Expressions, with counter-based states and length bound transitions to perform Regular Expression pattern matching. We evaluated our architecture on Regular Expressions taken from Snort rulesets. We were able to reduce the number of automata states between 50% to 85%. Additionally, we could reduce the number of transitions by a factor of 3 leading to further reduction in the memory requirements.
Virendra Singh
SIN2
2012 Derating based hardware optimizations in soft error tolerant designs
abstract
Ensuring reliable operation over an extended period of time is one of the biggest challenges facing present day electronic systems. The increased vulnerability of the components to atmospheric particle strikes poses a big threat in attaining the reliability required for various mission critical applications. Various soft error mitigation methodologies exist to address this reliability challenge. A general solution to this problem is to arrive at a soft error mitigation methodology with an acceptable implementation overhead and error tolerance level. This implementation overhead can then be reduced by taking advantage of various derating effects like logical derating, electrical derating and timing window derating, and/or making use of application redundancy, e.g. redundancy in firmware/software executing on the so designed robust hardware. In this paper, we analyze the impact of various derating factors and show how they can be profitably employed to reduce the hardware overhead to implement a given level of soft error robustness. This analysis is performed on a set of benchmark circuits using the delayed capture methodology. Experimental results show upto 23% reduction in the hardware overhead when considering individual and combined derating factors.
Prasanth Viswanathan Pillai, Virendra Singh, Rubin A. Parekhji
VTS2
2012 Defect Level and Fault Coverage in Coefficient Based Analog Circuit Testing
Suraj Sindia, Vishwani D. Agrawal, Virendra Singh
J. Electron. Test.3
2012 Parametric Fault Testing of Non-Linear Analog Circuits Based on Polynomial and V-Transform Coefficients
Suraj Sindia, Vishwani D. Agrawal, Virendra Singh
J. Electron. Test.3
2011 SSTKR: Secure and Testable Scan Design through Test Key Randomization
abstract
Scan test is the standard method, practiced by industry, that has consistently provided high fault coverage due to high controllability and high observability. The scan chain allows to control and observe the internal signals of a chip. However, this property also facilitates hackers to use scan architecture as a means to breach chip security. This paper addresses this issue by proposing a new method called Secure and testable Scan design through Test Key Randomization(SSTKR). SSTKR is a key based method to prevent hackers from stealing secret information. Linear Feedback Shift Register (LFSR) is used to generate authentication keys to be embedded in test vectors. Unique key is used for every test vector which prevents scan based side channel attacks effectively. Any attempt to steal secret information will lead to a randomized response. SSTKR has very low area and test time overhead without performance penalty. Our approach also facilitates in-field test of the chip.
Mohammed Abdul Razzaq, Virendra Singh, Adit D. Singh
Asian Test Symposium2
2011 Test and Diagnosis of Analog Circuits Using Moment Generating Functions
abstract
The function of a circuit under test (CUT) is represented as a transformation on the probability density function of its input excitation, which is a continuous random variable (RV) with Gaussian probability distribution. Probability moments of the output, now a transformed RV, are used as metrics for testing catastrophic and parametric faults in circuit components. The proposed use of probability moments as test metrics with white noise excitation as input addresses three important problems of analog circuit test, namely, it 1) reduces complexity of input signal design, 2) increases resolution of fault detection, and 3) reduces production test cost as it has no area overhead and may even marginally reduce the test time. We also propose a method to diagnose circuit elements with catastrophic faults based on unique relationships between specific moments of the output and circuit elements. We present a theoretical framework, test and diagnosis procedures and SPICE simulation results for a benchmark elliptic filter and a low noise amplifier. We are able to detect all catastrophic faults and single components that deviate from their nominal values by just over 10%. We diagnose all catastrophic faults in the example circuits.
Suraj Sindia, Vishwani D. Agrawal, Virendra Singh
Asian Test Symposium3
2011 Parallelizing TUNAMI-N1 Using GPGPU
abstract
We present a high performance tsunami-prediction system using General Purpose Graphics Processing Units (GPGPU). It is based on TUNAMI-N1, a Numerical Analysis Model for Investigation of near-field tsunamis. It uses linear shallow water wave equations, commonly accepted approximation for tsunami propagation, taking the input from a bathymetry file containing a large data set. Due to the largeness of the data set, the model is more amenable to parallelization. The system maps the TUNAMI-N1 model into the massively parallel GPU architecture using Nvidia CUDA framework. It employs multiple kernels that contain inherently parallel portion of the model and uses the concepts of data and hybrid parallelism to fully exploit the hardware capabilities of the GPUs. Experimental results show that our system achieves a speed up of six times.
Harsh Gidra, Israrul Haque, Nitin P. Kumar, M. Sargurunathan, Manoj Singh Gaur, Vijay Laxmi, Mark Zwolinski, Virendra Singh
HPCC8
2011 Adaptive execution assistance for multiplexed fault-tolerant chip multiprocessors
abstract
Relentless scaling of CMOS fabrication technology has made contemporary integrated circuits increasingly susceptible to transient faults, wearout-related permanent faults, intermittent faults and process variations. Therefore, mechanisms to mitigate the effects of decreased reliability are expected to become essential components of future general-purpose microprocessors. In this paper, we introduce a new throughput-efficient architecture for multiplexed fault-tolerant chip multiprocessors (CMPs). Our proposal relies on the new technique of adaptive execution assistance, which dynamically varies instruction outcomes forwarded from the leading core to the trailing core based on measures of trailing core performance. We identify policies and design low overhead hardware mechanisms to achieve this. Our work also introduces a new priority-based thread-scheduling algorithm for multiplexed architectures that improves multiplexed fault tolerant CMP throughput by prioritizing stalled threads. Through simulation-based evaluation, we And that our proposal delivers 17.2% higher throughput than perfect dual modular redundant (DMR) execution and outperforms previous proposals for throughput-efficient CMP architectures.
Pramod Subramanyan, Virendra Singh, Kewal K. Saluja, Erik Larsson
ICCD2
2011 Reduced overhead soft error mitigation using error control coding techniques
abstract
Soft errors are one of the biggest reliability challenges for present day electronic devices. With technology scaling, the contribution of soft errors to overall device failure is on the rise and it is becoming the dominant reliability failure mechanism. Several techniques exist for the detection and correction of soft errors. Reducing implementation overhead is one of the areas which researchers were focusing on, and several optimization techniques are being proposed. In this paper, we propose a novel methodology, using error detection and correction codes to reduce the implementation overhead. We extend the earlier work on delayed capture methodology, and divide the total number of flip-flops into various groups and calculate the check bits for each group. This method exploits the reduction in the fault space which is generated due to single event upsets, and illustrates how the detection and correction implementation overheads can be minimized. Experimental results highlight the effectiveness of this technique. As compared to the original implementation, 44.80% reduction in area is obtained, without sacrificing the coverage.
Prasanth Viswanathan Pillai, Virendra Singh, Rubin A. Parekhji
IOLTS2
2011 C-Routing: An adaptive hierarchical NoC routing methodology
abstract
Deterministic routing algorithms are easier to design and implement in NoC but these fail to adapt to congestion. Table based adaptive routing solutions are not scalable. As the number of nodes increases, the area required for routing table becomes a penalty. In this paper, we propose a new hierarchical cluster based adaptive routing called `C-Routing' in 2-D Mesh NoC. The solution reduces routing table size and provides deadlock freedom without use of virtual channels while ensuring livelock free routing. Routers in our method use intelligent routing to route information between the processing elements ensuring the correctness, deadlock freeness, and congestion handling. This method has been evaluated against other adaptive algorithms such as PROM, and Q-Routing etc. Results show that the proposed method performs better for given traffic patterns. C-routing uses adaptivity to avoid congestion by uniform distribution of traffic among the cores by sending flits over two different paths to the destination.
Manas Kumar Puthal, Virendra Singh, Manoj Singh Gaur, Vijay Laxmi
VLSI-SoC2
2011 Non-linear analog circuit test and diagnosis under process variation using V-Transform coefficients
abstract
Parametric fault testing of non-linear analog circuits based on a new mathematical transform is presented. The V-Transform acts on the polynomial expansion of the circuit's function. Its main properties are: 1) to make the polynomial coefficients monotonic, 2) to reduce masking of parametric faults due to process variation, and 3) to increase the sensitivity of polynomial coefficients to the circuit parameter variation, thus enhancing diagnostic resolution. We show that the sensitivity of V-Transform Coefficients (VTC) with respect to circuit parameter variation is up to 3 to 5 times greater than the sensitivity of polynomial coefficients. Fault diagnosis of parametric faults under process variation using VTC is then presented. We also propose a scheme to distinguish between circuit specifications failures due to process variation versus manufacturing defects which manifest as parametric faults. To validate our approach, we apply the test and diagnosis procedures to a benchmark fifth order elliptic filter. We use SPICE program for fault injection, with about 50,000 Monte Carlo simulation runs to demonstrate fault detection-diagnosis under process variation. The test scheme uncovers 95% of all injected single parametric faults whose sizes deviate 5% from the nominal values of circuit components corrected for process variation, while the procedure successfully diagnosed all component faults under ±3σ process variation with 88% confidence level.
Suraj Sindia, Vishwani D. Agrawal, Virendra Singh
VTS3
2010 Modified Scan Flip-Flop for Low Power Testing
abstract
Scanning of test vectors during testing causes unnecessary and excessive switching in the combinational circuit compared to that in the normal operation. In this paper, we propose a modified design of a scan flip-flop which eliminates the power consumed due to unnecessary switching in the combinational circuit during scan shift, with a little impact on performance. The new scan flip-flop disables the slave latch during scan, and uses an alternate low cost dynamic latch in the scan path instead. Methods for generating slave latch disable control signal are also presented.
Amit Mishra 0002, Nidhi Sinha, Satdev, Virendra Singh, Sreejit Chakravarty, Adit D. Singh
Asian Test Symposium4
2010 Multiplexed redundant execution: A technique for efficient fault tolerance in chip multiprocessors
abstract
Continued CMOS scaling is expected to make future microprocessors susceptible to transient faults, hard faults, manufacturing defects and process variations causing fault tolerance to become important even for general purpose processors targeted at the commodity market. To mitigate the effect of decreased reliability, a number of fault-tolerant architectures have been proposed that exploit the natural coarse-grained redundancy available in chip multiprocessors (CMPs). These architectures execute a single application using two threads, typically as one leading thread and one trailing thread. Errors are detected by comparing the outputs produced by these two threads. These architectures schedule a single application on two cores or two thread contexts of a CMP. As a result, besides the additional energy consumption and performance overhead that is required to provide fault tolerance, such schemes also impose a throughput loss. Consequently a CMP which is capable of executing 2n threads in non-redundant mode can only execute half as many (n) threads in fault-tolerant mode. In this paper we propose multiplexed redundant execution (MRE), a low-overhead architectural technique that executes multiple trailing threads on a single processor core. MRE exploits the observation that it is possible to accelerate the execution of the trailing thread by providing execution assistance from the leading thread. Execution assistance combined with coarse-grained multithreading allows MRE to schedule multiple trailing threads concurrently on a single core with only a small performance penalty. Our results show that MRE increases the throughput of fault-tolerant CMP by 16% over an ideal dual modular redundant (DMR) architecture.
Pramod Subramanyan, Virendra Singh, Kewal K. Saluja, Erik Larsson
DATE2
2010 Energy-efficient fault tolerance in chip multiprocessors using Critical Value Forwarding
abstract
Relentless CMOS scaling coupled with lower design tolerances is making ICs increasingly susceptible to wear-out related permanent faults and transient faults, necessitating on-chip fault tolerance in future chip microprocessors (CMPs). In this paper we introduce a new energy-efficient fault-tolerant CMP architecture known as Redundant Execution using Critical Value Forwarding (RECVF). RECVF is based on two observations: (i) forwarding critical instruction results from the leading to the trailing core enables the latter to execute faster, and (ii) this speedup can be exploited to reduce energy consumption by operating the trailing core at a lower voltage-frequency level. Our evaluation shows that RECVF consumes 37% less energy than conventional dual modular redundant (DMR) execution of a program. It consumes only 1.26 times the energy of a non-fault-tolerant baseline and has a performance overhead of just 1.2%.
Pramod Subramanyan, Virendra Singh, Kewal K. Saluja, Erik Larsson
DSN2
2010 Modified T-Flip-Flop based scan cell for RAS
abstract
Testing using a random access scan (RAS) design-for-test approach is experiencing renewed interest because of the potential for lower test application time, low power dissipation, and low test data volume compared to standard serial scan. In this paper we propose a significant modification and enhancement to the T-Flip-Flop based cell design for Random Access Scan (RAS). Importantly, the new RAS cell can allow the overlap of the test response read out with the loading of the next test input patterns within the same memory addressing cycle, thereby masking out the need for a separate memory cycle to read the test response in many cases. This can greatly reduce test application time. Experimental results show that the Modified T-Flip-Flop based scan cell is able to mask about 33% to 76% of reads. Further, this new RAS cell also eliminates the need for clock gating and additionally achieves reduction in gate overhead as much as about 20% compared to the existing T-flip-flop based RAS cell design.
Raghavendra Adiga, Gandhi Arpit, Virendra Singh, Kewal K. Saluja, Adit D. Singh
ETS3
2010 Scan cell reordering to minimize peak power during test cycle: A graph theoretic approach
abstract
Scan circuit is widely practiced DFT technology. The scan testing procedure consist of state initialization, test application, response capture and observation process. During the state initialization process the scan vectors are shifted into the scan cells and simultaneously the responses captured in last cycle are shifted out. During this shift operation the transitions that arise in the scan cells are propagated to the combinational circuit, which inturn create many more toggling activities in the combinational block and hence increases the dynamic power consumption. The dynamic power consumed during scan shift operation is much more higher than that of normal mode operation. Due to change in design characteristic the dynamic power dissipated during scan operation becomes an important issue. The average power and peak power are the standard metric to measure dynamic power. During scan test both average power and peak power are required to be within the specified power budget for safe testing of chip. Average power causes excessive heat dissipation where as peak power causes IR drop and cross talk problem. Particularly, the excessive peak power during test-cycle of at-speed testing is vulnerable. The excessive peak power causes high rate of current in the power and ground rails which decreases the supply voltage and causes ground bounce, this phenomenon is known as IR-drop. The larger IR-drop means the worse speed performance of circuit. This degradation in performance grows if circuit is operated at high frequency which is the case during at-speed testing. This degradation in performance leads to incorrect capture of responses and this results in to undesired yield loss. Hence, to avoid yield loss the the peak power minimization is necessary especially in case of narrow test-cycle. More over the minimization of peak power is also advantageous for parallel testing of multiple core to reduce test time. In this work we have focused on the problem of peak power consumption during test-cycle for at-speed testing. The methodology proposed in this work is based on scan cells reordering. Many direction has been explored to reduce peak power during test-cycle. One of the methodology on scan reordering is proposed by Bonhomme et al. The methodology is formulated as a global optimization problem and solved using simulated annealing approach. Although the simulated annealing can provides near optimal solution if it is allowed to run for sufficient number of iteration the graph theoretic formulation will wider the solution space for scan reordering methodology. With this motivation we are proposing a graph theoretic formulation for scan reordering methodology to minimize peak power during test-cycle. The overall approach consists of graph theoretic problem formulation and an algorithm to solve it. From given scan related informations viz. scan cells, possible scan path, and power consumption a complete vector-weighted graph is constructed. The vector-weight is a weight of an edge which keeps the information of peak power consumed by each test vector. On this graph a TSP (Travelling Sales Person) problem is formulated. The cost function in this formulation is peak power. The problem formulated is NP-complete. As the problem is NP-complete we have proposed a greedy based heuristic to solve it. The proposed heuristic consists of two parts. Part 1 to find a Hamiltonian cycle which consume less peak power from the constructed complete graph and Part 2 to find a Hamiltonian path having lower peak power from Hamiltonian cycle. The Part 1 of algorithm runs in polynomial time and the Part 2 runs in linear time. The memory space required to execute these algorithms is also linear. The experiment conducted on ITC99 and ISCAS89 benchmarks show that the proposed methodology is able to reduce appreciable percentage (around 55%) of peak power compared to. Overall, this paper has proposed a novel way of formulating a graph theoretic problem for scan reordering to minimize test-cycle peak power. The scan reordering methodology may incur nominal area overhead in terms of routing and may alter the delay fault coverage for at-speed skewed-load testing. In this work we have not taken these parameters into account. However, the proposed methodology can be extended to consider these parameters. One limitation of the scan reordering methodology is it is pattern dependent. If some additional pattern has to be added on top of the existing patterns the methodology will not be able to reduce peak power effectively. This issue needs further examination.
Jaynarayan T. Tudu, Erik Larsson, Virendra Singh, Hideo Fujiwara
ETS3
2010 SEU tolerant SRAM for FPGA applications
abstract
Modern integrated circuits require careful attention to the soft errors resulting into bit upsets, which are normally caused by alpha particle or neutron hits. These events, also referred to as single-event upsets (SEUs), will become more severe for future technologies. LUT-based FPGAs are heavily using SRAM and there is a growing concern on correct operations of such FPGAs. Although there have been researches on enhancing fault tolerance of such FPGAs, they are based on TMR (triple modular redundancy) and simply too costly for normal application. In this paper we propose a novel 10T SEU tolerant SRAM cell and discuss its modifications for storage of configuration bits in FPGA so that reasonable protection against soft errors can be achieved with small area increase.
Sudipta Sarkar, Anubhav Adak, Virendra Singh, Kewal K. Saluja, Masahiro Fujita 0004
FPT3
2010 Energy-efficient redundant execution for chip multiprocessors
abstract
Relentless CMOS scaling coupled with lower design tolerances is making ICs increasingly susceptible to wear-out related permanent faults and transient faults, necessitating on-chip fault tolerance in future chip microprocessors (CMPs). In this paper, we describe a power-efficient architecture for redundant execution on chip multiprocessors (CMPs) which when coupled with our per-core dynamic voltage and frequency scaling (DVFS) algorithm significantly reduces the energy overhead of redundant execution without sacrificing performance. Our evaluation shows that this architecture has a performance overhead of only 0.3% and consumes only 1.48 times the energy of a non-fault-tolerant baseline.
Pramod Subramanyan, Virendra Singh, Kewal K. Saluja, Erik Larsson
ACM Great Lakes Symposium on VLSI2
2010 Graph theoretic approach for scan cell reordering to minimize peak shift power
abstract
Scan circuit testing generally causes excessive switching activity compared to normal circuit operation. This excessive switching activity causes high peak and average power consumption. Higher peak power causes, supply voltage droop and excessive heat dissipation. This paper proposes a scan cell reordering methodology to minimize the peak power consumption during scan shift operation. The proposed methodology first formulate the problem as graph theoretic problem then solve it by a linear time heuristic. The experimental results show that the methodology is able to reduce up to 48% of peak power in compared to the solution provided by industrial tool.
Jaynarayan T. Tudu, Erik Larsson, Virendra Singh, Hideo Fujiwara
ACM Great Lakes Symposium on VLSI3
2010 Robust detection of soft errors using delayed capture methodology
abstract
With the scaling of technology node and voltage levels, the susceptibility of logic to soft errors is increasing. Hence it is very important to take care of soft errors in the combinational logic along with those in the sequential elements. In this paper, a novel method is proposed to detect the presence of soft errors in both combinational and sequential logic. In this method, flip-flops are grouped and parity is computed for each group twice - once at the input of the flip-flops and next at the output. Later, the parity at the inputs and outputs is compared to detect the presence of soft errors. The effectiveness of the technique is shown through experimental results.
Prasanth Viswanathan Pillai, Virendra Singh, Rubin A. Parekhji
IOLTS2
2010 Test application time minimization for RAS using basis optimization of column decoder
abstract
Random Access Scan, which addresses individual flip-flops in a design using a memory array like row and column decoder architecture, has recently attracted widespread attention, due to its potential for lower test application time, test data volume and test power dissipation when compared to traditional Serial Scan. This is because typically only a very limited number of random "care" bits in a test response need be modified to create the next test vector. Unlike traditional scan, most flip-flops need not be updated. Test application efficiency can be further improved by organizing the access by word instead of by bit. In this paper we present a new decoder structure that takes advantage of basis vectors and linear algebra to further significantly optimize test application in RAS by performing the write operations on multiple bits consecutively. Simulations performed on benchmark circuits show an average of 2-3 times speed up in test write time compared to conventional RAS.
A. Abhishek, Amanulla Khan, Virendra Singh, Kewal K. Saluja, Adit D. Singh
ISCAS3
2010 Genetic algorithm based topology generation for application specific Network-on-Chip
abstract
Network-on-Chip (NoC) has been proposed as a solution for the communication challenges of System-on-chip (SoC) design in nanoscale technologies. Application specific SoC design offers the opportunity for incorporating custom NoC architectures that are more suitable for a particular application, and do not necessarily conform to regular topologies. The aim is to generate a custom NoC that maximizes performance under the given resource constraints. The paper presents a heuristic technique based on genetic algorithm for synthesis of custom NoC architectures along with requisite routing tables with the objective to improve communication load distributions in the network subject to the resource constraints in such a way that the overall communication throughput and latency improves.
Naveen Choudhary, Manoj Singh Gaur, Vijay Laxmi, Virendra Singh
ISCAS4
2009 Leveraging Partially Enhanced Scan for Improved Observability in Delay Fault Testing
abstract
Enhanced scan design can significantly improve the fault coverage for two pattern delay tests at the cost of exorbitantly high area overhead. The redundant flip-flops introduced in the scan chains have traditionally only been used to launch the two-pattern delay test inputs, not to capture tests results. This paper presents a new, much lower cost partial enhanced scan methodology with both improved controllability and observability. Facilitating observation of some hard to observe internal nodes by capturing their response in the already available and underutilized redundant flip-flops improves delay fault coverage with minimal or almost negligible cost. Experimental results on ISCAS'89 benchmark circuits show significant improvement in TDF fault coverage for this new partial enhance scan methodology.
K. G. Deepak, Robinson Reyna, Virendra Singh, Adit D. Singh
Asian Test Symposium3
2009 Multi-tone Testing of Linear and Nonlinear Analog Circuits Using Polynomial Coefficients
abstract
A method of testing for parametric faults of analog circuits based on a polynomial representation of fault-free function of the circuit is presented. The response of the circuit under test (CUT) is estimated as a polynomial in the applied input voltage at relevant frequencies in addition to DC. Classification of CUT is based on a comparison of the estimated polynomial coefficients with those of the fault free circuit. This testing method requires no design for test hardware as might be added to the circuit by some other methods. The proposed method is illustrated for a benchmark elliptic filter. It is shown to uncover several parametric faults causing deviations as small as 5% from the nominal values.
Suraj Sindia, Virendra Singh, Vishwani D. Agrawal
Asian Test Symposium2
2009 Fault-tolerant average execution time optimization for general-purpose multi-processor system-on-chips
abstract
Fault-tolerance is due to the semiconductor technology development important, not only for safety-critical systems but also for general-purpose (non-safety critical) systems. However, instead of guaranteeing that deadlines always are met, it is for general-purpose systems important to minimize the average execution time (AET) while ensuring fault-tolerance. For a given job and a soft (transient) error probability, we define mathematical formulas for AET that includes bus communication overhead for both voting (active replication) and rollback-recovery with checkpointing (RRC). And, for a given multi-processor system-on-chip (MPSoC), we define integer linear programming (ILP) models that minimize AET including bus communication overhead when: (1) selecting the number of checkpoints when using RRC, (2) finding the number of processors and job-to-processor assignment when using voting, and (3) defining fault-tolerance scheme (voting or RRC) per job and defining its usage for each job. Experiments demonstrate significant savings in AET.
Mikael Väyrynen, Virendra Singh, Erik Larsson
DATE2
2009 On Minimization of Peak Power for Scan Circuit during Test
abstract
Scan circuit generally causes excessive switching activity compared to normal circuit operation. The higher switching activity in turn causes higher peak power supply current which results into supply voltage droop and eventually yield loss. This paper proposes an efficient methodology for test vector re-ordering to achieve minimum peak power supported by the given test vector set. The proposed methodology also minimizes average power under the minimum peak power constraint. A methodology to further reduce the peak power, below the minimum supported peak power, by inclusion of minimum additional vectors is also discussed. The paper defines the lower bound on peak power for a given test set. The results on several benchmarks shows that it can reduce peak power by up to 27%.
Jaynarayan T. Tudu, Erik Larsson, Virendra Singh, Vishwani D. Agrawal
ETS3
2009 DX-compactor: distributed X-compaction for SoCs
abstract
The emergence of System-on-Chip (SoC) devices has led to a complex on-chip interconnect structure that consumes significant area. Distributed compaction is a test response compaction scheme for an SoC that aims at reducing the area occupied for the purpose of testing the chip. This technique involves the design of compactors for individual cores on the chip. These are interconnected suitably to achieve the required functionality while reducing the area overhead by decreasing the length and the number of interconnects that are routed from scan chain outputs to the output pins. The distributed compaction technique matches the performance of an X-Compactor in terms of error detection and X-masking.
Reshma C. Jumani, Niraj Bharatkumar Jain, Virendra Singh, Kewal K. Saluja
ACM Great Lakes Symposium on VLSI3
2009 Polynomial coefficient based DC testing of non-linear analog circuits
abstract
DC testing of parametric faults in non-linear analog circuits based on polynomial approximation of the functionality of fault free circuit is presented. Classification of circuit under test (CUT) is based on comparison of estimates of polynomial coefficients with those of the fault free circuit. The method needs very little augmentation of circuit to make it testable as only output parameters are used for classification. Possible fault diagnosis using the proposed method in conjunction with sensitivity of polynomial coefficients is also presented.
Suraj Sindia, Virendra Singh, Vishwani D. Agrawal
ACM Great Lakes Symposium on VLSI2
2006 Instruction-Based Self-Testing of Delay Faults in Pipelined Processors
abstract
Aggressive processor design methodology using high-speed clock and deep submicrometer technology is necessitating the use of at-speed delay fault testing. Although nearly all modern processors use pipelined architecture, no method has been proposed in literature to model these for the purpose of test generation. This paper proposes a graph theoretic model of pipelined processors and develops a systematic approach to path delay fault testing of such processor cores using the processor instruction set. The proposed methodology generates test vectors under the extracted architectural constraints. These test vectors can be applied in functional mode of operation, hence, self-test becomes possible. Self-test in a functional mode can also be used for online periodic testing. Our approach uses a graph model for architectural constraint extraction and path classification. Test vectors are generated using constrained automatic test pattern generation (ATPG) under the extracted constraints. Finally, a test program consisting of an instruction sequence is generated for the application of generated test vectors. We applied our method to two example processors, namely a 16-bit 5-stage VPRO pipelined processor and a 32-bit pipelined DLX processor, to demonstrate the effectiveness of our methodology
Virendra Singh, Michiko Inoue, Kewal K. Saluja, Hideo Fujiwara
IEEE Trans. Very Large Scale Integr. Syst.1
2005 Testing Superscalar Processors in Functional Mode
abstract
This paper presents a methodology for testing a superscalar processor using functional mode of operation for the performance oriented delay faults. The functional mode test issues for superscalar are discussed. A graph based model is developed and used to develop for the generation of test programs.
Virendra Singh, Michiko Inoue, Kewal K. Saluja, Hideo Fujiwara
FPL1
2003 Software-Based Delay Fault Testing of Processor Cores
abstract
This paper presents a software-based self-testing methodology for delay fault testing. Delay faults affect the circuit functionality only when it can be activated in functional mode. A systematic approach or the generation of test vectors, which are applicable in functional mode, is presented. A graph theoretic model (represented by IE-Graph) is developed in order to model the datapath. A finite state machine model is used for the controller. These models are used for constraint extraction so that the generated test can be applied in functional mode.
Virendra Singh, Michiko Inoue, Kewal K. Saluja, Hideo Fujiwara
Asian Test Symposium1