Raghunandana K. K

dblp:331/7078 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 GPD: Predictive Control Flow Error Detection Leveraging Data Flow Error Detection Methods
abstract
Soft errors in General Purpose Graphics Processing Units (GPGPUs) result in data or control flow errors. Error detection and correction methods for data and control flow errors are orthogonal, and these methods incur separate area, power, and performance overheads. This paper proposes a low-overhead predictive control flow error detection method called GPGPU Predictive Detector (GPD), which leverages data flow error detection and correction methods to detect control flow errors. GPD is non-intrusive to application software and transparent to users. GPD is built on earlier work on data flow error detection and correction methods DDSR and TREFU. In GPD, DDSR and TREFU combined architecture protects all non-control flow instructions. The control flow error is detected by calculating the address of the instruction that succeeds the control flow instruction in advance and comparing it with the actual address it accesses. The effectiveness of GPD has been shown through a set of ISPASS-2009 and RODINIA benchmarks. Relative to a non-fault-tolerant GPGPU architecture, GPD has a performance overhead of 5% and average and peak power overheads of 4% and 3%, respectively. We prove through induction that the GPD provides fault coverage against GPGPU control flow and data flow errors.
Raghunandana K. K, Yogesh Prasad K. R, Matteo Sonza Reorda, Virendra Singh
IOLTS1
2024 TCC: GPGPU Architecture for Instruction Decoder and Control Flow Error Detection
abstract
The devices fabricated with the latest sub-nanometer technology node have a higher probability of parametric and wear-out failures, operational faults, and manufacturing defects, and these devices are more susceptible to intrinsic and extrinsic noise, resulting in soft errors. The parts with manufacturing defects are generally screened out during end-of-manufacturing tests. Thus, the soft errors during normal operations are of great concern. The soft errors in GPGPUs result into silent data corruption and control flow divergence errors. In order to deal with this, dual and triple modular redundancy architectures are used for soft error detection and correction, which result in large areas and power overheads. To overcome this, we propose a low overhead fault-tolerant microarchitecture called Trace Consistency Check (TCC) to detect the decoder and control flow divergence errors. The TCC is transparent to the application software. For error detection, we exploit the execution model of GPGPUs, where the warps of kernel executing in the streaming multiprocessor have temporal execution repetition. Hence, the instruction execution trace and control divergence paths across the warps are consistent. Inconsistency across warps for the same code region is attributed to decoder or control divergence errors. For error detection, new microarchitecture structures called Execution Trace buffer and Control Divergence Trace buffer were introduced to store and check the trace consistency across warps. The performance of TCC is evaluated through the ISPASS 2009 and RODINIA benchmarks. TCC's error detection capability and power overheads are evaluated. The simulation results show that TCC detects greater than 99% decoder and control flow errors with low power and no performance overheads.
Raghunandana K. K, Yogesh Prasad K. R, Matteo Sonza Reorda, Virendra Singh
DDECS1
2023 TREFU: An Online Error Detecting and Correcting Fault Tolerant GPGPU Architecture
abstract
General Purpose Graphics Processing Units (GPGPUs) are extensively used in high-performance applications/systems, whose execution times may vary from a few days to months. Many times, these systems are expected to provide high reliability and availability. On the other hand, the high-throughput GPGPUs are fabricated with the latest cutting-edge technology. The shrinking transistor feature size and aggressive voltage scaling resulted in increased susceptibility to soft errors. Hence, GPGPU execution results cannot be trusted. This necessitates the employment of error detection and correction methods for reliable results. To mitigate soft error effects in the GPGPU execution pipeline, we propose a fault-tolerant microarchitecture called Triple modular Redundant Execution with idle Functional Units (TREFU) to detect and correct errors online. The proposed method is transparent to the application software. A new microarchitecture structure, replay buffer, is introduced to store temporary operands and results and used as a checkpoint. On error detection, the data in the duplicate copy of the replay buffers are used for Triple Modular Redundant (TMR) execution and error correction. The effectiveness of TREFU is demonstrated through the ISPASS 2009 and RODINIA benchmarks. TREFU's performance and power overheads are evaluated for an error-free run and at various error rates of executed instructions ranging from 1 to 50K. The simulation results show that complete error detection and correction across all threads can be achieved with a mean performance overhead of 4%, an average power overhead of 4%, and a peak power overhead of 5%.
Raghunandana K. K, B. K. S. V. L. Varaprasad, Matteo Sonza Reorda, Virendra Singh
IOLTS1