Inderpreet Singh

dblp:86/7883 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
2since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 34% Parallel and multicore computing · 33% GPUs and heterogeneous computing · 29%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache coherence
0.212013
Cache coherence for GPU architectures · HPCA 2013
Memory systems › cache coherence › cache coherence protocol
GPU coherence protocol
0.212013
Cache coherence for GPU architectures · HPCA 2013
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy
0.212013
Cache coherence for GPU architectures · HPCA 2013
Parallel and multicore computing › transactional memory
conflict detection
0.112011
Hardware transactional memory for GPU architectures · MICRO 2011
GPUs and heterogeneous computing
GPU transactional memory
0.112011
Hardware transactional memory for GPU architectures · MICRO 2011
Parallel and multicore computing › transactional memory
hardware transactional memory
0.112011
Hardware transactional memory for GPU architectures · MICRO 2011
Parallel and multicore computing
transactional memory
0.112011
Hardware transactional memory for GPU architectures · MICRO 2011
Processor architecture and microarchitecture
chip multiprocessor
0.012013
Cache coherence for GPU architectures · HPCA 2013
Memory systems › cache coherence
directory-based coherence
0.012013
Cache coherence for GPU architectures · HPCA 2013
GPUs and heterogeneous computing › GPU microarchitecture
SIMT architecture
0.012011
Hardware transactional memory for GPU architectures · MICRO 2011

Methods — techniques the papers use, named apart from their topics

write-through protocol · 0.2synchronized counters · 0.2speculative validation · 0.1bloom filter validation · 0.1
YearPublicationVenuePosition
2026 Object detection for cross-linguistic vowel analysis: A novel language-agnostic method for forensic speech processing
Soham Gangopadhyay, Inderpreet Singh, Prateek Pandya, Ashish Mani, Sumit Goswami
Speech Commun.2
2024 Modified YOLOv5 for small target detection in aerial images
Inderpreet Singh, Geetika 0001
Multim. Tools Appl.1
2013 Cache coherence for GPU architectures
abstract
While scalable coherence has been extensively studied in the context of general purpose chip multiprocessors (CMPs), GPU architectures present a new set of challenges. Introducing conventional directory protocols adds unnecessary coherence traffic overhead to existing GPU applications. Moreover, these protocols increase the verification complexity of the GPU memory system. Recent research, Library Cache Coherence (LCC) [34, 54], explored the use of time-based approaches in CMP coherence protocols. This paper describes a time-based coherence framework for GPUs, called Temporal Coherence (TC), that exploits globally synchronized counters in single-chip systems to develop a streamlined GPU coherence protocol. Synchronized counters enable all coherence transitions, such as invalidation of cache blocks, to happen synchronously, eliminating all coherence traffic and protocol races. We present an implementation of TC, called TC-Weak, which eliminates LCC's trade-off between stalling stores and increasing L1 miss rates to improve performance and reduce interconnect traffic. By providing coherent L1 caches, TC-Weak improves the performance of GPU applications with inter-workgroup communication by 85% over disabling the non-coherent L1 caches in the baseline GPU. We also find that write-through protocols outperform a writeback protocol on a GPU as the latter suffers from increased traffic due to unnecessary refills of write-once data.
Inderpreet Singh, Arrvindh Shriraman, Wilson W. L. Fung, Mike O'Connor, Tor M. Aamodt
HPCA1
2011 Hardware transactional memory for GPU architectures
abstract
Graphics processor units (GPUs) are designed to efficiently exploit thread level parallelism (TLP), multiplexing execution of 1000s of concurrent threads on a relatively smaller set of single-instruction, multiple-thread (SIMT) cores to hide various long latency operations. While threads within a CUDA block/OpenCL workgroup can communicate efficiently through an intra-core scratchpad memory, threads in different blocks can only communicate via global memory accesses. Programmers wishing to exploit such communication have to consider data-races that may occur when multiple threads modify the same memory location. Recent GPUs provide a form of inter-block communication through atomic operations for single 32-bit/64-bit words. Although fine-grained locks can be constructed from these atomic operations, synchronization using locks is prone to deadlock. In this paper, we propose to solve these problems by extending GPUs to support transactional memory (TM). Major challenges include supporting 1000s of concurrent transactions and committing non-conflicting transactions in parallel. We propose KILO TM, a novel hardware TM design for GPUs that scales to 1000s of concurrent transactions. Without cache coherency hardware to depend on, it uses word-level, value-based conflict detection to avoid broadcast communication and reduce on-chip storage overhead. It employs speculative validation using a novel bloom filter organization to increase transaction commit parallelism. For a set of TM-enhanced GPU applications, KILO TM captures 59% of the performance of fine-grained locking, and is on average 128x faster than executing all transactions serially, for an estimated hardware area overhead of 0.5% of a commercial GPU.
Wilson W. L. Fung, Inderpreet Singh, Andrew Brownsword, Tor M. Aamodt
MICRO2
2009 Simulating Peer-to-Peer networks
abstract
The Gnutella protocol of peer-to-peer (P2P) networks has undergone several changes since its inception in the beginning of this century. However, despite the large number of revisions to the original version of the protocol, Gnutella suffers from serious problems of dead searches, complexity in study of network topology and network overloading. In this paper, we report the development of a new P2P simulator, PeerNS, which was built to study different problems of P2P networks and Gnutella, including those mentioned above. PeerNS works on actual P2P network statistics and, hence, it is very close to the real scenario. Moreover, we also discuss the implementation and the integration issues involved in using PeerNS to simulate our crawling-based algorithm, which could minimize the number of dead searches in the network and enhance the availability of information across the network.
Sanjay K. Dhurandher, Sudip Misra, Mohammad S. Obaidat, Inderpreet Singh, Raghu Agarwal, Bhuvnesh Bhambhani
AICCSA4
2009 Extracting the textual and temporal structure of supercomputing logs
abstract
Supercomputers are prone to frequent faults that adversely affect their performance, reliability and functionality. System logs collected on these systems are a valuable resource of information about their operational status and health. However, their massive size, complexity, and lack of standard format makes it difficult to automatically extract information that can be used to improve system management. In this work we propose a novel method to succinctly represent the contents of supercomputing logs, by using textual clustering to automatically find the syntactic structures of log messages. This information is used to automatically classify messages into semantic groups via an online clustering algorithm. Further, we describe a methodology for using the temporal proximity between groups of log messages to identify correlated events in the system. We apply our proposed methods to two large, publicly available supercomputing logs and show that our technique features nearly perfect accuracy for online log-classification and extracts meaningful structural and temporal message patterns that can be used to improve the accuracy of other log analysis techniques.
Sourabh Jain, Inderpreet Singh, Abhishek Chandra, Zhi-Li Zhang, Greg Bronevetsky
HiPC2
2009 On Increasing Information Availability in Gnutella-Like Peer-to-Peer Networks
abstract
In this paper, we address some of the problems such as dead searches, complexity in the study of network topology and network overloading that are associated with Gnutella and Gnutella-like peer-to-peer (P2P) networks. We use advanced heuristic parameters with information shuffling as a solution for them. We propose an advancement of Gnutella using the above-mentioned schemes. At a panoramic level, our work is founded on the following concepts: (a) Crawling the P2P networks to shuffle information, so that the knowledge is distributed over the whole network, and (b) Bringing the information within searchable hops of each network. These have been verified on a self-built P2P simulator, named PeerNS, which works on actual P2P network statistics and is, hence, very close to the actual scenario. The results obtained through simulation affirm that the nodes with extremely large number of dead searches benefit the most and are observed to have a sharp decrease in their dead search count after crawling a small part of the overall network.
Sudip Misra, Sanjay K. Dhurandher, Mohammad S. Obaidat, Inderpreet Singh, Bhuvnesh Bhambhani, Raghu Agarwal
ICC4