VLDB 2026 Research / reviewers in the wild / expert
Vishwesh Jatala
dblp:160/8460
· DBLP profile ↗
10ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0002-3105-922XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Cluster-Based Sampling Scheme for High-Performance Graph Neural NetworksabstractGraph Neural Networks (GNNs) have become a powerful toolbox for solving complex problems dealing with graph-structured data. Research has gained significant momentum to improve the accuracy of GNN models. However, GNN training is computationally expensive. Several popular GNN methods use random-based neighborhood sampling techniques to reduce the GNN training time. These techniques aim to sparsify the graph by sampling the neighborhood vertices randomly, but they do not consider the graph structure and associated neighborhood features while sampling. We propose a cluster-based sampling technique (CLING) that exploits the neighborhood structure and node features to address these challenges. To this, we first propose a weighted partitioning scheme that aims to partition the graph’s vertices by considering the similarity of vertex features and their neighborhood structure. Then, we sample the neighborhood vertices based on the identified clusters and the common neighborhood vertices. We implemented the proposed sampling scheme in CUDA using the DGL framework. We evaluated our approach on large-scale datasets and observed that GNN, when used with our sampling scheme, shows up to $2.3 \times$ speedup and an average of $1.7 \times$ speedup when compared to state-of-the-art implementations. Surendra Kumar Raut, Kishan Tamboli, Vishwesh Jatala |
ISPASS | 3 |
| 2026 | LSTC: Large-Scale Triangle Counting on Single GPUabstractTriangle counting in graphs has applications in a wide range of domains. For many years, researchers have been improving triangle counting performance by exploiting recent architectures, such as multicores and accelerators like GPUs. Improving the performance of triangle counting in GPU has several challenges: (1) GPUs are equipped very low memory, hence real-word large graphs can not be processed using the capacity of single GPU memory, (2) GPUs follow SIMD execution, whereas graphs exhibit irregular data parallelism, and (3) triangle counting incurs huge memory accesses, minimizing the number of global memory accesses is crucial for good performance. Kishan Tamboli, Vishwesh Jatala |
ICPE | 2 |
| 2025 | Multivariate Time Series Data Mining for Failure Prediction & Root Cause AnalysisabstractThis paper presents an analysis of a large-scale, high-dimensional industrial dataset containing over 2 million data points collected over several months. The dataset includes more than 200 failures of various types, each resulting from complex causes. Utilizing state-of-the-art unsupervised multivariate anomaly detection algorithms, a system was developed that predicts failures several minutes in advance, introducing a new evaluation metric termed lead time. This system also identifies the location and potential causes of these failures. Initial application of current anomaly detection algorithms yielded low F1 scores, due to either low recall or low precision. To address this issue, a visual reasoning framework was created to reduce false positives by analyzing motifs and discords in the top-N anomalous signals contributing to the anomaly. We find that the human-in-the-loop approach enhances the precision and F1-score of multivariate algorithms by leveraging human judgment for final decision-making. Another key finding is that the LSTM-based anomaly detection algorithm achieves sufficient lead time for the industrial failure prediction studied in this paper. Naman Agarwal, Vishwesh Jatala, Aniket Saha |
ICASSP | 3 |
| 2023 | Entropy Aware Training for Fast and Accurate Distributed GNNabstractSeveral distributed frameworks have been developed to scale Graph Neural Networks (GNNs) on billion-size graphs. On several benchmarks, we observe that the graph partitions generated by these frameworks have heterogeneous data distributions and class imbalance, affecting convergence, and resulting in lower performance than centralized implementations. We holistically address these challenges and develop techniques that reduce training time and improve accuracy. We develop an Edge-Weighted partitioning technique to improve the micro average F1 score (accuracy) by minimizing the total entropy. Furthermore, we add an asynchronous personalization phase that adapts each compute-host’s model to its local data distribution. We design a class-balanced sampler that considerably speeds up convergence. We implemented our algorithms on the DistDGL framework and observed that our training techniques scale much better than the existing training approach. We achieved a (2-3x) speedup in training time and 4% improvement on average in micro-F1 scores on 5 large graph benchmarks compared to the standard baselines. Dhruv Deshmukh, Gagan Raj Gupta 0001, Manisha Chawla, Vishwesh Jatala, Anirban Haldar |
ICDM | 4 |
| 2022 | Joint Partitioning and Sampling Algorithm for Scaling Graph Neural NetworkabstractGraph Neural Network (GNN) has emerged as a popular toolbox for solving complex problems on graph data structures. Graph neural networks use machine learning techniques to learn the vector representations of nodes and/or edges. Learning these representations demands a huge amount of memory and computing power. The traditional shared-memory multiprocessors are insufficient to meet real-world data’s computing requirements; hence, research has gained momentum toward distributed GNN.Scaling the distributed GNN has the following challenges: (1) the input graph needs to be efficiently partitioned, (2) the cost of communication between compute nodes should be reduced, and (3) the sampling strategy should be efficiently chosen to minimize the loss in accuracy. To address these challenges, we propose a joint partitioning and sampling algorithm, which partitions the input graph with weighted METIS and uses a bias sampling strategy to minimize total communication costs.We implemented our approach using the DistDGL framework and evaluated it using several real-world datasets. We observe that our approach (1) shows an average reduction in communication overhead by 53%, (2) requires less partitioning time to partition a graph, (3) shows improved accuracy, (4) shows a speed up of 1.5x on OGB-Arxiv dataset, when compared to the state-of-the-art DistDGL implementation. Manohar Lal Das, Vishwesh Jatala |
HIPC | 2 |
| 2020 | A Study of Graph Analytics for Massive Datasets on Distributed Multi-GPUsabstractThere are relatively few studies of distributed GPU graph analytics systems in the literature and they are limited in scope since they deal with small data-sets, consider only a few applications, and do not consider the interplay between partitioning policies and optimizations for computation and communication.In this paper, we present the first detailed analysis of graph analytics applications for massive real-world datasets on a distributed multi-GPU platform and the first analysis of strong scaling of smaller real-world datasets. We use D-IrGL, the state-of-the-art distributed GPU graph analytical framework, in our study. Our evaluation shows that (1) the Cartesian vertex-cut partitioning policy is critical to scale computation out on GPUs even at a small scale, (2) static load imbalance is a key factor in performance since memory is limited on GPUs, (3) device-host communication is a significant portion of execution time and should be optimized to gain performance, and (4) asynchronous execution is not always better than bulk-synchronous execution. Vishwesh Jatala, Roshan Dathathri, Gurbinder Gill, Loc Hoang, V. Krishna Nandivada, Keshav Pingali |
IPDPS | 1 |
| 2019 | Gluon-Async: A Bulk-Asynchronous System for Distributed and Heterogeneous Graph AnalyticsabstractDistributed graph analytics systems for CPUs, like D-Galois and Gemini, and for GPUs, like D-IrGL and Lux, use a bulk-synchronous parallel (BSP) programming and execution model. BSP permits bulk-communication and uses large messages which are supported efficiently by current message transport layers, but bulk-synchronization can exacerbate the performance impact of load imbalance because a round cannot be completed until every host has completed that round. Asynchronous distributed graph analytics systems circumvent this problem by permitting hosts to make progress at their own pace, but existing systems either use global locks and send small messages or send large messages but do not support general partitioning policies such as vertex-cuts. Consequently, they perform substantially worse than bulk-synchronous systems. Moreover, none of their programming or execution models can be easily adapted for heterogeneous devices like GPUs. In this paper, we design and implement a lock-free, non-blocking, bulk-asynchronous runtime called Gluon-Async for distributed and heterogeneous graph analytics. The runtime supports any partitioning policy and uses bulk-communication. We present the bulk-asynchronous parallel (BASP) model which allows the programmer to utilize the runtime by specifying only the abstract communication required. Applications written in this model are compared with the BSP programs written using (1) D-Galois and D-IrGL, the state-of-the-art distributed graph analytics systems (which are bulk-synchronous) for CPUs and GPUs, respectively, and (2) Lux, another (bulk-synchronous) distributed GPU graph analytical system. Our evaluation shows that programs written using BASP-style execution are on average ~1.5x faster than those in D-Galois and D-IrGL on real-world large-diameter graphs at scale. They are also on average ~12x faster than Lux. To the best of our knowledge, Gluon-Async is the first asynchronous distributed GPU graph analytics system. Roshan Dathathri, Gurbinder Gill, Loc Hoang, Vishwesh Jatala, Keshav Pingali, V. Krishna Nandivada, Hoang-Vu Dang, Marc Snir |
PACT | 4 |
| 2018 | Reducing GPU Register File Energy
Vishwesh Jatala, Jayvant Anantpur, Amey Karkare |
Euro-Par | 1 |
| 2017 | Scratchpad Sharing in GPUsabstractGeneral-Purpose Graphics Processing Unit (GPGPU) applications exploit on-chip scratchpad memory available in the Graphics Processing Units (GPUs) to improve performance. The amount of thread level parallelism (TLP) present in the GPU is limited by the number of resident threads, which in turn depends on the availability of scratchpad memory in its streaming multiprocessor (SM). Since the scratchpad memory is allocated at thread block granularity, part of the memory may remain unutilized. In this article, we propose architectural and compiler optimizations to improve the scratchpad memory utilization. Our approach, called Scratchpad Sharing , addresses scratchpad under-utilization by launching additional thread blocks in each SM. These thread blocks use unutilized scratchpad memory and also share scratchpad memory with other resident blocks. To improve the performance of scratchpad sharing, we propose Owner Warp First (OWF) scheduling that schedules warps from the additional thread blocks effectively. The performance of this approach, however, is limited by the availability of the part of scratchpad memory that is shared among thread blocks. We propose compiler optimizations to improve the availability of shared scratchpad memory. We describe an allocation scheme that helps in allocating scratchpad variables such that shared scratchpad is accessed for short duration. We introduce a new hardware instruction, relssp , that when executed releases the shared scratchpad memory. Finally, we describe an analysis for optimal placement of relssp instructions, such that shared scratchpad memory is released as early as possible, but only after its last use, along every execution path. We implemented the hardware changes required for scratchpad sharing and the relssp instruction using the GPGPU-Sim simulator and implemented the compiler optimizations in Ocelot framework. We evaluated the effectiveness of our approach on 19 kernels from 3 benchmarks suites: CUDA-SDK, GPGPU-Sim, and Rodinia. The kernels that under-utilize scratchpad memory show an average improvement of 19% and maximum improvement of 92.17% in terms of the number of instruction executed per cycle when compared to the baseline approach, without affecting the performance of the kernels that are not limited by scratchpad memory. Vishwesh Jatala, Jayvant Anantpur, Amey Karkare |
ACM Trans. Archit. Code Optim. | 1 |
| 2016 | Improving GPU Performance Through Resource SharingabstractGraphics Processing Units (GPUs) consisting of Streaming Multiprocessors (SMs) achieve high throughput by running a large number of threads and context switching among them to hide execution latencies. The number of thread blocks, and hence the number of threads that can be launched on an SM, depends on the resource usage--e.g. number of registers, amount of shared memory--of the thread blocks. Since the allocation of threads to an SM is at the thread block granularity, some of the resources may not be used up completely and hence will be wasted. Vishwesh Jatala, Jayvant Anantpur, Amey Karkare |
HPDC | 1 |