Yash Malhotra

dblp:440/7760 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2026
0009-0005-9482-6341ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 50% GPUs and heterogeneous computing · 25% Electronic design automation · 25%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › fast fourier transform
distributed FFT
1.012026
torch_cufft: Extending PyTorch FFT Capacity with Multi-GPU and Descriptor-Resident cuFFTXt Execution · HPDC 2026
GPUs and heterogeneous computing
multi-GPU computing
1.012026
torch_cufft: Extending PyTorch FFT Capacity with Multi-GPU and Descriptor-Resident cuFFTXt Execution · HPDC 2026
High-performance computing
scientific computing
1.012026
torch_cufft: Extending PyTorch FFT Capacity with Multi-GPU and Descriptor-Resident cuFFTXt Execution · HPDC 2026
Electronic design automation
spectral methods
1.012026
torch_cufft: Extending PyTorch FFT Capacity with Multi-GPU and Descriptor-Resident cuFFTXt Execution · HPDC 2026
Machine learning › Deep learning architectures and training › neural operator
fourier neural operator
0.312026
torch_cufft: Extending PyTorch FFT Capacity with Multi-GPU and Descriptor-Resident cuFFTXt Execution · HPDC 2026

Methods — techniques the papers use, named apart from their topics

descriptor-resident execution · 2.0cuFFT · 2.0
YearPublicationVenuePosition
2026 torch_cufft: Extending PyTorch FFT Capacity with Multi-GPU and Descriptor-Resident cuFFTXt Execution
abstract
Large scientific images and spectral-learning workloads often require two-dimensional Fast Fourier Transforms (2D FFTs) that exceed single-GPU memory. This matters for Fourier Neural Operators (FNOs) and Transform Once (T1)-style models, where frequency-domain computation is central to the learning workflow. PyTorch provides convenient FFT APIs, but scaling these transforms across multiple GPUs requires lower-level libraries and careful memory-layout management.
Yash Malhotra, Sanmukh R. Kuppannagari
HPDC1