Jinghan Yao

dblp:239/5092 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0002-7129-9508ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 From Skew to Symmetry: Node-Interconnect Multi-Path Balancing with Execution-time Planning for Modern GPU Clusters
Jinghan Yao, Kaushik Kandadi Suresh, Bharath Ramesh 0005, Hari Subramoni, Dhabaleswar K. Panda 0001
IPDPS1
2025 Design and Optimization of GPU-Aware MPI Allreduce Using Direct Sendrecv Communication
abstract
Modern GPU-accelerated high-performance computing (HPC) and deep learning (DL) applications rely heavily on collective communication, particularly the Allreduce operation. As systems scale to hundreds or thousands of GPUs, conventional algorithms such as Ring or vendor libraries like NCCL struggle to sustain performance for large scale due to communication bottlenecks and algorithmic dependencies. In this work, we propose a novel GPU-aware Allreduce design using a Direct Sendrecv algorithm with throttling to improve scalability and bandwidth utilization across various scales and interconnects. To further reduce overhead, we introduce computation-communication overlap and kernel fusion techniques. The design also extends to CPU-staging scenarios for small messages. Evaluations on large-scale GPU systems demonstrate that our designs outperform baseline NCCL implementations by up to 40% at medium message sizes. In application-level evaluations, the proposed design achieves up to 7% improvement in nanoGPT training and 27% improvement in the Amber HPC simulation.
Chen-Chun Chen, Jinghan Yao, Hari Subramoni, Dhabaleswar K. Panda 0001
ICPP2
2025 Unified Designs of Multi-Rail-Aware MPI Allreduce and Alltoall Operations Across Diverse GPU and Interconnect Systems
abstract
The growing demand for computing power in highperformance computing is driving the adoption of diverse accelerators and interconnect networks in modern exascale clusters. Within a node, device interconnects like NVLink, Infinity Fabric, and$\mathbf{X}^{e}$Link, along with IPC techniques, provide high throughput in dense GPU environments. Recently, multi-rail interconnects, such as InfiniBand, Slingshot, and Omni-Path, have enabled highbandwidth communication between nodes. Additionally, highperformance computing applications impose significant demands on collective operations such as Allreduce and Alltoall. Therefore, designing efficient and scalable MPI runtimes for diverse system architectures at large scales is essential. In this paper, we propose unified designs to optimize MPI Allreduce and Alltoall operations using multi-rail-aware, two-level algorithms. The designs support a variety of GPU and interconnect combinations, including NVIDIA, AMD, and Intel GPUs, across InfiniBand, Slingshot, and Omni-Path networks, in the modern dense GPU systems. We optimized the Allreduce operation using a persistent device buffer for device-side reduction and employed an early-triggered, pipelined approach to overlap computation with communication. Additionally, we leveraged the device buffer with IPC techniques as a shared buffer to enhance the two-level Alltoall algorithm, designing both PUSH and PULL variants. We evaluate the advantages of our designs through benchmark and applicationlevel tests on the IsambardAI, Frontier, Cardinal, and Stampede3 systems. In benchmark evaluations, the proposed Allreduce design shows a$2.8 \mathrm{x}, 2.9 \mathrm{x}$, and 2.9 x performance improvement at 1 GB with 32 NVIDIA, 64 AMD, and 64 Intel GPUs, respectively. Additionally, the proposed Alltoall design demonstrates a 1.3 x and 1.05 x improvement at 4 MB with 32 NVIDIA and 64 AMD GPUs, respectively. In application-level evaluations, the proposed Allreduce design demonstrates a$2 x$performance improvement in Amber, while the Alltoall design shows a 1.4x performance gain in heFFTe, both tested on 32 H100 and GH200 GPUs with Infiniband and Slingshot-11 interconnects, respectively.
Chen-Chun Chen, Jinghan Yao, Hari Subramoni, Dhabaleswar K. Panda 0001
IPDPS2
2025 A High-Density Transcranial Electrical Stimulation System on Chip with Real-Time Bio-Impedance Sensing
abstract
This paper presents a system on chip (SoC) designed for high-density transcranial electrical stimulation (HD-TES) with real-time bio-impedance sensing. Multiple chips can work collaboratively to generate arbitrary waveforms for HD-TES including temporal interference (TI) stimulation. This design implements a combined digital calibration and analog switch mechanism for common-mode voltage holding (CMVH), ensuring long-time safety of HD-TES. The SoC can perform impedance measurement (IM) through transcranial alternating current stimulation (tACs) allowing to monitor the biological impedance variations during tACs. To ensure the safety of current stimulation, switched-capacitor overcurrent protection (OCP) circuit is integrated. Measured results show that the SoC can apply arbitrary stimulation currents within the voltage range of ±13V. The amplitude range of the current is from - 2.5mA to 2.5mA with a 10 μA minimum step. The common-mode voltage of the electrode remains stable near 0V when operating in HD-TES mode. Impedance measurement error is less than 5% within the extensive range of 0.5kΩ to 500kΩ.
Shaokai Yuan, Jinghan Yao, Jianzheng Li, Yajie Qin
ISCAS2
2024 HyperSack: Distributed Hyperparameter Optimization for Deep Learning using Resource-Aware Scheduling on Heterogeneous GPU Systems
abstract
Hyperparameter Optimization (HPO) can unlock the full potential of Deep Learning (DL) models; however, it is considered one of the most compute-intensive tasks in the DL do-main due to multi-dimensional search spaces and complex neural network architectures. A common method for accelerating HPO workloads is parallelizing training jobs on multiple computing devices, such as modern GPUs in High-Performance Computing (HPC) environments. Nonetheless, existing HPO parallelization strategies underutilize powerful GPU devices, like the NVIDIA AI00 and HI00, especially for training lightweight Deep Neu-ral Networks. Resource-sharing mechanisms can improve G PU utilization; nonetheless, naive adaptations for HPO workloads lead to poor performance. Therefore, we propose HyperSack-a distributed HPO framework for dynamic and resource-aware scheduling on heterogeneous GPU-based HPC systems with resource elasticity and fault tolerance. HyperSack reduces the execution time of HPO workloads by orchestrating the placement of DL training jobs on GPU devices with different computational capabilities. It supports different hardware architectures, HPO workloads, and scheduling policies. Our evaluations on vision and language models HPO workloads show up to 2.8x performance improvement in execution time on AI00 GPUs, 4.0x on HI00 GPUs, and 3.9x on a combination of 12 AI00 and 4 HI00 GPUs using HyperSack over standard HPO parallelization methods.
Nawras Alnaasan, Bharath Ramesh 0005, Jinghan Yao, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda 0001
HiPC3
2024 Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
abstract
In the realm of large language models (LLMs) like the Generative Pre-trained Transformer (GPT), the Mixture of Experts (MoE) paradigm has emerged as a powerful technique for enhancing model expressiveness and accuracy. However, the deployment of GPT MoE models for parallel inference on distributed systems presents significant challenges, primarily due to the extensive Alltoall communication required for expert routing and aggregation. This communication bottleneck exacerbates the already complex computational landscape, hindering the efficient utilization of high-performance computing resources. In this paper, we propose a lightweight optimization technique called ExFlow, to largely accelerate the inference of these MoE models. We take a new perspective on alleviating the communication overhead by exploiting the inter-layer expert affinity. Unlike previous methods, our solution can be directly applied to pre-trained MoE models without any fine-tuning or accuracy degradation. By proposing a context-coherent expert parallelism on distributed systems, our ExFlow design only uses one Alltoall communication to deliver the same functionality while previous methods all require two Alltoalls. By carefully examining the conditional probability in tokens’ routing across multiple layers, we proved that pre-trained GPT MoE models implicitly exhibit a strong inter-layer expert affinity. We then design an efficient integer programming model to precisely capture such features and show that by properly placing the experts on corresponding GPUs, we can reduce up to 67% of tokens’ cross-GPU routing latency on various hardware configurations and topologies. Our solution beats the cutting-edge Deepspeed-MoE in GPT MoE models with experts from 8 to 64, with up to 2.2x improvement in inference throughput. To the best of our knowledge, this is the first work in leveraging inter-layer expert affinity to accelerate the inference of GPT MoE models. We further provide a detailed study of how the model implicitly acquires this expert affinity at the very early training stage and how this affinity evolves and stabilizes during training.
Jinghan Yao, Quentin Anthony, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda 0001
IPDPS1
2023 Flover: A Temporal Fusion Framework for Efficient Autoregressive Model Parallel Inference
abstract
Autoregressive models, despite their commendable performance in a myriad of generative tasks, face challenges stemming from their inherently sequential structure. Inference on these models, by design, harnesses a temporal dependency, where the current token's probability distribution is conditioned on preceding tokens. This inherent characteristic severely impedes computational efficiency during inference as a typical inference request can require more than thousands of tokens, where generating each token requires a load of entire model weights, making the inference more memory-bound. The large overhead becomes profound in real deployment where requests arrive randomly, necessitating various generation lengths. Existing solutions, such as dynamic batching and concurrent instances, introduce significant response delays and bandwidth contention, falling short of achieving optimal latency and throughput. To address these shortcomings, we propose Flover - a temporal fusion framework for efficiently inferring multiple requests in parallel. We deconstruct the general generation pipeline into pre-processing and token generation, and equip the framework with a dedicated work scheduler for fusing the generation process temporally across all requests. By orchestrating the token-level parallelism, Flover exhibits optimal hardware efficiency and significantly spares the system resources. By further employing a fast buffer reordering algorithm that allows memory eviction of finished tasks, it brings over 11× inference speedup on GPT and 16 × on LLAMA compared to the cutting-edge solutions provided by NVIDIA FasterTransformer. Crucially, by leveraging the advanced tensor parallel technique, Flover proves efficacious across diverse computational landscapes, from single-GPU setups to distributed scenarios, thereby offering robust performance optimization that adapts to variable use cases.
Jinghan Yao, Nawras Alnaasan, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda 0001
HiPC1
2023 A Novel Framework for Efficient Offloading of Communication Operations to Bluefield SmartNICs
abstract
Smart Network Interface Cards (SmartNICs) such as NVIDIA’s BlueField Data Processing Units (DPUs) provide advanced networking capabilities and processor cores, enabling the offload of complex operations away from the host. In the context of MPI, prior work has explored the use of DPUs to offload non-blocking collective operations. The limitations of current state-of-the-art approaches are twofold: They only work for a pre-defined set of algorithms/communication patterns and have degraded communication latency due to staging data between the DPU and the host. In this paper, we propose a framework that supports the offload of any communication pattern to the DPU while achieving low communication latency with perfect overlap. To achieve this, we first study the limitations of higher-level programming models such as MPI in expressing the offload of complex communication patterns to the DPU. We present a new set of APIs to alleviate these shortcomings and support any generic communication pattern. Then, we analyze the bottlenecks involved in offloading communication operations to the DPU and propose efficient designs for a few candidate communication patterns. To the best of our knowledge, this is the first framework providing both efficient and generic communication offload to the DPU. Our proposed framework outperforms state-of-the-art staging-based offload solutions by 47% in Alltoall micro-benchmarks, and at the application level, we see improvements up to 60% in P3DFFT and 15% in HPL on 512 processes.
Kaushik Kandadi Suresh, Benjamin Michalowicz, Bharath Ramesh 0005, Nicholas Contini, Jinghan Yao, Shulei Xu, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda 0001
IPDPS5
2021 SOFT: Softmax-free Transformer with Linear Complexity
abstract
Vision transformers (ViTs) have pushed the state-of-the-art for various visual recognition tasks by patch-wise image tokenization followed by self-attention. However, the employment of self-attention modules results in a quadratic complexity in both computation and memory usage. Various attempts on approximating the self-attention computation with linear complexity have been made in Natural Language Processing. However, an in-depth analysis in this work shows that they are either theoretically flawed or empirically ineffective for visual recognition. We further identify that their limitations are rooted in keeping the softmax self-attention during approximations. Specifically, conventional self-attention is computed by normalizing the scaled dot-product between token feature vectors. Keeping this softmax operation challenges any subsequent linearization efforts. Based on this insight, for the first time, a softmax-free transformer or SOFT is proposed. To remove softmax in self-attention, Gaussian kernel function is used to replace the dot-product similarity without further normalization. This enables a full self-attention matrix to be approximated via a low-rank matrix decomposition. The robustness of the approximation is achieved by calculating its Moore-Penrose inverse using a Newton-Raphson method. Extensive experiments on ImageNet show that our SOFT significantly improves the computational efficiency of existing ViT variants. Crucially, with a linear complexity, much longer token sequences are permitted in SOFT, resulting in superior trade-off between accuracy and complexity.
Jinghan Yao, Junge Zhang, Xiatian Zhu, Hang Xu 0004, Weiguo Gao, Chunjing Xu, Tao Xiang 0002, Li Zhang 0040
NeurIPS2
2021 SPRNet: Single-Pixel Reconstruction for One-Stage Instance Segmentation
abstract
Object instance segmentation is one of the most fundamental but challenging tasks in computer vision, and it requires the pixel-level image understanding. Most existing approaches address this problem by adding a mask prediction branch to a two-stage object detector with the region proposal network (RPN). Although producing good segmentation results, the efficiency of these two-stage approaches is far from satisfactory, restricting their applicability in practice. In this article, we propose a one-stage framework, single-pixel reconstruction net (SPRNet), which performs efficient instance segmentation by introducing a single-pixel reconstruction (SPR) branch to off-the-shelf one-stage detectors. The added SPR branch reconstructs the pixel-level mask from every single pixel in the convolution feature map directly. Using the same ResNet-50 backbone, SPRNet achieves comparable mask AP with Mask R-CNN at a higher inference speed and gains all-round improvements on box AP at every scale compared with RetinaNet.
Jun Yu 0002, Jinghan Yao, Jian Zhang 0026, Zhou Yu 0001, Dacheng Tao
IEEE Trans. Cybern.2