EDBT 2026 Demo / reviewers in the wild / expert
Zhihang Tang
dblp:161/7257
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0003-4643-5483ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | µShare: Non-Intrusive Kernel Co-Locating on NVIDIA GPUsabstractThe hardware scheduler on NVIDIA GPUs is highly inefficient in utilizing micro-architectural hardware resources. It places blocks from the same kernel within the same GPU Streaming Multiprocessor (SM) core, resulting in a stacking colocating problem, where identical blocks are placed within the same SM core, saturating only a subset of intra-SM hardware resources while leaving others underutilized. The primary challenge in addressing this issue is that the NVIDIA hardware is closed-source, preventing us from directly modifying the hardware scheduler. To bridge the semantic gap between the resource demands of kernels and the scheduler, we introduce µ Share, which enables intra-SM scattered colocating of kernels through a non-intrusive half-plus blocksize shaping method. It shapes the blocksize of kernels to a halfplus blocksize (i.e., slightly more than half of the SM's thread capacity), scattering identical blocks of the same kernel across different SMs. It further adopts a time-shifted launching method to reduce intra-SM resource contention. Compared to state-of-the-art systems, µ Share does not require intrusive modifications to hardware or kernel code, yet it can still improve inference throughput by 26.90%-54.09% and increases low-level hardware utilization by 38.53%-61.15%. Wenhao Huang 0005, Zhaolin Duan, Laiping Zhao, Yuhao Zhang 0006, Yichi Chen 0001, Zhihang Tang, Kang Chen 0001, Deze Zeng, Wenxin Li 0001, Keqiu Li |
HPCA | 9 |
| 2026 | HiAsCC: Hierarchical Asynchronous Collective Communication Method for Large Model Training
Zhihang Tang, Bo He 0003, Qi Qi 0001, Yulong Tao, Jingyu Wang 0001, Laiping Zhao, Keqiu Li |
ICDCS | 1 |
| 2025 | HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language ModelsabstractSongtao Jiang, Yan Zhang, Yeying Jin, Zhihang Tang, Yangyang Wu, Yang Feng, Jian Wu, Zuozhu Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Songtao Jiang, Yan Zhang 0004, Yeying Jin, Zhihang Tang, Yang Feng 0011, Jian Wu 0001, Zuozhu Liu |
ACL (1) | 4 |
| 2025 | Shuffle-Exchange: Enhancing Collective Communication Efficiency for Large Model TrainingabstractTraining large models in parallel by GPU clusters significantly accelerates the computation in each iteration. However, the frequent collective communication for synchronizing the huge number of gradients poses a scalability challenge, whose performance gradually becomes the bottleneck as the number of workers increases. Ring-reduce is a favorable architecture since it can balance the communication and computation load among workers. In this paper, we discover that the communication resources are underutilized when using the ring-reduce synchronization method in clusters. Accordingly, Shuffle-Exchange Synchronization (SES), a novel method is proposed to improve the communication efficiency for distributed large model training. SES organizes all the worker nodes into several groups, within which they perform small-scale ring-reduce synchronizations during each iteration. To achieve better convergence performance, a gradient correction operation is integrated into SES. Experiments in 16 workers on a real-world industrial computing platform, show that SES can accelerate the large model training to 1.97× without losing model performance. Zhihang Tang, Bo He 0003, Qi Qi 0001, Jingyu Wang 0001, Laiping Zhao |
ICDCS | 1 |
| 2025 | Efficient Scheduling for Multiple Distributed DNN Training Tasks in Resource-Constrained Edge NetworksabstractThe increasing parameter size of Deep Neural Networks (DNNs) has significantly enhanced model performance. As large-scale DNN models typically require partitioning into multiple blocks for distributed training, existing research has predominantly focused on offline scheduling for individual or batched training tasks. However, the stochastic arrival of such tasks in edge networks poses a critical challenge for efficiently scheduling them in resource-constrained edge clusters. In this paper, aiming to minimize the average training completion time across all tasks, we first extract DNN operator graphs and partition them into coarse-grained subgraphs using a max-flow mincut algorithm. Then, we formulate the online scheduling problem for multiple distributed DNN training tasks as a Markov Decision Process (MDP) and propose a reinforcement learning-based (RL-based) solution. Extensive experiments comparing our method with three conventional baselines (FIFO, SJF, and Greedy) under diverse configurations show that our approach reduces the average training completion time by 12.95%, demonstrating its effectiveness in resource-constrained edge environments with dynamic workloads. Zhihang Tang, Weiqi Yue, Baofu Wu, Binbin Huang 0006, Laiping Zhao, Keqiu Li |
ICPADS | 1 |
| 2025 | V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis
Shujian Gao, Zhihang Tang, Xiaotang Gai, Jian Wu 0001, Zuozhu Liu |
MICCAI (5) | 5 |
| 2024 | RFaaS: Function Scheduling Across Heterogeneous Clusters
Zhihang Tang, Zezheng Mao, Laiping Zhao, Keqiu Li |
NPC (1) | 1 |