VLDB 2026 Research / reviewers in the wild / expert
Zhiheng Hu
dblp:332/0698
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy SelectionabstractTraining large language models faces frequent interruptions due to various faults, demanding robust fault-tolerance. Existing backup-free methods, such as redundant computation, dynamic parallelism, and data rerouting, each incur performance penalties, whether from ongoing overhead, lengthy reconfigurations, or post-recovery inefficiencies. We propose Chameleon, an adaptive fault-tolerant system that intelligently selects optimal recovery strategies when a failure occurs. Chameleon achieves this through a unified performance model, expedient execution plan search, accurate performance estimation, and efficient communication optimizations. Experiments on a 32-card cluster show that Chameleon maintains a performance gap of within 11.00% between post-recovery and failure-free training, while preserving model convergence and efficient memory usage. Compared to state-of-the-art methods, Chameleon achieves up to 1.229x and 1.355x higher average throughput than Oobleck and Recycle, respectively. Zhibin Wang 0002, Haoran Xia, Junhe Lu, Qianyu Jiang, Rong Gu 0001, Hengxi Xu, Xinjing Huang, Guanghuan Fang, Zhiheng Hu, Yongjin Cai, Chen Tian 0001 |
INFOCOM | 11 |
| 2026 | A Theoretical Framework on Real-Time Communication and Information Estimation in Ultra Large-Scale 6G C-V2X NetworksabstractThe emergence of sixth generation communication (6G) wireless networks is set to revolutionize vehicular communication by enabling ultra-reliable, low-latency, and high-capacity connectivity in cellular vehicle-to-everything (C-V2X) environments. This paper presents a theoretical framework on novel cooperative vehicular communication and information perception algorithms for large-scale 6G C-V2X networks while leveraging integrated space-air-ground communication system. Specifically, we address key challenges in real-time information exchange and fusion among multiple vehicles. Utilizing inequality theory and functional mapping theory, we derive an upper bound on channel capacity for a fixed number of relays and propose a low-complexity, multi-class relay selection algorithm. Furthermore, we introduce an optimal mobile edge computing (MEC) based correspondence strategy to improve vehicle-to-vehicle communication, alongside an efficient information estimation algorithm to facilitate real-time data sharing. Our simulation results confirm that the proposed algorithms significantly outperform existing cooperative vehicular schemes in terms of channel capacity, while the developed evaluation theory ensures accurate cooperative perception with reduced computational complexity. The proposed framework and theoretical contributions offer a foundational basis for 6G C-V2X networks. Zi Long Liu 0001, Haishi Wang, Wei Huang 0010, Chaojie Gu, Zhiheng Hu, Md. Noor-A-Rahim |
IEEE Trans. Wirel. Commun. | 6 |
| 2025 | Using Analytical Performance/Power Model and Fine-Grained DVFS to Enhance AI Accelerator Energy EfficiencyabstractRecent advancements in deep learning have significantly increased AI processors' energy consumption, which is becoming a critical factor limiting AI development. Dynamic Voltage and Frequency Scaling (DVFS) stands as a key method in power optimization. However, due to the latency of DVFS control in AI processors, previous works typically apply DVFS control at the granularity of a program's entire duration or sub-phases, rather than at the level of AI operators. Yijia Zhang 0002, Fuchun Wei, Bingqiang Wang, Yanlin Liu, Zhiheng Hu, Xiaoxin Xu, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Chen Tian 0001 |
ASPLOS (1) | 6 |
| 2025 | Squeezing Operator Performance Potential for the Ascend ArchitectureabstractWith the rise of deep learning, many companies have developed domain-specific architectures (DSAs) optimized for AI workloads, with Ascend being a representative. To fully realize the operator performance on Ascend, effective analysis and optimization is urgently needed. Compared to GPU, Ascend requires users to manage operations manually, leading to complex performance issues that require precise analysis. However, existing roofline models face challenges of visualization complexity and inaccurate performance assessment. To address these needs, we introduce a component-based roofline model that abstracts components to capture operator performance, thereby effectively identifying bottleneck components. Furthermore, through practical operator optimization case studies, we illustrate a comprehensive process of optimization based on roofline analysis, summarizing common performance issues and optimization strategies. Finally, extensive end-to-end optimization experiments demonstrate significant model speed improvements, ranging from 1.07× to 2.15×, along with valuable insights from practice. Zhibin Wang 0002, Guyue Liu, Yongzhong Wang, Fuchun Wei, Zhiheng Hu, Yanlin Liu, Yaoyuan Wang, Wan-Chun Dou, Guihai Chen, Chen Tian 0001 |
ASPLOS (2) | 10 |
| 2025 | Optimal Real-time Communication in 6G Ultra-Massive V2X Mobile NetworksabstractThis paper introduces a novel cooperative vehicular communication algorithm tailored for future 6G ultra-massive vehicle-to-everything (V2X) networks leveraging integrated space-air-ground communication systems. Specifically, we address the challenge of real-time information exchange among rapidly moving vehicles. We demonstrate the existence of an upper bound on channel capacity given a fixed number of relays, and propose a low-complexity relay selection heuristic algorithm. Simulation results verify that our proposed algorithm achieves superior channel capacities compared to existing cooperative vehicular communication approaches. Zi Long Liu 0001, Zeping Sui, Wei Huang 0010, Md. Noor-A-Rahim, Haishi Wang, Zhiheng Hu |
VTC2025-Fall | 7 |
| 2024 | BEVTemp: Enhancing Vision-Based Roadside 3D Object Detection with Temporal Information
Gaoyuan Miao, Rong Quan, Cong Pan 0001, Zhiheng Hu, Jie Qin 0004 |
PRICAI (3) | 4 |
| 2023 | ISmallNet: Densely Nested Network with Label Decoupling for Infrared Small Target DetectionabstractSmall targets are often submerged in cluttered backgrounds of infrared images. Conventional detectors tend to generate false alarms, while CNN-based detectors lose small targets in deep layers. To this end, we propose iSmallNet, a multi-stream densely nested network with label decoupling for infrared small object detection. On the one hand, to fully exploit the shape information of small targets, we decouple the original labeled ground-truth (GT) map into an interior map and a boundary one. The GT map, in collaboration with the two additional maps, tackles the unbalanced distribution of small object boundaries. On the other hand, two key modules are delicately designed and incorporated into the proposed network to boost the overall performance. First, to maintain small targets in deep layers, we develop a multi-scale nested interaction module to explore a wide range of context information. Second, we develop an interior-boundary fusion module to integrate multi-granularity information. Experiments on NUAA-SIRST and NUDT-SIRST clearly show the superiority of iSmallNet over 11 state-of-the-art detectors. Zhiheng Hu, Yongzhen Wang 0001, Peng Li 0064, Haoran Xie 0001, Mingqiang Wei |
ICASSP | 1 |