Huitian Wang

dblp:257/5881 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Balanced Sparse Tree: A Scalable Network Topology for Large Language Models
abstract
The development of large language models (LLMs) has catalyzed unprecedented demand on the computing network, specifically for large-scale, few-hops, and low-latency, which directly underpin LLM task efficiency. However, mainstream topologies such as Clos suffer from costs and latency, while topologies with good scalability have symmetric or collective communication issues. In order to achieve a favorable balance among design metrics, we propose a novel topology named the Balanced Sparse Tree (BST), which is a topology characterized by symmetric design and sparse connections, motivated by hypergraph theory and Steiner Systems. Its degree-diameter upper-bound approaches the Moore Bound for Bipartite Biregular graphs, larger than other known dia-meter-2 topologies. Furthermore, we incorporate differentiated routing, deadlock freedom, and topology-affined deployment into BST. Testbed experiments, simulations, together with modeling analysis, demonstrate the superiority of BST over the state-of-the-art in network scale, latency, bandwidth, and cost. With equivalent scales, BST outperforms Clos with a 50% cost reduction while maintaining comparable performance for AI workloads. Furthermore, BST delivers a 3.9%–11.8% gain in collective communications and has 13.4% improvement over state-of-the-art topologies.
Shaoteng Liu, Dejun Kong 0001, Huitian Wang, Hongji Dong, Fuguang Huang, Xiaofeng Gao 0001, Bingyang Liu, Guihai Chen
SIGCOMM3
2023 Multi-Exit DNN Inference Acceleration Based on Multi-Dimensional Optimization for Edge Intelligence
abstract
Edge intelligence, as a prospective paradigm for accelerating DNN inference, is mostly implemented by model partitioning which inevitably incurs the large transmission overhead of DNN's intermediate data. A popular solution introduces multi-exit DNNs to reduce latency by enabling early exits. However, existing work ignores the correlation between exit settings and synergistic inference, causing incoordination of device-to-edge. To address this issue, this paper first investigates the bottlenecks of executing multi-exit DNNs in edge computing and builds a novel model for inference acceleration with exit selection, model partition, and resource allocation. To tackle the intractable coupling subproblems, we propose a Multi-exit DNN inference Acceleration framework based on Multi-dimensional Optimization (MAMO). In MAMO, the exit selection subproblem is first extracted from the original problem. Then, bidirectional dynamic programming is employed to determine the optimal exit setting for an arbitrary multi-exit DNN. Finally, based on the optimal exit setting, a DRL-based policy is developed to learn joint decisions of model partition and resource allocation. We deploy MAMO on a real-world testbed and evaluate its performance in various scenarios. Extensive experiments show that it can adapt to heterogeneous tasks and dynamic networks, and accelerate DNN inference by up to 13.7x compared with the state-of-the-art.
Fang Dong 0001, Huitian Wang, Dian Shen, Zhaowu Huang, Qiang He 0001, Jinghui Zhang 0001, Liangsheng Wen
IEEE Trans. Mob. Comput.2
2022 Enabling Latency-Sensitive DNN Inference via Joint Optimization of Model Surgery and Resource Allocation in Heterogeneous Edge
abstract
Nowadays, edge computing is widely adopted to resolve the emerging deep neural networks (DNNs)-driven intelligence scenarios with the requirement of low-latency and high-accuracy, which includes heterogeneous end devices and DNNs. In such scenarios, the influx of data and computation into a shared edge server incurs prohibitive latency. Thus, we exploit the advantage of Multi-exit DNNs (ME-DNNs) that tasks can exit early at appropriate depths to save inference time. However, naively using ME-DNNs in the heterogeneous edge still fails to deliver fast inference due to improper model surgery and resource allocation.
Zhaowu Huang, Fang Dong 0001, Dian Shen, Huitian Wang, Xiaolin Guo, Shucun Fu
ICPP4
2021 Enabling Low Latency Edge Intelligence based on Multi-exit DNNs in the Wild
abstract
In recent years, deep neural networks (DNNs) have witnessed a booming of artificial intelligence Internet of Things applications with stringent demands across high accuracy and low latency. A widely adopted solution is to process such computation-intensive DNNs inference tasks with edge computing. Nevertheless, existing edge-based DNN processing methods still cannot achieve acceptable performance due to the intensive transmission data and unnecessary computation. To address the above limitations, we take the advantage of Multi-exit DNNs (ME-DNNs) that allows the tasks to exit early at different depths of the DNN during inference, based on the input complexity. However, naively deploying ME-DNNs in edge still fails to deliver fast and consistent inference in the wild environment. Specifically, 1) at the model-level, unsuitable exit settings will increase additional computational overhead and will lead to excessive queuing delay; 2) at the computation-level, it is hard to sustain high performance consistently in the dynamic edge computing environment. In this paper, we present a Low Latency Edge Intelligence Scheme based on Multi-Exit DNNs (LEIME) to tackle the aforementioned problem. At the model-level, we propose an exit setting algorithm to automatically build optimal ME-DNNs with lower time complexity; At the computation-level, we present a distributed offloading mechanism to fine-tune the task dispatching at runtime to sustain high performance in the dynamic environment, which has the property of close-to-optimal performance guarantee. Finally, we implement a prototype system and extensively evaluate it through testbed and large-scale simulation experiments. Experimental results demonstrate that LEIME significantly improves applications' performance, achieving 1.1–18.7 × speedup in different situations.
Zhaowu Huang, Fang Dong 0001, Dian Shen, Junxue Zhang 0001, Huitian Wang, Guangxing Cai, Qiang He 0001
ICDCS5
2021 Dynamic Path Based DNN Synergistic Inference Acceleration in Edge Computing Environment
abstract
Deep Neural Networks (DNNs) have achieved excellent performance in intelligent applications. Nevertheless, it is elusive for devices with limited resources to support computationally intensive DNNs, while employing the cloud may lead to prohibitive latency. Better solutions are exploiting edge computing and reducing unnecessary computation. Multi-exit DNN based on the early exit mechanism has an impressive effect in the latter, and in edge computing paradigm, model partition on multi-exit chain DNNs is proved to accelerate inference effectively. However, despite reducing computations to some extent, multiple exits may lead to instability of performance due to variable sample quality, performance inferior to the original model especially in the worst case. Furthermore, nowadays DNNs are universally characterized by a directed acyclic graph (DAG), complicating the partition of multi-exit DNN exceedingly. To solve the issues, in this paper, considering online exit prediction and model execution optimization for multi-exit DNN, we propose a Dynamic Path based DNN Synergistic inference acceleration framework (DPDS), where exit designators are designed to avoid iterative entry for exits; to further promote computational synergy in the edge, the multi-exit DNN is dynamically partitioned according to network environment to achieve fine-grained computing offloading. Experimental results show that DPDS can significantly accelerate DNN inference by 1.87× to 6.78×.
Huitian Wang, Fang Dong 0001, Wei Zhao 0023
ICPADS3
2019 ADDA: Adaptive Distributed DNN Inference Acceleration in Edge Computing Environment
abstract
Implementing intelligent mobile applications on IoT devices with DNN technology has become an inevitable trend. Due to the limitations of the size of DNN model deployed onto end devices and the instability of wide-area network transmission, either End-only mode or Cloud-only mode cannot guarantee the reasonable latency and recognition accuracy simultaneously. A better solution is to exploit the edge computing, where the existing edge computing execution framework and offloading mechanism for DNN inference suffer unnecessary computational overheads and underutilized computing capacity of end and edge. To address these shortcomings, an adaptive distributed DNN inference acceleration framework for edge computing environment is proposed in this paper, where DNN computation path optimization and DNN computation partition optimization are taken into consideration. The evaluations demonstrate that our method can effectively accelerate the DNN inference compared to the state-of-the-art methods.
Huitian Wang, Guangxing Cai, Zhaowu Huang, Fang Dong 0001
ICPADS1