Zitian Zhao

dblp:229/3263 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0001-8605-0938ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 50% GPUs and heterogeneous computing · 50%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Computer networks
1 paper
Edge and fog computing · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference acceleration
1.012026
EdgeSD: Efficient Speculative Decoding With Vision-Decoding Disaggregation for MLLM Inference in Edge-Cloud Networks · IEEE Trans. Mob. Comput. 2026
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding
1.012026
EdgeSD: Efficient Speculative Decoding With Vision-Decoding Disaggregation for MLLM Inference in Edge-Cloud Networks · IEEE Trans. Mob. Comput. 2026
Edge and fog computing
edge-cloud collaboration
1.012026
EdgeSD: Efficient Speculative Decoding With Vision-Decoding Disaggregation for MLLM Inference in Edge-Cloud Networks · IEEE Trans. Mob. Comput. 2026
GPUs and heterogeneous computing
GPU computing
0.912025
Improving Tridiagonalization Performance on GPU Architectures · PPoPP 2025
GPUs and heterogeneous computing
GPU kernel optimization
0.912025
Improving Tridiagonalization Performance on GPU Architectures · PPoPP 2025
High-performance computing
numerical linear algebra
0.912025
Improving Tridiagonalization Performance on GPU Architectures · PPoPP 2025
High-performance computing › numerical linear algebra
tridiagonalization
0.912025
Improving Tridiagonalization Performance on GPU Architectures · PPoPP 2025
Machine learning › Efficient and distributed learning
model compression
0.312026
EdgeSD: Efficient Speculative Decoding With Vision-Decoding Disaggregation for MLLM Inference in Edge-Cloud Networks · IEEE Trans. Mob. Comput. 2026
Machine learning › Efficient and distributed learning › model compression › token compression
token merging
0.312026
EdgeSD: Efficient Speculative Decoding With Vision-Decoding Disaggregation for MLLM Inference in Edge-Cloud Networks · IEEE Trans. Mob. Comput. 2026

Methods — techniques the papers use, named apart from their topics

speculative decoding · 2.0parallel delta-stepping · 2.0image token merging · 2.0double blocking band reduction · 0.9bulge chasing · 0.9
YearPublicationVenuePosition
2026 EdgeSD: Efficient Speculative Decoding With Vision-Decoding Disaggregation for MLLM Inference in Edge-Cloud Networks
abstract
The deployment of multimodal large language models (MLLMs) in edge-cloud networks faces critical challenges, including computational resource heterogeneity, memory bottlenecks, and bandwidth constraints. To address these issues, we propose EdgeSD, a novel framework that accelerates MLLM inference by integrating speculative decoding (SD) with edge-cloud collaboration. First, EdgeSD decouples the vision encoding and decoding processes of the draft MLLM across heterogeneous edge servers (ESs). This disaggregation architecture overcomes single-node memory constraints, enabling optimized resource utilization and high-resolution input processing. Second, to resolve the communication bottleneck and computational burden inherent in this distributed architecture, EdgeSD integrates a bandwidth-aware dynamic image token merging (ITM) method. Unlike general pruning techniques, this EdgeSD-specific ITM method focuses on minimizing inter-ES transmission latency for vision-decoding disaggregation while maintaining draft quality. Third, to optimize SD efficiency on consumer-grade ESs, EdgeSD employs an adaptive and scalable token tree structure solved using a parallel delta-stepping algorithm. This structure maximizes the number of accepted tokens under strict edge latency constraints. Extensive experiments on six multimodal datasets and five benchmarks with various MLLM pairs demonstrate that EdgeSD achieves substantial acceleration and throughput gains in edge-cloud collaboration scenarios using a lightweight draft MLLM, achieving 3.04-5.12x speedup compared to baseline methods.
Hualong Huang, Wenhan Zhan, Hancong Duan, Kai Peng 0002, Geyong Min, Zijia Zhao, Zitian Zhao, Yalan Ye
IEEE Trans. Mob. Comput.7
2025 Improving Tridiagonalization Performance on GPU Architectures
abstract
Tridiagonalization, which is a key step in symmetric eigenvalue decomposition (EVD), aims to convert a symmetric matrix to a tridiagonal form. In Nvidia's cuSOLVER library, the FP64 precision tridiagonalization process only reach 2.1 TFLOPs out of 67 TFLOPs on H100 GPU, and it consumes a significant portion of the elapsed time in the entire EVD process, accounting for over 97%. Thus, improving the tridiagonalization performance is crucial on accelerating EVD. In this paper, we analyze the reasons behind the suboptimal performance of tridiagonalization on GPU architectures, and we propose a new double blocking band reduction algorithm along with an implementation of GPU-based bulge chasing to improve the tridiagonalization performance. Through experimental evaluation, the proposed FP64 precision tridiagonalization method yields up to 19.6 TFLOPs which is 9.3x and 5.2x faster compared cuSOVLER and MAGMA, respectively.
Zhekai Duan, Zitian Zhao, Saiqi Zheng, Qiao Li 0001, Xu Jiang 0004, Shaoshuai Zhang
PPoPP3
2025 Dynamic Model Deployment, Batch Scheduling, and Resource Allocation in MLLM-Enabled Edge-Cloud Networks: A Multiagent Two-Timescale DRL Approach
abstract
The deployment of multimodal large language models (MLLMs) on resource-constrained mobile devices poses significant challenges due to their high computational demands. This paper introduces a novel two-timescale optimization framework for efficient MLLM inference in Edge-Cloud networks, addressing the problem of multi-timescale resource management by jointly optimizing slow-timescale MLLMs deployment decisions and fast-timescale batch scheduling, GPU resource allocation, and bandwidth allocation under dynamic network conditions and spatiotemporal request heterogeneity. Our key innovation is a hierarchical twin delayed deep deterministic policy gradient (HALTD3) algorithm that integrates attention mechanisms and long short-term memory networks to optimize slow-timescale MLLMs deployment and fast-timescale resource allocation, minimizing weighted system costs including deployment cost, end-to-end latency, and energy consumption, while meeting stringent quality-of-service requirements. Extensive experiments demonstrate that the HALTD3 algorithm substantially outperforms baseline methods in reducing system costs across diverse MLLM workloads and dynamic network scenarios, validating its effectiveness for practical edge-cloud collaborative inference.
Hualong Huang, Yongkang Du, Wenhan Zhan, Hancong Duan, Kai Peng 0002, Yamin Cheng, Yalan Ye, Zitian Zhao
IEEE Internet Things J.8
2023 Distributed Dependent Task Offloading in CPU-GPU Heterogenous MEC: A Federated Reinforcement Learning Approach
abstract
Mobile edge computing (MEC) has emerged as a promising paradigm to enable computation-intensive and latency-sensitive mobile applications by offloading tasks to proximal edge servers. This paper proposes a novel federated reinforcement learning framework called Transformer-based Federated Soft Actor-Critic (TFSAC) to address a joint computation offloading and resource scheduling problem in a CPU-GPU heterogeneous MEC network while preserving privacy. Specifically, a graph attention network (GAT) extracts high-dimensional features from the task dependency graph. Rather than simply averaging weights, TFSAC applies transformer encoders to learn contextual relationships between agents and enable selective aggregation of relevant knowledge during federated model training to preserve agents’ privacy. Experiments on real-world trace data demonstrate TFSAC’s superiority over benchmarks in maximizing quality-of-service (QoS) across configurations.
Hualong Huang, Zhekai Duan, Wenhan Zhan, Zhi Wang 0020, Zitian Zhao
TrustCom6
2023 Rethinking vision transformer through human-object interaction detection
Yamin Cheng, Zitian Zhao, Hancong Duan
Eng. Appl. Artif. Intell.2
2019 A lighten CNN-LSTM model for speaker verification on embedded devices
Zitian Zhao, Hancong Duan, Geyong Min, Zilei Huang, Xian Zhuang, Hao Xi, Meirong Fu
Future Gener. Comput. Syst.1