Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zhekai Duan

dblp:348/6573 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0002-0283-8419ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 50% GPUs and heterogeneous computing · 50%
Computer networks
1 paper
Edge and fog computing · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 87% Vision and language · 13%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU computing
0.912025
Improving Tridiagonalization Performance on GPU Architectures · PPoPP 2025
GPUs and heterogeneous computing
GPU kernel optimization
0.912025
Improving Tridiagonalization Performance on GPU Architectures · PPoPP 2025
High-performance computing
numerical linear algebra
0.912025
Improving Tridiagonalization Performance on GPU Architectures · PPoPP 2025
High-performance computing › numerical linear algebra
tridiagonalization
0.912025
Improving Tridiagonalization Performance on GPU Architectures · PPoPP 2025
Machine learning › Efficient and distributed learning › model deployment
edge deployment
0.812024
Self-adapting Large Visual-Language Models to Edge Devices Across Visual Modalities · ECCV (28) 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
Self-adapting Large Visual-Language Models to Edge Devices Across Visual Modalities · ECCV (28) 2024
Edge and fog computing › mobile edge computing
computation offloading
0.812024
Mobility-Aware Computation Offloading With Load Balancing in Smart City Networks Using MEC Federation · IEEE Trans. Mob. Comput. 2024
Edge and fog computing
mobile edge computing
0.812024
Mobility-Aware Computation Offloading With Load Balancing in Smart City Networks Using MEC Federation · IEEE Trans. Mob. Comput. 2024
Computer vision › Vision and language › vision-language model
vision-language model adaptation
0.212024
Self-adapting Large Visual-Language Models to Edge Devices Across Visual Modalities · ECCV (28) 2024

Methods — techniques the papers use, named apart from their topics

transformer · 1.5lyapunov optimization · 1.5deep deterministic policy gradient · 1.5clustering · 1.5double blocking band reduction · 0.9bulge chasing · 0.9knowledge distillation · 0.8
YearPublicationVenuePosition
2025 Improving Tridiagonalization Performance on GPU Architectures
abstract
Tridiagonalization, which is a key step in symmetric eigenvalue decomposition (EVD), aims to convert a symmetric matrix to a tridiagonal form. In Nvidia's cuSOLVER library, the FP64 precision tridiagonalization process only reach 2.1 TFLOPs out of 67 TFLOPs on H100 GPU, and it consumes a significant portion of the elapsed time in the entire EVD process, accounting for over 97%. Thus, improving the tridiagonalization performance is crucial on accelerating EVD. In this paper, we analyze the reasons behind the suboptimal performance of tridiagonalization on GPU architectures, and we propose a new double blocking band reduction algorithm along with an implementation of GPU-based bulge chasing to improve the tridiagonalization performance. Through experimental evaluation, the proposed FP64 precision tridiagonalization method yields up to 19.6 TFLOPs which is 9.3x and 5.2x faster compared cuSOVLER and MAGMA, respectively.
Zhekai Duan, Zitian Zhao, Saiqi Zheng, Qiao Li 0001, Xu Jiang 0004, Shaoshuai Zhang
PPoPP2
2025 Human-object interaction detector with unsupervised domain adaptation
Yamin Cheng, Zhekai Duan, Hualong Huang, Zhi Wang 0020
Knowl. Based Syst.2
2024 Self-adapting Large Visual-Language Models to Edge Devices Across Visual Modalities
Kaiwen Cai, Zhekai Duan, Gaowen Liu, Charles Fleming, Xiaoxuan Lu 0001
ECCV (28)2
2024 Battery-Care Resource Allocation and Task Offloading in Multi-Agent Post-Disaster MEC Environment
abstract
Being an up-and-coming application scenario of mobile edge computing (MEC), the post-disaster rescue suffers multitudinous computing-intensive tasks but unstably guaranteed network connectivity. In rescue environments, quality of service (QoS), such as task execution delay, energy consumption and battery state of health (SoH), is of significant meaning. This paper studies a multi-user post-disaster MEC environment with unstable 5G communication, where device-to-device (D2D) link communication and dynamic voltage and frequency scaling (DVFS) are adopted to balance each user's requirement for task delay and energy consumption. A battery degradation evaluation approach to prolong battery lifetime is also presented. The distributed optimization problem is formulated into a mixed cooperative-competitive (MCC) multi-agent Markov decision process (MAMDP) and is tackled with recurrent multi-agent Proximal Policy Optimization (rMAPPO). Extensive simulations and comprehensive comparisons with other representative algorithms clearly demonstrate the effectiveness of the proposed rMAPPO-based offloading scheme.
Yiwei Tang, Hualong Huang, Wenhan Zhan, Geyong Min, Zhekai Duan, Yuchuan Lei
WCNC5
2024 Optimal service caching, pricing and task partitioning in mobile edge computing federation
Hualong Huang, Zhekai Duan, Wenhan Zhan, Geyong Min, Kai Peng 0002
Future Gener. Comput. Syst.2
2024 Mobility-Aware Computation Offloading With Load Balancing in Smart City Networks Using MEC Federation
abstract
Internet-of-Things (IoT) has played a critical role in developing sustainable smart cities and emerging numerous latency-sensitive IoT applications. Mobile edge computing (MEC) federation has the capability to incorporate a transparent resource management approach, which enables the sharing and utilization of MEC services from edge infrastructure providers (EIPs) and provides agile access services to mobile devices (MDs). In this paper, we investigate the joint optimization problem of computation offloading, task migration, and resource allocation in the MEC federation. The objective is to minimize the weighted sum of latency and energy consumption while maintaining load balancing under the constraint of the long-term migration cost budget of EIPs. To address the problem, we decompose it into two sub-problems: 1) the MDs clustering sub-problem and 2) the sub-problem of joint computation offloading, task migration, and resource allocation. Firstly, an MDs clustering matching (MDCM) algorithm is proposed to cluster the MDs in edge servers (ESs) according to the differences in channel gains. Afterward, the second sub-problem is simplified by the Lyapunov optimization technique, and then we propose a Transformer-based mobility prediction model and a decentralized deep deterministic policy gradient (DDPG)-based framework to solve it. Extensive simulation results demonstrate the cost-efficiency of the proposed algorithm.
Hualong Huang, Wenhan Zhan, Geyong Min, Zhekai Duan, Kai Peng 0002
IEEE Trans. Mob. Comput.4
2023 Distributed Dependent Task Offloading in CPU-GPU Heterogenous MEC: A Federated Reinforcement Learning Approach
abstract
Mobile edge computing (MEC) has emerged as a promising paradigm to enable computation-intensive and latency-sensitive mobile applications by offloading tasks to proximal edge servers. This paper proposes a novel federated reinforcement learning framework called Transformer-based Federated Soft Actor-Critic (TFSAC) to address a joint computation offloading and resource scheduling problem in a CPU-GPU heterogeneous MEC network while preserving privacy. Specifically, a graph attention network (GAT) extracts high-dimensional features from the task dependency graph. Rather than simply averaging weights, TFSAC applies transformer encoders to learn contextual relationships between agents and enable selective aggregation of relevant knowledge during federated model training to preserve agents’ privacy. Experiments on real-world trace data demonstrate TFSAC’s superiority over benchmarks in maximizing quality-of-service (QoS) across configurations.
Hualong Huang, Zhekai Duan, Wenhan Zhan, Zhi Wang 0020, Zitian Zhao
TrustCom2