Ziyi Han

dblp:215/7445 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 3 first-author · 5 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Faster, Smaller, and Smarter: Task-Aware Expert Merging for Online MoE Inference
Ziyi Han, Xutong Liu 0002, Ruiting Zhou, Xiangxiang Dai, John C. S. Lui
INFOCOM1
2026 Online Scheduling With Trajectory Prediction for Collaborative DNN Inference in Vehicular Networks
abstract
In recent years, deep neural networks (DNNs) have been extensively utilized to provide vehicular intelligent services. Given the limited computing capabilities of vehicles, collaborative vehicle-edge DNN inference has emerged as a promising approach. This method partitions the DNN, then distributes parts to the vehicle or the edge,e.g.,roadside unit (RSU), for sequential inferences. However, determining the optimal DNN partition is challenging due to the uneven load distribution of DNN models and varying road traffic conditions. Moreover, vehicle movement can cause loss of inference results if vehicles leave the RSU signal coverage. To this end, we propose a novel online learning-based collaborative DNN Inference frameworkMCI.MCIutilizes multiple RSUs to assist vehicles with sequential inference and ensure reliable data transmission. To reduce learning cost,MCIdesigns a trajectory prediction to analyze vehicle context before making decisions. Then,MCIcombines the classical EXP4 and LinUCB algorithms to learn system dynamics and make effective scheduling decisions. We prove thatMCIachieves a sublinear regret bound of$O(T^{3/4} \sqrt {\log T})$. Extensive experimental results show thatMCIreduces latency by up to 68% and has a lower failure rate, compared to state-of-the-art algorithms.
Ziyi Han, Ruiting Zhou, Haisheng Tan, John C. S. Lui
IEEE Trans. Netw.1
2025 Spatial-Aware Anchor Growth for 3D Gaussian Field Reconstruction
Ziyi Han, Penglin Li
ICIC (15)2
2024 SAFE: Intelligent Online Scheduling for Collaborative DNN Inference in Vehicular Network
abstract
Recent years have witnessed a widespread use of deep neural networks (DNNs) in providing various intelligent services, and vehicular networks are no exception. Given the limited computing capabilities of vehicles, collaborative vehicle-edge DNN inference has emerged as a viable alternative. This approach employs DNN partitioning, where a part of DNN is computed on vehicles, and the other part on the edge, e.g., roadside unit (RSU), aiming to enhance the inference accuracy and reduce the inference latency. In this setting, deriving an optimal DNN partitioning scheme becomes critical, yet challenging given the constant movement of vehicles and the highly dynamic wireless connections. Furthermore, vehicles may move out of the signal coverage of an RSU, making it difficult to receive the inference results. To this end, we propose a two-stage intelligent scheduling framework named Soft Actor-critic for discrete actions (SAC-D) based collaborative DNN inference FramEwork (SAFE). SAFE engages multiple RSUs to assist vehicles in completing inference tasks sequentially and ensuring reliable data transmission. It can learn the dynamic vehicular network and make scheduling decisions to minimize the overall latency of vehicle inference tasks. Extensive experimental results show that SAFE can reduce up to 80% of the overall latency with a lower failure rate, compared to four baselines.
Ruiting Zhou, Ziyi Han, Zhi Zhou 0006, Wei Wang 0030
CSCWD2
2024 Efficient Online DNN Inference with Continuous Learning in Edge Computing
abstract
Compressed edge DNN models usually experience decreasing model accuracy when performing inference due to data drift. To maintain the inference accuracy, retraining models with continuous learning is usually employed in the edge. However, online edge DNN inference with continuous learning faces new challenges. First, introducing retraining jobs leads to resource competition with the existing edge inference tasks, which will affect the inference latency. Second, retraining jobs and inference tasks exhibit significant differences in workload and latency requirements. These two jobs cannot adopt the same scheduling policy. To overcome the challenges, we propose an Online scheduling algorithm for INference with Continuous learning (OINC). OINC minimizes the weighted sum of the latency of inference tasks and the completion time of retraining jobs with limited edge resources, while ensuring the satisfaction of the inference task’s service level objective (SLO) and meeting the deadlines of retraining jobs. OINC first reserves a portion of resources to complete all current inference tasks and allocates the remaining resources to retraining jobs. Subsequently, based on the reserved resource ratio, OINC invokes two sub-algorithms to select edges and allocate resources for each inference task and retraining job respectively. Compared with six state-of-the-art algorithms, OINC can reduce the weighted sum by up to 23.7%, and increase the success rate by up to 35.6%.
Ruiting Zhou, Lei Jiao 0002, Ziyi Han, Jieling Yu
IWQoS4
2024 Multi-view and region reasoning semantic enhancement for image-text retrieval
Wengang Cheng, Ziyi Han, Lifang Wu
Multim. Syst.2
2024 Eris: An Online Auction for Scheduling Unbiased Distributed Learning Over Edge Networks
abstract
The emergence of edge intelligence has made smart IoT services (e.g.,video/audio surveillance, autonomous driving and smart city) a reality. To ensure the quality of service, edge service providers train unbiased models of distributed machine learning jobs over the local datasets collected by edge networks, and usually adopt the parameter server (PS) architecture. However, the training ofunbiased distributed learning(UDL) depends on geo-distributed data and edge resources, bringing a new challenge for service providers: how to effectively schedule and price UDL jobs such that the long-term system utility (i.e.,social welfare) can be maximized. In this paper, we propose an online auction-based scheduling algorithmEris, which determines the data workload, the number and the placement of concurrent workers and PSs for each arriving UDL job, and dynamically prices limited edge resources based on current resource consumption.Erisapplies a primal-dual framework which calls an efficient dual subroutine to schedule UDL jobs, achieving a good competitive ratio and pseudo-polynomial time complexity. To evaluate the effectiveness ofEris, we implement both a testbed and a large-scaled simulator. The results demonstrate thatErisoutperforms and achieves up to 44% more social welfare compared to state-of-the-art algorithms in today's cloud system.
Jinlong Pang, Ziyi Han, Ruiting Zhou, Renli Zhang, John C. S. Lui
IEEE Trans. Mob. Comput.2
2024 InSS: An Intelligent Scheduling Orchestrator for Multi-GPU Inference With Spatio-Temporal Sharing
abstract
As the applications of AI proliferate, it is critical to increase the throughput of online DNN inference services. Multi-process service (MPS) improves the utilization rate of GPU resources by spatial-sharing, but it also brings unique challenges. First, interference between co-located DNN models deployed on the same GPU must be accurately modeled. Second, inference tasks arrive dynamically online, and each task needs to be served within a bounded time to meet the service-level objective (SLO). Third, the problem of fragments has become more serious. To address the above three challenges, we propose anIntelligentScheduling orchestrator for multi-GPU inference servers with spatio-temporalSharing (InSS), aiming to maximize the system throughput.InSSexploits two key innovations: i) An interference-aware latency analytical model which estimates the task latency. ii) A two-stage intelligent scheduler is tailored to jointly optimize the model placement, GPU resource allocation and adaptively decides batch size by coupling the latency analytical model. Our prototype implementation on four NVIDIA A100 GPUs shows thatInSScan improve the throughput by up to 86% compared to the state-of-the-art GPU schedulers, while satisfying SLOs. We further show the scalability ofInSSon 64 GPUs.
Ziyi Han, Ruiting Zhou, Cheng-Zhong Xu 0001, Renli Zhang
IEEE Trans. Parallel Distributed Syst.1
2023 An Efficient Visible Light Positioning and Rotation Estimation System Using Two LEDs and a Photodiode Array
abstract
Existing visible light positioning systems suffer from high computational complexity or cannot output rotation estimation results, making it difficult to support indoor navigation. This paper introduces an indoor positioning system with two beacon light-emitting diodes (LEDs) and a photodiode array at the receiver. The photodiode array can estimate the angles of arrival of the light signals from the beacon LEDs, and the user coordinates can be expressed as closed-form functions of the LED coordinates and the measured light directional vectors. We also carry out asymptotic error analysis for the positioning algorithm, and the analytical results reveals important insights for the system design. Simulation results show that the system can achieve centimeter-level accuracy and low average rotation estimation error.
Yongbin Gong, Di Miao, Yuzheng Yang, Ziyi Han, Jingrui Li, Bingcheng Zhu, Lanting Fang, Liang Chen 0007
WCNC5
2022 Online scheduling algorithms for unbiased distributed learning over wireless edge networks
Jinlong Pang, Ziyi Han, Ruiting Zhou, Haisheng Tan, Yue Cao 0002
J. Syst. Archit.2
2021 Online Scheduling Unbiased Distributed Learning over Wireless Edge Networks
abstract
To realize high quality smart IoT services, such as intelligent video surveillance in Auto Driving and Smart City, tremendous amount of distributed machine learning jobs train unbiased models in wireless edge networks, adopting the parameter server (PS) architecture. Due to the large datasets collected geo-distributedly, the training of unbiased distributed learning (UDL) brings high response latency and bandwidth consumption. In this paper, we propose an online scheduling algorithm, Okita, to minimize both the latency cost and bandwidth cost in UDL. Okita schedules UDL jobs at each time slot to jointly decide the execution time window, the amount of training data, the number and the location of concurrent workers and PSs in each site. To evaluate the practical performance of Okita, we implement a testbed based on Kubernetes. Extensive experiments and simulations show that Okita can reduce up to 60% of total cost, compared with the state-of-the-art schedulers in cloud systems.
Ziyi Han, Ruiting Zhou, Jinlong Pang, Yue Cao 0002, Haisheng Tan
ICPADS1
2018 A remotely keyed file encryption scheme under mobile cloud computing
Li Yang 0005, Ziyi Han, Zhengan Huang, Jianfeng Ma 0001
J. Netw. Comput. Appl.2
2018 Efficient Multifactor Two-Server Authenticated Scheme under Mobile Cloud Computing
abstract
Because the authentication method based on username‐password has the disadvantage of easy disclosure and low reliability and the excess password management degrades the user experience tremendously, the user is eager to get rid of the bond of the password in order to seek a new way of authentication. Therefore, the multifactor biometrics‐based user authentication wins the favor of people with advantages of simplicity, convenience, and high reliability. Now the biometrics‐based (especially the fingerprint information) authentication technology has been extremely mature, and it is universally applied in the scenario of the mobile payment. Unfortunately, in the existing scheme, biometric information is stored on the server side. As thus, once the server is hacked by attackers to cause the leakage of the fingerprint information, it will take a deadly threat to the user privacy. Aiming at the security problem due to the fingerprint information in the mobile payment environment, we propose a novel multifactor two‐server authenticated scheme under mobile cloud computing (MTSAS). In the MTSAS, it divides the authentication method and authentication means; in the meanwhile, the user’s biometric characteristics cannot leave the user device. Thus, MTSAS avoids the fingerprint information disclosure, protects user privacy, and improves the security of the user data. In the same time, considering user actual requirements, different authentication factors depending on the privacy level of authentication are chosen. Security analysis proves that MTSAS has achieved the authentication purpose and met security requirements by the BAN logic. In comparison with other schemes, the result shows that MTSAS not only has the reasonable computational efficiency, but also keeps the superior communication cost.
Ziyi Han, Li Yang 0005, Sen Mu, Qiang Liu 0031
Wirel. Commun. Mob. Comput.1