Li Zhang 0133

dblp:89/5992-133 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0003-0779-8310ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Computer networks · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Enabling Efficient Transmission of Satellite-to-Ground Downlinks via Throughput Prediction
Geyang Li, Li Zhang 0133, Chuanxiu Chi, Shangguang Wang
INFOCOM2
2026 Prototyping and Analyzing Mobile SoC Clusters as Modern Edge Servers
abstract
The rapidly growing edge computing platforms, coupled with the still imperfect edge infrastructure, present an excellent opportunity for the emergence of new edge hardware. However, it remains unclear whether alternative architectures built from energy-efficient mobile System-on-Chips (SoCs) can meet the stringent performance, cost, and energy demands of modern edge workloads. In this paper, we propose a new type of edge server composed of 60 Qualcomm Snapdragon 865 mobile SoCs in a 2U rack, referred to as SoC Cluster. We demonstrate its successful deployment on existing edge cloud platforms and its ability to natively serve mobile cloud gaming services. Despite the emergence of new hardware on edge platforms and its successful operation in serving mobile cloud gaming, our trace analysis revealed low hardware utilization and significant dynamic fluctuations in usage. To assess its broader applicability, we conducted the first measurement study of SoC Cluster to reveal its ability to run two popular and modern edge applications: deep learning inference and video transcoding. We developed a cross-platform benchmark suite to evaluate throughput, latency, power consumption, and application-specific metrics like video quality. We then directly compare SoC Cluster with a traditional edge server equipped with Intel CPUs and NVIDIA GPUs in terms of energy efficiency, space efficiency, and monetary cost. Results show that SoC Cluster exhibits up to 6.5? higher energy efficiency and 7.7? higher space efficiency. We also disclose its limitations in serving computation-intensive workloads such as large deep learning models. The outcomes provide insightful implications and offer practical direction for refining SoC Cluster toward broader deployment in edge scenarios.
Li Zhang 0133, Boqing Shi, Xiang Li 0067, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang, Mengwei Xu 0001
IEEE Trans. Mob. Comput.1
2025 ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents
abstract
Recent advancements in integrating large language models (LLMs) with application programming interfaces (APIs) have gained significant interest in both academia and industry. Recent work demonstrates that these API-based agents exhibit relatively strong autonomy and planning capabilities. However, their ability to handle multi-dimensional difficulty levels, diverse task types, and real-world demands remains unknown. In this paper, we introduce \textsc{ShortcutsBench}, a large-scale benchmark for the comprehensive evaluation of API-based agents in solving real-world complex tasks. \textsc{ShortcutsBench} includes a wealth of real APIs from Apple Inc., refined user queries, human-annotated high-quality action sequences, detailed parameter filling values, and parameters requesting necessary input from the system or user. We revealed how existing benchmarks~/~datasets struggle to accommodate the advanced reasoning capabilities of existing more intelligent LLMs. Moreover, our extensive evaluation of agents built with $5$ leading open-source (size $\geq$ 57B) and $5$ closed-source LLMs (e.g. Gemini-1.5-Pro and GPT-4o-mini) with varying intelligence level reveals significant limitations of existing API-based agents in the whole process of handling complex queries related to API selection, parameter filling, and requesting necessary input from the system and the user. These findings highlight the great challenges that API-based agents face in effectively fulfilling real and complex user queries. All datasets, code, experimental logs, and results are available at https://github.com/EachSheep/ShortcutsBench
Haiyang Shen, Desong Meng, Dongqi Cai 0001, Li Zhang 0133, Mengwei Xu 0001, Yun Ma 0002
ICLR6
2025 AndroidWMSearch: Mobile Agents Tree Search with World Model
abstract
Mobile agents powered by large language models (LLMs) have demonstrated remarkable potential in automating operations on mobile devices. Recent studies have demonstrated that incorporating tree search methods and increasing testtime computation can enhance an agent's multi-step reasoning and planning capabilities. However, unlike simulated sandbox environments, Android is a dynamic environment with many irreversible operations, making tree search backtracking less feasible on the Android platform. To address this challenge, we propose AndroidWMSearch, a novel agent tree search framework that leverages a world model to emulate the Android environment. This framework allows the agent to evaluate and rank candidate actions through simulation before actual execution. We systematically explore this paradigm by: (1) Proposing a model-based Android tree search framework, AndroidWMSearch, in which LLMs are utilized both as world models and value functions. (2) Training specialized LLMs to act as world models, utilizing a scalable data synthesis pipeline for the training process. On the AndroidWorld benchmarks, our AndroidWMSearch surpasses the T3A agent by 4.7%, underscoring the effectiveness of our proposed framework. Moreover, utilizing our AndroidWM-7B, which is specifically trained for Android environments, as the world model results in a 3.0% performance gain compared to employing GPT-4o. These findings highlight the importance and efficacy of training a dedicated world model tailored for mobile agents.
Xianqing Jia, Li Zhang 0133, Mengwei Xu 0001
ICPADS2
2024 SoCFlow: Efficient and Scalable DNN Training on SoC-Clustered Edge Servers
abstract
SoC-Cluster, a novel server architecture composed of massive mobile system-on-chips (SoCs), is gaining popularity in industrial edge computing due to its energy efficiency and compatibility with existing mobile applications. However, we observe that the deployed SoC-Cluster servers are not fully utilized, because the hosted workloads are mostly user-triggered and have significant tidal phenomena. To harvest the free cycles, we propose to co-locate deep learning tasks on them.
Daliang Xu, Mengwei Xu 0001, Chiheng Lou, Li Zhang 0133, Gang Huang 0001, Xin Jin 0008, Xuanzhe Liu
ASPLOS (1)4
2024 Poster: Efficient and Accurate Mobile Task Automation through Learning from Code
abstract
With the emergence and continuous prosperity of large language models (LLMs), artificial intelligence (AI) agents have experienced rapid advancements. Most mobile AI agents merely imitate human operations, executing actions based on the human user interface (UI). The restricted input impairs the efficiency and accuracy of mobile tasks. We propose an unexplored approach: learning from the source code. Source code is the plain interaction for mobile applications, which can be used to enhance the UI understanding of mobile agents, improve action execution accuracy, and reduce the average action completion steps. The implementation of the agent prototype is preliminary evaluated on 5 open-source applications and 22 tasks, reducing the average number of task completion steps by 54%.
Shihe Wang, Li Zhang 0133, Mengwei Xu 0001
MobiSys2
2024 LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Task Automation
abstract
The emergent large language/multimodal models facilitate the evolution of mobile agents, especially in mobile UI task automation. However, existing evaluation approaches, which rely on human validation or established datasets to compare agent-predicted actions with predefined action sequences, are unscalable and unfaithful. To overcome these limitations, this paper presents LlamaTouch, a testbed for on-device mobile UI task execution and faithful, scalable task evaluation. By observing that the task execution process only transfers UI states, LlamaTouch employs a novel evaluation approach that only assesses whether an agent traverses all manually annotated, essential application/system states. LlamaTouch comprises three key techniques: (1) On-device task execution that enables mobile agents to interact with realistic mobile environments for task execution. (2) Fine-grained UI component annotation that merges pixel-level screenshots and textual screen hierarchies to explicitly identify and precisely annotate essential UI components with a rich set of designed annotation primitives. (3) A multi-level application state matching algorithm that utilizes exact and fuzzy matching to accurately detect critical information in each screen, even with unpredictable UI layout/content dynamics. LlamaTouch currently incorporates four mobile agents and 496 tasks, encompassing both tasks in the widely-used datasets and our self-constructed ones to cover more diverse mobile applications. Evaluation results demonstrate LlamaTouch’s high faithfulness of evaluation in real-world mobile environments and its better scalability than human validation. LlamaTouch also enables easy task annotation and integration of new mobile agents. Code and dataset are publicly available at https://github.com/LlamaTouch/LlamaTouch.
Li Zhang 0133, Shihe Wang, Xianqing Jia, Zhihan Zheng, Yunhe Yan, Longxi Gao, Yuanchun Li 0003, Mengwei Xu 0001
UIST1
2024 More is Different: Prototyping and Analyzing a New Form of Edge Server with Massive Mobile SoCs
Li Zhang 0133, Zhe Fu 0005, Boqing Shi, Xiang Li 0067, Rujin Lai, Chenyang Yang 0004, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang, Mengwei Xu 0001
USENIX ATC1
2024 High-density Mobile Cloud Gaming on Edge SoC Clusters
Li Zhang 0133, Shangguang Wang, Mengwei Xu 0001
USENIX ATC1
2024 Efficient, Scalable, and Sustainable DNN Training on SoC-Clustered Edge Servers
abstract
In the realm of industrial edge computing, a novel server architecture known as SoC-Cluster, characterized by its aggregation of numerous mobile systems-on-chips (SoCs), has emerged as a promising solution owing to its enhanced energy efficiency and seamless integration with prevalent mobile applications. Despite its advantages, the utilization of SoC-Cluster servers remains unsatisfactory, primarily attributed to the tidal patterns of user-initiated workloads. To address such inefficiency, we introduceSoCFlow+, a pioneering framework designed to facilitate the co-location of deep learning training tasks on SoC-Cluster servers, thereby optimizing resource utilization.SoCFlow+incorporates three novel techniques tailored to mitigate the inherent limitations of commercial SoC-Cluster servers. First, it employs group-wise parallelism complemented by delayed aggregation, a strategy engineered to enhance the training efficiency and scalability of deep learning models, effectively circumventing network bottlenecks. Second, it integrates a data-parallel mixed-precision training algorithm, optimized to exploit the heterogeneous processing capabilities inherent to mobile SoCs fully. Third,SoCFlow+employs an underclocking-aware workload re-balanacing mechanism to tackle the training performance degradation caused by the thermal control of mobile SoCs. Through rigorous experimental validation,SoCFlow+achieves a convergence speedup ranging from 1.6× to 740× across 32 SoCs, compared to conventional benchmarks. Furthermore, when juxtaposed with commodity GPU servers (e.g., NVIDIA V100) under identical power constraints,SoCFlow+not only exhibits comparable training speed but also achieves a remarkable reduction in energy consumption by a factor of 2.31× to 10.23×, all while preserving convergence accuracy.
Mengwei Xu 0001, Daliang Xu, Chiheng Lou, Li Zhang 0133, Gang Huang 0001, Xin Jin 0008, Xuanzhe Liu
IEEE Trans. Mob. Comput.4
2022 Position Paper: Renovating Edge Servers with ARM SoCs
abstract
Edge servers are key to the success of edge computing. Compared to cloud servers, edge servers suffer from a more constrained and costly electricity supply due to their dense, near-population deployment. Towards higher energy efficiency, we propose an extreme design of edge servers - SoC-Cluster that consists of massive, inter-connected ARM SoCs. Indeed, such SoC-Clusters have already been adopted to serve the cloud gaming application in the wild. In this paper, we present a concrete implementation of a COTS SoC-Cluster and its hardware specifications. We then discuss the potential killer applications that such SoC-Cluster can well serve and the major challenges to be solved. We also dive deep into two of such applications (live video transcoding and deep learning serving) and carry out a measurement study to demystify the application performance of SoC-Cluster. The results reveal that, compared to traditional servers, SoC-Cluster not only can reduce energy consumption but even deliver higher workload throughput in certain scenarios. Finally, we conclude the paper and discuss the primary research directions that can be explored by our community from applications, software, and hardware aspects.
Mengwei Xu 0001, Li Zhang 0133, Shangguang Wang
SEC2
2021 From cloud to edge: a first look at public edge platforms
abstract
Public edge platforms have drawn increasing attention from both academia and industry. In this study, we perform a first-of-its-kind measurement study on a leading public edge platform that has been densely deployed in China. Based on this measurement, we quantitatively answer two critical yet unexplored questions. First, from end users' perspective, what is the performance of commodity edge platforms compared to cloud, in terms of the end-to-end network delay, throughput, and the application QoE. Second, from the edge service provider's perspective, how are the edge workloads different from cloud, in terms of their VM subscription, monetary cost, and resource usage. Our study quantitatively reveals the status quo of today's public edge platforms, and provides crucial insights towards developing and operating future edge services.
Mengwei Xu 0001, Zhe Fu 0005, Xiao Ma 0009, Li Zhang 0133, Feng Qian 0001, Shangguang Wang, Xuanzhe Liu
Internet Measurement Conference4