VLDB 2026 Research / reviewers in the wild / expert
Yuanchun Li 0003
dblp:87/4523-3
· DBLP profile ↗
59ranked-venue papers
7as first author
52since 2021 · last 2026
0000-0002-1591-2526ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 25 · 25 since 2021Software engineering, systems software and programming languages · 16 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Agentic Textual Graph Retrieval with Synthetic Stepwise SupervisionabstractIntegrating textual graphs into Large Language Models (LLMs) is promising for complex graph-based QA. However, a key bottleneck is retrieving informative yet compact subgraphs that fit the LLM context. Existing retrievers often struggle, relying either on shallow embedding similarity or costly interactive policies that require excessive supervision. To address these challenges, we introduce an agentic textual graph reasoning framework featuring an LLM-based retriever trained with synthetic stepwise supervision. Rather than relying on final answer rewards which often yield sparse and unstable signals, we optimize the retriever by evaluating each step against offline-extracted golden subgraphs. Our approach distills golden subgraphs via a specialized data synthesis pipeline to formulate dense rewards, facilitating a two-stage training scheme that effectively learns the interactive graph exploration policy. Based on extensive experiments on three common datasets in comparison with seven strong baselines, our approach achieves an average improvement of 15.6% in accuracy and 17.2% in F1 score. The advantage is even higher in more complicated multi-hop reasoning tasks. Ge Chang 0002, Jinbo Su, Yuhao Shang, Huiwen Zheng, Hongli Ma, Yuanchun Li 0003, Yunxin Liu 0001 |
ACL (1) | 9 |
| 2026 | Benchmarking LLM's Capability in Reasoning over Conflicting Web ReferencesabstractLarge language models (LLMs) integrated with retrieval-augmented generation (RAG) have become a dominant framework for building intelligent assistants.In real-world applications such as ChatGPT with web search, the retrieved document often comes from diverse, potentially unreliable sources and may contain inconsistent claims.Unlike traditional search engines that rely on users to manually compare information, LLM-based systems typically feed all retrieved content into the model's context, requiring LLMs to autonomously identify, differentiate, and reason over conflicting viewpoints.Unlike mainstream LLM evaluation tasks like math and code generation that are primarily focused on reasoning with factual context, question-answering with multi-source references requires fundamentally different capabilities to identify and reason over knowledge contradictions.In this paper, we introduce CONFRAG, a benchmark for evaluating LLMs' reasoning capability over real-world conflicting documents retrieved from the web.It consists of 1,814 real-world questions, each paired with an average of 9.58 retrieved paragraphs from heterogeneous online sources.A total of 57.2% of the questions exhibit explicit contradictions.We further propose three structured evaluation tasks, answer clustering, answer coverage, and reason coverage, to quantify a model's ability to organize and explain contradictory content.Experiments with state-of-the-art models such as GPT-4.1 and Claude-3-7-Sonnet reveal substantial performance gaps, highlighting the need for more targeted research in contradiction-aware question answering.To the best of our knowledge, CONFRAG is the first benchmark specifically designed to evaluate contradiction-aware reasoning on real-world long web documents. Yizhen Yuan, Yuanchun Li 0003, Yunxin Liu 0001 |
ACL (1) | 4 |
| 2026 | VoLLM: Smoothness-aware Serving of LLM-powered Voice Q&A via Adaptive Preemption
Yuanchun Li 0003, Ju Ren 0001, Lichen Pang, Shansong Yang, Yunxin Liu 0001 |
IPDPS | 2 |
| 2026 | Mobile GUI Agents under Real-world Threats: Are We There Yet?abstractRecent years have witnessed a rapid development of mobile GUI agents powered by large language models (LLMs), which can autonomously execute diverse device-control tasks based on natural language instructions. The increasing accuracy of these agents on standard benchmarks has raised expectations for large-scale real-world deployment, and there are already several commercial agents released and used by early adopters. However, are we really ready for GUI agents integrated into our daily devices as system building blocks? We argue that an important pre-deployment validation is missing to examine whether the agents can maintain their performance under real-world threats. Specifically, unlike existing common benchmarks that are based on simple static app contents (they have to do so to ensure environment consistency between different tests), real-world apps are filled with contents from untrustworthy third parties, such as advertisement emails, user-generated posts and medias, etc. These contents may inevitably appear in the agents' observation space and influence the task execution process. Systematic investigation of this problem is challenging since the real-world app contents are significantly skewed—testing on normal real-world apps usually cannot uncover any potential risk since most app contents are benign. To this end, we introduce a scalable app content instrumentation framework to enable flexible and targeted content modifications within existing applications. Leveraging this framework, we create a test suite comprising both a dynamic task execution environment and a static dataset of challenging GUI states. The dynamic environment encompasses 122 reproducible tasks, and the static dataset consists of over 3,000 scenarios constructed from commercial apps. We perform experiments on both open-source and commercial GUI agents. Our findings reveal that all examined agents can be significantly degraded due to third-party contents, with an average misleading rate of 42.0% and 36.1% in dynamic and static environments respectively. The framework and benchmark has been released at https://agenthazard.github.io. Guohong Liu 0002, Jialei Ye, Wei Liu 0302, Pengzhi Gao, Jian Luan 0001, Yuanchun Li 0003, Yunxin Liu 0001 |
MobiSys | 7 |
| 2026 | AgentProg: Empowering Long-Horizon GUI Agents with Program-guided Context ManagementabstractThe rapid development of mobile GUI agents has stimulated growing research interest in long-horizon task automation. However, building agents for these tasks faces a critical bottleneck: the reliance on ever-expanding interaction history incurs substantial context overhead. Existing context management and compression techniques often fail to preserve vital semantic information, leading to degraded task performance. We propose AgentProg, a program-guided approach for agent context management that reframes the interaction history as a program with variables and control flow. By organizing information according to the structure of program, this structure provides a principled mechanism to determine which information should be retained and which can be discarded. We further integrate a global belief state mechanism inspired by Belief MDP framework to handle partial observability and adapt to unexpected environmental changes. Experiments on AndroidWorld and our extended long-horizon task suite demonstrate that AgentProg has achieved state-of-the-art success rates on these benchmarks. More importantly, it maintains robust performance on long-horizon tasks while baseline methods experience catastrophic degradation. Our system is open-sourced at https://github.com/MobileLLM/AgentProg. Shizuo Tian, Hao Wen 0004, Shanhui Zhao, Guohong Liu 0002, Ju Ren 0001, Yunxin Liu 0001, Yuanchun Li 0003 |
MobiSys | 9 |
| 2026 | An Efficient Context Management System for On-Device LLMaaSabstractLarge language models (LLMs) are renovating the mobile AI, catalyzing novel mobile applications such as UI task automation. A new paradigm of mobile AI ecosystem emerges in the LLM era: LLM as a mobile OS service (LLMaaS), where LLM runs as a system service and exposes its functionality (language understanding and generation) to third-party apps. As a giant step towards on-device LLMaaS, this work presents Libra, a system that tackles with challenge in managing persistent LLM contexts (KV cache) under tight memory constraint. Libra manages the LLM contexts based on the fine-grained, chunk-wise, globally-optimized KV cache compression and swapping. Specifically, it integrates three novel techniques: (1) Tolerance-Aware Compression applies different compression rates to each chunk based on their attention scores. (2) Swapping-Recompute Pipeline simultaneously swaps and recomputes the LLM contexts to improve the hardware resource utilization. (3) Chunk Lifecycle Management judiciously determines when and what to evict to reduce the context switching overhead. Through comprehensive experiments on various edge devices, Libra reduces the context switching latency by up to 20 × and on average 9.7 × compared to competitive baselines. Wangsong Yin, Mengwei Xu 0001, Yuanchun Li 0003, Xuanzhe Liu |
SenSys | 3 |
| 2026 | Training With Integer-Only Arithmetic: Energy-Efficient Federated Learning With Mobile DSP OffloadingabstractAI is making mobile applications increasingly cooler, but also introduces serious privacy risks due to the extensive user data collection. Federated learning (FL), as a privacy-preserving machine learning paradigm, enables mobile devices to collaboratively learn a shared prediction model while keeping all training data on devices. However, a key obstacle towards practical cross-device FL training is the huge energy consumption, especially for lightweight mobile devices. Prior literature mostly optimizes the convergence speed and network communication cost. In this work, we first perform the experimental analysis of improving FL performance through low-precision training with energy-friendly Digital Signal Processor (DSP) on mobile devices. Then, we demonstrate that directly integrating the state-of-the-art INT8 (8-bit integer) training algorithm and classic FL protocols will significantly degrade the model accuracy. Finally, we propose a novel FL protocol, namelyQ-FedUpdate, incorporates two critical techniques: error-compensated aggregation and pipelined batch quantization. The former can ensure the tiny model updates be accumulated and take effects, and the latter can improve the DSP cache hit rate to reduce the context switching. Extensive experiments show that,Q-FedUpdatecan effectively reduce the on-device energy consumption by 21×, and accelerate the FL convergence by 6.1× with only 2% accuracy loss. Jinliang Yuan, Daliang Xu, Mengwei Xu 0001, Yuanchun Li 0003, Xuanzhe Liu, Yunhao Liu 0001, Shangguang Wang |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | PFHAR: Practically Adopting Multi-Modal Foundation Model for Human Activity Recognition Through Edge-Cloud Collaborative LearningabstractMulti-modal human activity recognition (HAR) is a key technology for a wide range of applications and has received widespread attention in recent years. However, the difficulty of achieving generalizability in multi-modal sensing models, combined with heterogeneous and unlabeled downstream data, significantly hinders their broader adoption. In this work, we proposePFHAR, a unified framework for practically adopting multi-modal foundation HAR model to target user groups.PFHARuses a novel dynamic masked contrastive learning method to pre-train a foundation model on various heterogeneous public HAR datasets, ensuring strong generalizability across different modal combinations. It then adopts semi-supervised edge-cloud collaborative learning to fine-tune the pre-trained model with heterogeneous and unlabeled local data, adapting it for the target user group. Our evaluations on public and self-collected datasets demonstrate thatPFHARsignificantly outperforms SOTA baselines in both the pre-training and edge-cloud collaborative fine-tuning stages. Zhengyuan Zhang 0001, Dong Zhao 0001, Guanzhou Zhu, Chunliang Li, Yuanchun Li 0003, Huadong Ma |
IEEE Trans. Mob. Comput. | 7 |
| 2026 | Bi-Level Bandwidth Coordination for Multiple Video Inference at the EdgeabstractHigh-definition (HD) cameras for surveillance and road traffic have experienced tremendous growth, demanding intensive computation resources for real-time analytics. Recently, offloading frames from the front-end device to the back-end edge server has shown great promise. In multi-stream competitive environments, efficient bandwidth management and proper scheduling are crucial to ensure both high inference accuracy and high throughput. To achieve this goal, we propose BiSwift, a bi-level framework that scales the concurrent real-time video analytics by a novel adaptive hybrid codec integrated with multi-level pipelines, and a global bandwidth controller for multiple video streams. The lower-level front-back-end collaborative mechanism (called adaptive hybrid codec) locally optimizes the accuracy and accelerates end-to-end video analytics for a single stream. The upper-level scheduler aims to accuracy fairness among multiple streams via the global bandwidth controller. The evaluation of BiSwift shows that BiSwift is able to real-time object detection on 9 streams with an edge device only equipped with an NVIDIA RTX3070 (8G) GPU. BiSwift improves 10%~21% accuracy and presents$1.2\sim 9\times $throughput compared with the state-of-the-art video analytics pipelines. Haipeng Dai 0001, Jinghan Chen, Liang Mi, Weijun Wang 0001, Yuanchun Li 0003, Tingting Yuan 0001, Yuben Qu, Yunxin Liu 0001, Xiaoming Fu 0001, Guihai Chen |
IEEE Trans. Netw. | 5 |
| 2025 | GUI-Xplore: Empowering Generalizable GUI Agents with One ExplorationabstractGUI agents hold significant potential to enhance the experience and efficiency of human-device interaction. However, current methods face challenges in generalizing across applications (apps) and tasks, primarily due to two fundamental limitations in existing datasets. First, these datasets overlook developer-induced structural variations among apps, limiting the transferability of knowledge across diverse software environments. Second, many of them focus solely on navigation tasks, which restricts their capacity to represent comprehensive software architectures and complex user interactions. To address these challenges, we introduce GUI-Xplore, a dataset meticulously designed to enhance cross-application and cross-task generalization via an exploration-and-reasoning framework. GUI-Xplore integrates pre-recorded exploration videos providing contextual insights, alongside five hierarchically structured downstream tasks designed to comprehensively evaluate GUI agent capabilities. To fully exploit GUI-Xplore’s unique features, we propose Xplore-Agent, a GUI agent framework that combines Action-aware GUI Modeling with Graph-Guided Environment Reasoning. Further experiments indicate that Xplore-Agent achieves a 10% improvement over existing methods in unfamiliar environments, yet there remains significant potential for further enhancement towards truly generalizable GUI agents.1 Shanhui Zhao, Hao Wen 0004, Samith Va, Mengwei Xu 0001, Yuanchun Li 0003 |
CVPR | 7 |
| 2025 | Empower Vision Applications with LoRA LMMabstractLarge Multimodal Models (LMMs) have shown significant progress in various complex vision tasks with the solid linguistic and reasoning capacity inherited from large language models (LMMs). Low-rank adaptation (LoRA) offers a promising method to integrate external knowledge into LMMs, compensating for their limitations on domain-specific tasks. However, the existing LoRA model serving is excessively computationally expensive and causes extremely high latency. In this paper, we present an end-to-end solution that empowers diverse vision tasks and enriches vision applications with LoRA LMMs. Our system, VaLoRA, enables accurate and efficient vision tasks by 1) an accuracy-aware LoRA adapter generation approach that generates LoRA adapters rich in domain-specific knowledge to meet application-specific accuracy requirements, 2) an adaptive-tiling LoRA adapters batching operator that efficiently computes concurrent heterogeneous LoRA adapters, and 3) a flexible LoRA adapter orchestration mechanism that manages application requests and LoRA adapters to achieve the lowest average response latency. We prototype VaLoRA on five popular vision tasks on three LMMs. Experiment results reveal that VaLoRA improves 24-62% of the accuracy compared to the original LMMs and reduces 20-89% of the latency compared to the state-of-the-art LoRA model serving systems. Liang Mi, Weijun Wang 0001, Wenming Tu, Qingfeng He, Xinyu Fang, Yazhu Dong, Yuanchun Li 0003, Meng Li 0010, Haipeng Dai 0001, Guihai Chen, Yunxin Liu 0001 |
EuroSys | 9 |
| 2025 | LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile AppsabstractLarge language models (LLMs) have opened new opportunities for automated mobile app exploration, an important and challenging problem that used to suffer from the difficulty of generating meaningful UI interactions. However, existing LLM-based exploration approaches rely heavily on LLMs to generate actions in almost every step, leading to a huge cost of token fees and computational resources. We argue that such extensive usage of LLMs is neither necessary nor effective, since many actions during exploration do not require, or may even be biased by the abilities of LLMs. Further, based on the insight that a precise and compact knowledge plays the central role for effective exploration, we introduce LLM-Explorer, a new exploration agent designed for efficiency and affordability. LLM-Explorer uses LLMs primarily for maintaining the knowledge instead of generating actions, and knowledge is used to guide action generation in a LLM-less manner. Based on a comparison with 5 strong baselines on 20 typical apps, LLM-Explorer was able to achieve the fastest and highest coverage among all automated app explorers, with over 148x lower cost than the state-of-the-art LLM-based approach. Shanhui Zhao, Hao Wen 0004, Wenjie Du 0004, Cheng Liang 0006, Yunxin Liu 0001, Xiaozhou Ye, Ye Ouyang, Yuanchun Li 0003 |
MobiCom | 8 |
| 2025 | AutoDroid-V2: Boosting SLM-based GUI Agents via Code GenerationabstractLarge language models (LLMs) have brought exciting new advances to mobile UI agents, a long-standing research field that aims to complete arbitrary natural language tasks through mobile UI interactions. However, existing UI agents usually demand powerful large language models that are difficult to be deployed locally on end-users' devices, raising huge concerns about user privacy and centralized serving cost. Inspired by the remarkable coding abilities of recent small language models (SLMs), we propose to convert the UI task automation problem to a code generation problem, which can be effectively solved by an on-device SLM and efficiently executed with an on-device code interpreter. Unlike normal coding tasks that can be extensively pre-trained with public datasets, generating UI automation code is challenging due to the diversity, complexity, and variability of target apps. Therefore, we adopt a document-centered approach that automatically builds fine-grained API documentation for each app and generates diverse task samples based on this documentation. By guiding the agent with the synthetic documents and task samples, it learns to generate precise and efficient scripts to complete unseen tasks. Based on detailed comparisons with state-of-the-art mobile UI agents, our approach effectively improves the mobile task automation with significantly higher success rates and lower latency/token consumption. Code is open-sourced at https://github.com/MobileLLM/AutoDroid-V2. Hao Wen 0004, Shizuo Tian, Borislav Pavlov, Wenjie Du 0004, Ge Chang 0002, Shanhui Zhao, Yunxin Liu 0001, Ya-Qin Zhang, Yuanchun Li 0003 |
MobiSys | 11 |
| 2025 | Region-based Content Enhancement for Efficient Video Analytics at the Edge
Weijun Wang 0001, Liang Mi, Shaowei Cen, Haipeng Dai 0001, Yuanchun Li 0003, Xiaoming Fu 0001, Yunxin Liu 0001 |
NSDI | 5 |
| 2025 | Serving MoE Models on Resource-Constrained Edge Devices via Dynamic Expert SwappingabstractMixture of experts (MoE) is a popular technique in deep learning that improves model capacity with conditionally-activated parallel neural network modules (experts). However, serving MoE models in resource-constrained latency-critical edge scenarios is challenging due to the significantly increased model size and complexity. In this paper, we first analyze the behavior pattern of MoE models in continuous inference scenarios, which leads to three key observations about the expert activations, including temporal locality, exchangeability, and skippable computation. Based on these observations, we introduce PC-MoE, an inference framework for resource-constrained continuous MoE model serving. The core of PC-MoE is a new data structure,Parameter Committee, that intelligently maintains a subset of important experts in use to reduce resource consumption. To evaluate the effectiveness of PC-MoE, we conduct experiments using state-of-the-art MoE models on common computer vision and natural language processing tasks. The results demonstrate optimal trade-offs between resource consumption and model accuracy achieved by PC-MoE. For instance, on object detection tasks with the Swin-MoE model, our approach can reduce memory usage and latency by 42.34% and 18.63% with only 0.10% accuracy degradation. Yuanchun Li 0003, Weijun Wang 0001, Linghe Kong, Yunxin Liu 0001 |
IEEE Trans. Computers | 2 |
| 2025 | TimelyNet: Adaptive Neural Architecture for Autonomous Driving with Dynamic DeadlineabstractTo maintain driving safety, the execution of neural network-based autonomous driving pipelines must meet the dynamic deadlines in response to the changing environment and vehicle’s velocity. To this end, this article proposes a real-time neural architecture adaptation approach, called TimelyNet, which uses a supernet to replace the most compute-intensive neural network module in an existing end-to-end autonomous driving pipeline. From the supernet, TimelyNet samples subnets with varying inference latency levels to meet the dynamic deadlines during run-time driving without fine-tuning. Specifically, TimelyNet employs a one-shot prediction method that jointly uses a lookup table and an invertible neural network to periodically determine the optimal hyperparameters of a subnet to meet its execution deadline while achieving the highest possible accuracy. The lookup table stores multiple subnet architectures with different latencies, while the invertible neural network models the distribution of the optimal subnet architecture given the latency. Extensive evaluation based on hardware-in-the-loop CARLA simulations shows that TimelyNet-integrated driving pipelines achieve the best driving safety, characterized by the lowest wrong-lane driving rate and zero collisions, compared with several baselines, including the state-of-the-art driving pipelines. Duc Van Le, Yuanchun Li 0003, Yunxin Liu 0001, Rui Tan 0001 |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2025 | Efficient and Adaptive Diffusion Model Inference Through Lookup Table on Mobile DevicesabstractDiffusion models have revolutionized image synthesis applications. Many studies focus on using approximate computation such as model quantization to reduce inference costs on mobile devices. However, due to their extensive model parameters and autoregressive inference fashion, the overhead of diffusion models remains high, which is challenging for mobile devices to handle. To reduce the inference overhead of diffusion models on mobile devices, we proposeLUT-Diff, an algorithm-system co-design specifically tailored for mobile device diffusion model inference optimization.LUT-Diffoptimizes using lookup tables and can efficiently generate a series of lookup table candidates for diffusion models without end-to-end training. During inference,LUT-Diffadaptively selects the best inference strategy based on the application/user's latency budget. Additionally,LUT-Diffincludes a parallel inference engine that rapidly completes model inference through CPU-GPU co-scheduling. Extensive experiments demonstrate thatLUT-Diffcan generate images comparable to the original model, with an up to 0.012 MSE in generated images.LUT-Diffcan also achieve up to 9.1× inference acceleration and reduce the inference memory footprint by up to 70.9% compared to baseline methods. Moreover,LUT-Diffcan save at least 3281× the learning cost of lookup tables. Qipeng Wang 0001, Shiqi Jiang 0002, Yifan Yang 0004, Ruiqi Liu 0001, Yuanchun Li 0003, Ting Cao 0003, Xuanzhe Liu |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | Squeezer: Efficient Multi-DNN Inference for Edge Video Analytics via Cross-Model SchedulingabstractVideo analytics at the edge is becoming increasingly prevalent in many scenarios, such as smart campuses and intelligent factories. These applications often consist of multiple subtasks, which necessitates the optimization for multi-DNN (Deep Neural Network) inference. Due to limited consideration over cross-model scheduling, current practices cannot fully leverage available computing resources, leading to suboptimal performance. To address this, we propose Squeezer, a multiDNN serving framework that holistically schedules multiple DNN models on an edge server with a single GPU. Squeezer decouples the cross-model scheduling into a two-layered approach, which involves (1) balanced operator grouping which partitions operators of multiple DNN models into groups, significantly reducing the scheduling complexity and (2) kernel scheduler which orchestrates parallel execution within each group by considering the interplay among kernels running in parallel, thereby enabling cross-model optimizations in multi-DNN inference. Performance evaluation results demonstrate that Squeezer outperforms state-of-the-art baselines, achieving up to 1.91× improvement in system throughput. Lingxiao Ma, Ziyan Fu 0001, Yuanchun Li 0003, Ju Ren 0001, Yaoxue Zhang, Yunxin Liu 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | AdaWiFi, Collaborative WiFi Sensing for Cross-Environment AdaptationabstractDeep learning (DL) based Wi-Fi sensing has witnessed great development in recent years. Although decent results have been achieved in certain scenarios, Wi-Fi based activity recognition is still difficult to deploy in real smart homes due to the limited cross-environment adaptability, i.e. a well-trained Wi-Fi sensing neural network in one environment is hard to adapt to other environments. To address this challenge, we proposeAdaWiFi, a DL-based Wi-Fi sensing framework that allows multiple Internet-of-Things (IoT) devices to collaborate and adapt to various environments effectively. The key innovation ofAdaWiFiincludes a collective sensing model architecture that utilizes complementary information between distinct devices and avoids the biased perception of individual sensors and an accompanying model adaptation technique that can transfer the sensing model to new environments with limited data. We evaluate our system on a public dataset and a custom dataset collected from three complex sensing environments. The results demonstrate thatAdaWiFiis able to achieve significantly better sensing adaptation effectiveness (e.g. 30% higher accuracy with one-shot adaptation) as compared with state-of-the-art baselines. Naiyu Zheng, Yuanchun Li 0003, Shiqi Jiang 0002, Yuanzhe Li 0001, Rongchun Yao, Chuchu Dong, Zhimeng Yin 0001, Yunxin Liu 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Anatomizing Deep Learning Inference in Web BrowsersabstractWeb applications have increasingly adopted Deep Learning (DL) through in-browser inference , wherein DL inference performs directly within Web browsers. The actual performance of in-browser inference and its impacts on the Quality of Experience ( QoE ) remain unexplored, and urgently require new QoE measurements beyond traditional ones, e.g., mainly focusing on page load time. To bridge this gap, we make the first comprehensive performance measurement of in-browser inference to date. Our approach proposes new metrics to measure in-browser inference: responsiveness, smoothness, and inference accuracy. Our extensive analysis involves 9 representative DL models across Web browsers of 50 popular PC devices and 20 mobile devices. The results reveal that in-browser inference exhibits a substantial latency gap, averaging 16.9 times slower on CPU and 4.9 times slower on GPU compared to native inference on PC devices. The gap on mobile CPU and mobile GPU is 15.8 times and 7.8 times, respectively. Furthermore, we identify contributing factors to such latency gap, including underutilized hardware instruction sets, inherent overhead in the runtime environment, resource contention within the browser, and inefficiencies in software libraries and GPU abstractions. Additionally, in-browser inference imposes significant memory demands, at times exceeding 334.6 times the size of the DL models themselves, partly attributable to suboptimal memory management. We also observe that in-browser inference leads to a significant 67.2% increase in the time it takes for GUI components to render within Web browsers, significantly affecting the overall user QoE of Web applications reliant on this technology. Qipeng Wang 0001, Shiqi Jiang 0002, Zhenpeng Chen 0001, Yuanchun Li 0003, Aoyu Li, Yun Ma 0002, Ting Cao 0003, Xuanzhe Liu |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2024 | SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory BudgetabstractRui Kong, Yuanchun Li, Qingtian Feng, Weijun Wang, Xiaozhou Ye, Ye Ouyang, Linghe Kong, Yunxin Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yuanchun Li 0003, Qingtian Feng, Weijun Wang 0001, Xiaozhou Ye, Ye Ouyang, Linghe Kong, Yunxin Liu 0001 |
ACL (1) | 2 |
| 2024 | TESLA: Thermally Safe, Load-Aware, and Energy-Efficient Cooling Control System for Data CentersabstractThe increasing demand for artificial intelligence and cloud computing has led to skyrocketing energy consumption of data centers (DCs). This paper focuses on tackling this energy challenge through cooling control system optimization, which aims to ensure thermal safety with minimal cooling energy consumption. Current industry practice involves human operators, while many data-driven methods have also been proposed. However, human intervention often results in unnecessary energy consumption, particularly in the face of fluctuating server loads, whereas existing data-driven methods struggle to maintain thermal safety in practice. To overcome these issues, we propose TESLA, a thermally safe, load-aware, and energy-efficient cooling control system for data centers. TESLA employs a novel data-driven framework that integrates domain knowledge to predict DC temperature and cooling energy under dynamic server load. Based on these predictions, a Bayesian optimizer (BO) finds the energy-optimal settings for the cooling system at every control step. Besides cooling energy, BO’s optimization objective also includes minimizing cooling interruption that causes rapid temperature rise within the data center and leads to thermal safety violations. We deploy TESLA on a real data-center testbed and show that it achieves on average <?TeX $10.1\%$?> Math 1 cooling energy saving relative to a fixed cooling system parameter setting and no thermal safety violation relative to previous data-driven methods. Hanfei Geng, Yuanzhe Li 0001, Jichao Leng, Xianyuan Zhan, Yuanchun Li 0003, Feng Zhao 0001, Yunxin Liu 0001 |
ICPP | 7 |
| 2024 | AutoDroid: LLM-powered Task Automation in AndroidabstractMobile task automation is an attractive technique that aims to enable voice-based hands-free user interaction with smartphones. However, existing approaches suffer from poor scalability due to the limited language understanding ability and the non-trivial manual efforts required from developers or endusers. The recent advance of large language models (LLMs) in language understanding and reasoning inspires us to rethink the problem from a model-centric perspective, where task preparation, comprehension, and execution are handled by a unified language model. In this work, we introduce AutoDroid, a mobile task automation system capable of handling arbitrary tasks on any Android application without manual efforts. The key insight is to combine the commonsense knowledge of LLMs and domain-specific knowledge of apps through automated dynamic analysis. The main components include a functionality-aware UI representation method that bridges the UI with the LLM, exploration-based memory injection techniques that augment the app-specific domain knowledge of LLM, and a multi-granularity query optimization module that reduces the cost of model inference. We integrate AutoDroid with off-the-shelf LLMs including online GPT-4/GPT-3.5 and on-device Vicuna, and evaluate its performance on a new benchmark for memory-augmented Android task automation with 158 common tasks. The results demonstrated that AutoDroid is able to precisely generate actions with an accuracy of 90.9%, and complete tasks with a success rate of 71.3%, outperforming the GPT-4-powered baselines by 36.4% and 39.7%. Hao Wen 0004, Yuanchun Li 0003, Guohong Liu 0002, Shanhui Zhao, Toby Jia-Jun Li, Shiqi Jiang 0002, Yunhao Liu 0001, Yunxin Liu 0001 |
MobiCom | 2 |
| 2024 | FlexNN: Efficient and Adaptive DNN Inference on Memory-Constrained Edge DevicesabstractDue to the popularity of deep neural networks (DNNs) and considerations over network overhead, data privacy, and inference latency, there is a growing interest in deploying DNNs to edge devices in recent years. However, the limited memory becomes a major bottleneck for on-device DNN deployment, making it crucial to reduce the memory footprint of DNN. The mainstream model customization solutions require intensive deployment efforts and may lead to severe accuracy degradation, and existing deep learning (DL) frameworks don't take memory as a priority. Besides, recent works to enhance the memory management scheme cannot be directly applied because of several challenges, including the unbalanced memory footprint across layers, the inevitable overhead of memory management, and the memory budget dynamicity. To tackle these challenges, we introduce FlexNN, an efficient and adaptive memory management framework for DNN inference on memory-constrained devices. FlexNN uses a slicing-loading-computing joint planning approach, to achieve optimal memory utilization and minimal memory management overhead. We implemented FlexNN atop NCNN, and conducted comprehensive evaluations with common model architectures on various devices. The results have shown that our approach is able to adapt to different memory constraints with optimal latency-memory trade-offs. For example, FlexNN can reduce the memory consumption by 93.81% with only a 3.64% increase in latency, as compared with the original NCNN on smartphones. Yuanchun Li 0003, Yuanzhe Li 0001, Ting Cao 0003, Yunxin Liu 0001 |
MobiCom | 2 |
| 2024 | Poster: Enabling Agent-centric Interaction on Smartphones with LLM-based UI ReassemblingabstractIn this poster, we introduce a novel dynamic user interface (UI) specifically designed for mobile devices powered by large language models (LLMs) agents. The advent of LLMs has led to a surge in deploying LLM-based agents on personal and Internet of Things (IoT) devices, with the aim of facilitating various daily tasks through device manipulation. However, this integration poses a significant challenge: how to intelligently and flexibly select and present information both during and after the execution of tasks, ensuring users are well-informed about the operations and can access the desired results conveniently. To address this challenge, we propose a UI reassembling method. This method allows for analyzing and strategically combining different mobile applications and their UI components, enabling the dynamic construction and adjustment of UIs tailored to user needs. Our prototype exhibits promising performance, with the UI selection module achieving an F1 score of 0.74. This innovative approach opens up exciting possibilities of new user-device interaction paradigm, leveraging the capabilities of LLMs to enhance the user experience in handling mobile and IoT devices. Hao Wen 0004, Wenjie Du 0004, Yuanchun Li 0003, Yunxin Liu 0001 |
MobiSys | 3 |
| 2024 | Empowering In-Browser Deep Learning Inference on Edge Through Just-In-Time Kernel OptimizationabstractWeb is increasingly becoming the primary platform to deliver AI services onto edge devices, making in-browser deep learning (DL) inference more prominent. Nevertheless, the heterogeneity of edge devices, combined with the underdeveloped state of Web hardware acceleration practices, hinders current in-browser inference from achieving its full performance potential on target devices. Fucheng Jia, Shiqi Jiang 0002, Ting Cao 0003, Tianrui Xia, Yuanchun Li 0003, Qipeng Wang 0001, Ju Ren 0001, Yunxin Liu 0001, Lili Qiu, Mao Yang 0004 |
MobiSys | 7 |
| 2024 | LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Task AutomationabstractThe emergent large language/multimodal models facilitate the evolution of mobile agents, especially in mobile UI task automation. However, existing evaluation approaches, which rely on human validation or established datasets to compare agent-predicted actions with predefined action sequences, are unscalable and unfaithful. To overcome these limitations, this paper presents LlamaTouch, a testbed for on-device mobile UI task execution and faithful, scalable task evaluation. By observing that the task execution process only transfers UI states, LlamaTouch employs a novel evaluation approach that only assesses whether an agent traverses all manually annotated, essential application/system states. LlamaTouch comprises three key techniques: (1) On-device task execution that enables mobile agents to interact with realistic mobile environments for task execution. (2) Fine-grained UI component annotation that merges pixel-level screenshots and textual screen hierarchies to explicitly identify and precisely annotate essential UI components with a rich set of designed annotation primitives. (3) A multi-level application state matching algorithm that utilizes exact and fuzzy matching to accurately detect critical information in each screen, even with unpredictable UI layout/content dynamics. LlamaTouch currently incorporates four mobile agents and 496 tasks, encompassing both tasks in the widely-used datasets and our self-constructed ones to cover more diverse mobile applications. Evaluation results demonstrate LlamaTouch’s high faithfulness of evaluation in real-world mobile environments and its better scalability than human validation. LlamaTouch also enables easy task annotation and integration of new mobile agents. Code and dataset are publicly available at https://github.com/LlamaTouch/LlamaTouch. Li Zhang 0133, Shihe Wang, Xianqing Jia, Zhihan Zheng, Yunhe Yan, Longxi Gao, Yuanchun Li 0003, Mengwei Xu 0001 |
UIST | 7 |
| 2024 | Towards Energy-efficient Federated Learning via INT8-based Training on Mobile DSPsabstractAI is making the Web an even cooler place, but also introduces serious privacy risks due to the extensive user data collection. Federated learning (FL), as a privacy-preserving machine learning paradigm, enables mobile devices to collaboratively learn a shared prediction model while keeping all training data on devices. However, a key obstacle towards practical cross-device FL training is huge energy consumption, especially for lightweight mobile devices. In this work, we perform the first-of-its-kind analysis of improving FL performance through low-precision training with an energy-friendly Digital Signal Processor (DSP) on mobile devices. We first demonstrate that directly integrating the state-of-the-art INT8 (8-bit integer) training algorithm and classic FL protocols will significantly degrade the model accuracy. Moreover, we observe that there are still unavoidable frequent quantization operations on devices that cause extreme load stress on DSP-enabled INT8 training. To address the above challenges, we present Q-FedUpdate, an FL framework that efficiently preserves model accuracy with ultra-low energy consumption. It maintains a global full-precision model and allows the tiny model updates to be continuously accumulated, instead of being erased by the quantization. Furthermore, it introduces pipelining technology to parallel CPU-based quantization and DSP-enabled training, which reduces the floating-point computation overhead of frequent data quantization. Extensive experiments show that Q-FedUpdate can effectively reduce the on-device energy consumption by 21×, and accelerate the FL convergence by 6.1× with only 2% accuracy loss. Jinliang Yuan, Shangguang Wang, Daliang Xu, Yuanchun Li 0003, Mengwei Xu 0001, Xuanzhe Liu |
WWW | 5 |
| 2024 | Seamless Cross-Edge Service Migration for Real-Time Rendering ApplicationsabstractSeamless cross-edge migration for real-time rendering applications is challenging. The strong interactive nature of real-time rendering applications demands a downtime lower than 15ms to achieve an imperceptible migration. Existing methods based on virtual machine migration and container migration suffer from unpleasant downtime brought by dirty page retransmission-induced repeated memory data copy and the shared storage failure-induced extensive disk data copy. In this paper, we propose Cloud-assisted Service Migration (CSM) which leverages cloud-edge collaboration to achieve seamless service migration for real-time rendering applications. CSM improves service migration user experience in three folds: First, it introduces a dual rendering mechanism to bypass the peer-to-peer data copy and compresses the freezing stage. Second, a user equipment-centric session switch mechanism is proposed to save time by well coordinating application session switches and 5G user plane session switches. Third, a smooth switching mechanism is leveraged to prevent unpleasant frame flickers during session switching. We implement CSM in edge-rendering multiplayer games and deploy it on a 5G test bed with a full-stack user plane protocol stack. The evaluation results show that CSM can reduce downtime to < 14ms and the service migration process is user imperceptible. Yuanzhe Li 0001, Shangguang Wang, Yuanchun Li 0003, Ao Zhou 0001, Mengwei Xu 0001, Xiao Ma 0009, Yunxin Liu 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | HiMoDepth: Efficient Training-Free High-Resolution On-Device Depth PerceptionabstractHigh-resolution depth estimation, with a minimum resolution of$1280\times 960$, is essential for achieving more immersive experiences in on-device 3D vision applications. However, implementing high-resolution solutions on resource-limited mobile devices presents significant challenges, such as the need for additional expensive depth sensors, computation-intensive machine learning models requiring large-scale datasets, or the need for device motion while the target object remains stationary. In this study, we propose HiMoDepth, an efficient training-free high-resolution depth estimation system that utilizes widely-available on-device dual cameras. HiMoDepth consists of two modules: 1) homogenizing the on-device heterogeneous cameras by iteratively cropping the Field-of-Views to make the focal length of the cameras equal and filtering out the out-of-sync frames based on time stamps, and 2) designing a hierarchical mobile GPU-friendly stereo matching method that effectively reduces the latency of stereo matching with high-resolution depth maps by using efficient data layout, reducing the number of memory accesses, and searching the corresponding pixel over a coarse-to-fine hierarchy. We implement HiMoDepth on multiple commodity mobile devices and conduct comprehensive evaluations. Experimental results show that HiMoDepth significantly outperforms the baselines in both accuracy and running speed on mobile devices that support high-resolution depth maps. Ju Ren 0001, Bangwen He, Youngki Lee 0001, Ting Cao 0003, Yuanchun Li 0003, Yaoxue Zhang, Yunxin Liu 0001 |
IEEE Trans. Mob. Comput. | 8 |
| 2024 | CamoNet: On-Device Neural Network Adaptation With Zero Interaction and Unlabeled Data for Diverse Edge EnvironmentsabstractDeploying deep learning models to edge devices for low-latency and privacy-preserving applications has become a trend. To adapt to heterogeneous devices and data, it is significant to generate customized models. However, existing model adaptation approaches require edge devices to make interactions (collecting hardware information or local data) with the cloud, which raises privacy concerns, increases communication costs, and burdens the cloud. By contrast, we proposeCamoNet, a universal on-device model adaptation framework with zero interaction between devices and the cloud. InCamoNet, a lightweight on-device neural architecture search module is utilized to quickly generate a customized model for subsequent on-device training, followed by an on-device contrastive transfer learning module to effectively leverage unlabeled data for fine-tuning the customized model. Extensive experimental results show thatCamoNetcan effectively run on various edge devices. Compared with the SOTA model adaptation approaches,CamoNetachieves significant accuracy improvement by 25.2% on average for image classification, 10.1% on average for object detection, and reduces the training memory by 4.8-11.4×. We will open-source our models and tools for edge AI developers. Zhengyuan Zhang 0001, Dong Zhao 0001, Renhao Liu, Kuo Tian, Yuxing Yao, Yuanchun Li 0003, Huadong Ma |
IEEE Trans. Mob. Comput. | 6 |
| 2023 | FedSlice: Protecting Federated Learning Models from Malicious Participants with Model SlicingabstractCrowdsourcing Federated learning (CFL) is a new crowdsourcing development paradigm for the Deep Neural Network (DNN) models, also called “software 2.0”. In practice, the privacy of CFL can be compromised by many attacks, such as free-rider attacks, adversarial attacks, gradient leakage attacks, and inference attacks. Conventional defensive techniques have low efficiency because they deploy heavy encryption techniques or rely on Trusted Execution Environments (TEEs). To improve the efficiency of protecting CFL from these attacks, this paper proposes FedSlice to prevent malicious participants from getting the whole server-side model while keeping the performance goal of CFL. FedSlice breaks the server-side model into several slices and delivers one slice to each participant. Thus, a malicious participant can only get a subset of the server-side model, preventing them from effectively conducting effective attacks. We evaluate FedSlice against these attacks, and results show that FedSlice provides effective defense: the server-side model leakage is reduced from 100% to 43.45%, the success rate of adversarial attacks is reduced from 100% to 11.66%, the average accuracy of membership inference is reduced from 71.91% to 51.58%, and the data leakage from shared gradients is reduced to the level of random guesses. Besides, FedSlice only introduces less than 2% accuracy loss and about 14% computation overhead. To the best of our knowledge, this is the first paper to discuss defense methods against these attacks to the CFL framework. Ziqi Zhang 0017, Yuanchun Li 0003, Yifeng Cai, Ding Li 0001, Yao Guo 0001, Xiangqun Chen |
ICSE | 2 |
| 2023 | Privacy as a Resource in Differentially Private Federated LearningabstractDifferential privacy (DP) enables model training with a guaranteed bound on privacy leakage, therefore is widely adopted in federated learning (FL) to protect the model update. However, each DP-enhanced FL job accumulates privacy leakage, which necessitates a unified platform to enforce a global privacy budget for each dataset owned by users. In this work, we present a novel DP-enhanced FL platform that treats privacy as a resource and schedules multiple FL jobs across sensitive data. It first introduces a novel notion of device-time blocks for distributed data streams. Such data abstraction enables fine-grained privacy consumption composition across multiple FL jobs. Regarding the non-replenishable nature of the privacy resource (that differs it from traditional hardware resources like CPU and memory), it further employs an allocation-then-recycle scheduling algorithm. Its key idea is to first allocate an estimated upper-bound privacy budget for each arrived FL job, and then progressively recycle the unused budget as training goes on to serve further FL jobs. Extensive experiments show that our platform is able to deliver up to 2.1× as many completed jobs while reducing the violation rate by up to 55.2% under limited privacy budget constraint. Jinliang Yuan, Shangguang Wang, Shihe Wang, Yuanchun Li 0003, Xiao Ma 0009, Ao Zhou 0001, Mengwei Xu 0001 |
INFOCOM | 4 |
| 2023 | Evaluating and Enhancing the Robustness of Federated Learning System against Realistic Data CorruptionabstractFederated learning (FL) has emerged as a prominent paradigm enabling collaborative model training without transmitting local data, thereby safeguarding data privacy. However, the practical implementation of FL systems on these devices faces a significant challenge: the heterogeneous corruption of data on individual clients, leading to unanticipated accuracy degradation during real-world deployment. In this work, we first introduce a realistic data corruption simulation framework to test the robustness of FL systems. In this framework, an in-depth analysis of potential data corruption patterns occurring on devices is conducted, followed by the construction of individual datasets with varying corruption types and degrees. Such data corruption results in the robustness degradation of conventional FL protocol (FedAVG) significantly higher than centralized learning (CL). Atop this key observation, we propose an adaptive FL protocol that emulates the CL training process. The protocol leverages imbalanced client data sampling to mitigate the negative impact of data corruption. Furthermore, a hybrid aggregation strategy is designed to accelerate model convergence and reduce additional communication overhead. Extensive experiments validate the effectiveness of our approach in enhancing the robustness of FL systems against client data corruption, which achieves up to 12% higher converge accuracy than FedAVG-based systems with acceptable overhead. Chen Yang 0043, Yuanchun Li 0003, Jinliang Yuan, Qibo Sun, Shangguang Wang, Mengwei Xu 0001 |
ISSRE | 2 |
| 2023 | ReSPlay: Improving Cross-Platform Record-and-Replay with GUI Sequence MatchingabstractRecord-and-replay is an important testing technique to ensure the quality of mobile applications (apps in short). State-of-the-art record-and-replay approaches are typically based on widget matching, which has shown limited effectiveness, especially on devices with different platforms and resolutions, due to the difficulty in matching widgets with subtle visual differences. Our key observation is that, even if two widgets look similar, the resulting screenshot sequences can still be very different during execution. Thus, instead of matching GUI widgets directly, we are able to find the correct replay actions by comparing the resulting GUI screenshot sequences, which can be better distinguished across different platforms, thus potentially improving the record-and-replay efficiency through GUI exploration and comparison.This paper proposes a general record-and-replay framework called ReSPlay, which leverages a more robust visual feature, GUI sequences, to guide replaying more accurately. ReSPlay pre-trains a deep reinforcement learning model, SDP-Net, offline from random app traces. Specifically, SDP-Net is trained to search a particular path from GUI transition graphs to learn an optimal policy to locate the target operation positions by maximizing the possibilities to reach the target GUI sequence. Finally, the trained SDP-Net is used to search for potential event traces with high rewards and replicate them on the target device for replay. We evaluate our proposed framework on multiple real devices. Experimental results show that the overall average replay accuracy of ReSPlay on devices across different OSes, GUI styles, and resolutions is 28.12% higher than the state-of-the-art baselines. Linna Wu, Yuanchun Li 0003, Ziqi Zhang 0017, Hanwen Lei, Ding Li 0001, Yao Guo 0001, Xiangqun Chen |
ISSRE | 3 |
| 2023 | PatchBackdoor: Backdoor Attack against Deep Neural Networks without Model ModificationabstractBackdoor attack is a major threat to deep learning systems in safety-critical scenarios, which aims to trigger misbehavior of neural network models under attacker-controlled conditions. However, most backdoor attacks have to modify the neural network models through training with poisoned data and/or direct model editing, which leads to a common but false belief that backdoor attack can be easily avoided by properly protecting the model. In this paper, we show that backdoor attacks can be achieved without any model modification. Instead of injecting backdoor logic into the training data or the model, we propose to place a carefully-designed patch (namely backdoor patch) in front of the camera, which is fed into the model together with the input images. The patch can be trained to behave normally at most of the time, while producing wrong prediction when the input image contains an attacker-controlled trigger object. Our main techniques include an effective training method to generate the backdoor patch and a digital-physical transformation modeling method to enhance the feasibility of the patch in real deployments. Extensive experiments show that PatchBackdoor can be applied to common deep learning models (VGG, MobileNet, ResNet) with an attack success rate of 93% to 99% on classification tasks. Moreover, we implement PatchBackdoor in real-world scenarios and show that the attack is still threatening. Yizhen Yuan, Shenghao Xie 0002, Yuanchun Li 0003, Yunxin Liu 0001 |
ACM Multimedia | 4 |
| 2023 | AdaptiveNet: Post-deployment Neural Architecture Adaptation for Diverse Edge EnvironmentsabstractDeep learning models are increasingly deployed to edge devices for real-time applications. To ensure stable service quality across diverse edge environments, it is highly desirable to generate tailored model architectures for different conditions. However, conventional pre-deployment model generation approaches are not satisfactory due to the difficulty of handling the diversity of edge environments and the demand for edge information. In this paper, we propose to adapt the model architecture after deployment in the target environment, where the model quality can be precisely measured and private edge data can be retained. To achieve efficient and effective edge model generation, we introduce a pretraining-assisted on-cloud model elastification method and an edge-friendly on-device architecture search method. Model elastification generates a high-quality search space of model architectures with the guidance of a developer-specified oracle model. Each subnet in the space is a valid model with different environment affinity, and each device efficiently finds and maintains the most suitable subnet based on a series of edge-tailored optimizations. Extensive experiments on various edge devices demonstrate that our approach is able to achieve significantly better accuracy-latency tradeoffs (e.g. 46.74% higher on average accuracy with a 60% latency budget) than strong baselines with minimal overhead (13 GPU hours in the cloud and 2 minutes on the edge server). Hao Wen 0004, Yuanchun Li 0003, Zunshuai Zhang, Shiqi Jiang 0002, Xiaozhou Ye, Ye Ouyang, Yunxin Liu 0001 |
MobiCom | 2 |
| 2023 | ConvReLU++: Reference-based Lossless Acceleration of Conv-ReLU Operations on Mobile CPUabstractMany activation values of Convolutional Neural Networks (CNNs) are zeros due to ReLU (Rectified Linear Unit), one of the most common activation functions used in modern neural networks. Since ReLU outputs are zero for all negative inputs, existing CNN acceleration approaches estimate zero outputs to skip redundant computation, which has to sacrifice accuracy for efficiency and leads to dilemma trade-offs and cockamamie configuration. In this paper, we introduce a lossless acceleration method ConvReLU++ for CNN inference on mobile devices, which accurately detects and skips zero-outputs for speedup without failures. The key to early negative detection is adopting reference-based upper-bounds calculation. This ensures that as soon as the intermediate results become negative, the final results are guaranteed to be negative. Upon detection, the remaining computation can be skipped and the following ReLU output can be simply set to zero. We rigorously prove the losslessness property of ConvReLU++, analyze the theoretical FLOPs reduction, and show the compatibility of our method with vector-level parallelism on mobile platforms. We implement ConvReLU++ in popular mobile inference frameworks and evaluate it on common deep vision tasks. The results demonstrate that ConvReLU++ can achieve 2.90% to 8.91% latency reduction over the original inference framework on edge devices without sacrificing accuracy. Our code can be found at https://github.com/monster119120/conv_relu_plus_plus. Yuanchun Li 0003, Yizhen Yuan, Linghe Kong |
MobiSys | 2 |
| 2023 | Understanding the Impact of Quantum Noise on Quantum ProgramsabstractQuantum computing is expected to introduce the next era of computing speed and power, and its software - quantum program is gaining increasing research interest in the software engineering community. A significant characteristic of quantum computing is the existence of noise. Unlike classical computers where the output of a program is usually deterministic, the execution of a quantum program may be affected by quantum noise. Such a difference may cause difficulties or misunderstandings for developers shifting from classical programming to quantum programming. To understand the impact of quantum noise on quantum programs and its implications for software developers, we conduct a series of studies with real-world quantum programs and quantum computing environments. Specifically, we first measure and analyze the noise in a real quantum computer by testing it with a basic quantum program. We find that a non-neglectable amount of quantum noise generally exists in real quantum computers. Then we investigate the robustness of quantum programs against different quantum noises by testing 18 real-world quantum programs and 50,000 randomly generated quantum circuits in simulated and real environments. We observe that quantum noise can significantly influence the correctness of quantum programs, and different quantum circuit structures show diverse sensitivity patterns under the same noise. Based on the observations, we build a machine learning model to predict the fidelity of a quantum program under certain quantum noise. The model achieves a small average fidelity prediction error, meaning the impact of noise can be precisely estimated statistically. Zhonghao Pan, Yang Feng 0003, Yunxin Liu 0001, Yuanchun Li 0003 |
SANER | 5 |
| 2023 | Environment-aware Testing for DNN-based Smart-home WiFi Sensing SystemsabstractWiFi-based human activity recognition is a promising sensing application in smart homes due to the low cost, wide availability, and privacy preservation of WiFi devices. However, pushing WiFi sensing technology to industry-scale deployment is difficult due to its poor robustness against environment differences. How to systematically test such sensing system is crucial to improve its practicality, and is also challenging because the sensing performance is significantly influenced by the underlying physical environments. In this paper, we introduce the problem of testing environment-dependent sensing systems, including how to measure test coverage and how to effectively generate data to improve the coverage. We describe our initial attempts on examining test sufficiency with environment-neuron joint coverage and improving the coverage through targeted environment variations and signal transformations. Our experiments have demonstrated the higher effectiveness of using environment-neuron coverage to represent test sufficiency, as compared with using the conventional neuron coverage. Meanwhile, the coverage-guided sensing data generation can lead to higher accuracy of the sensing system under changing environments. Naiyu Zheng, Chuchu Dong, Yuanzhe Li 0001, Yunxin Liu 0001, Yuanchun Li 0003 |
SANER | 7 |
| 2023 | PatchCensor: Patch Robustness Certification for Transformers via Exhaustive TestingabstractIn the past few years, Transformer has been widely adopted in many domains and applications because of its impressive performance. Vision Transformer (ViT), a successful and well-known variant, attracts considerable attention from both industry and academia thanks to its record-breaking performance in various vision tasks. However, ViT is also highly nonlinear like other classical neural networks and could be easily fooled by both natural and adversarial perturbations. This limitation could pose a threat to the deployment of ViT in the real industrial environment, especially in safety-critical scenarios. How to improve the robustness of ViT is thus an urgent issue that needs to be addressed. Among all kinds of robustness, patch robustness is defined as giving a reliable output when a random patch in the input domain is perturbed. The perturbation could be natural corruption, such as part of the camera lens being blurred. It could also be a distribution shift, such as an object that does not exist in the training data suddenly appearing in the camera. And in the worst case, there could be a malicious adversarial patch attack that aims to fool the prediction of a machine learning model by arbitrarily modifying pixels within a restricted region of an input image. This kind of attack is also called physical attack, as it is believed to be more real than digital attack. Although there has been some work on patch robustness improvement of Convolutional Neural Network, related studies on its counterpart ViT are still at an early stage as ViT is usually much more complex with far more parameters. It is harder to assess and improve its robustness, not to mention to provide a provable guarantee. In this work, we propose PatchCensor, aiming to certify the patch robustness of ViT by applying exhaustive testing. We try to provide a provable guarantee by considering the worst patch attack scenarios. Unlike empirical defenses against adversarial patches that may be adaptively breached, certified robust approaches can provide a certified accuracy against arbitrary attacks under certain conditions. However, existing robustness certifications are mostly based on robust training, which often requires substantial training efforts and the sacrifice of model performance on normal samples. To bridge the gap, PatchCensor seeks to improve the robustness of the whole system by detecting abnormal inputs instead of training a robust model and asking it to give reliable results for every input, which may inevitably compromise accuracy. Specifically, each input is tested by voting over multiple inferences with different mutated attention masks, where at least one inference is guaranteed to exclude the abnormal patch. This can be seen as complete-coverage testing, which could provide a statistical guarantee on inference at the test time. Our comprehensive evaluation demonstrates that PatchCensor is able to achieve high certified accuracy (e.g., 67.1% on ImageNet for 2%-pixel adversarial patches), significantly outperforming state-of-the-art techniques while achieving similar clean accuracy (81.8% on ImageNet). The clean accuracy is the same as vanilla ViT models. Meanwhile, our technique also supports flexible configurations to handle different adversarial patch sizes by simply changing the masking strategy. Yuheng Huang 0004, Lei Ma 0003, Yuanchun Li 0003 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2022 | Representational Continuity for Unsupervised Continual Learning
Divyam Madaan, Jaehong Yoon, Yuanchun Li 0003, Yunxin Liu 0001, Sung Ju Hwang |
ICLR | 3 |
| 2022 | ReMoS: Reducing Defect Inheritance in Transfer Learning via Relevant Model SlicingabstractTransfer learning is a popular software reuse technique in the deep learning community that enables developers to build custom models (students) based on sophisticated pretrained models (teachers). However, like vulnerability inheritance in traditional software reuse, some defects in the teacher model may also be inherited by students, such as well-known adversarial vulnerabilities and backdoors. Reducing such defects is challenging since the student is unaware of how the teacher is trained and/or attacked. In this paper, we propose ReMoS, a relevant model slicing technique to reduce defect inheritance during transfer learning while retaining useful knowledge from the teacher model. Specifically, ReMoS computes a model slice (a subset of model weights) that is relevant to the student task based on the neuron coverage information obtained by profiling the teacher model on the student task. Only the relevant slice is used to finetune the student model, while the irrelevant weights are retrained from scratch to minimize the risk of inheriting defects. Our experiments on seven DNN defects, four DNN models, and eight datasets demonstrate that ReMoS can reduce inherited defects effectively (by 63% to 86% for CV tasks and by 40% to 61% for NLP tasks) and efficiently with minimal sacrifice of accuracy (3% on average). Ziqi Zhang 0017, Yuanchun Li 0003, Jindong Wang 0001, Ding Li 0001, Yao Guo 0001, Xiangqun Chen, Yunxin Liu 0001 |
ICSE | 2 |
| 2022 | MobiDepth: real-time depth estimation using on-device dual camerasabstractReal-time depth estimation is critical for the increasingly popular augmented reality and virtual reality applications on mobile devices. Yet existing solutions are insufficient as they require expensive depth sensors or motion of the device, or have a high latency. We propose MobiDepth, a real-time depth estimation system using the widely-available on-device dual cameras. While binocular depth estimation is a mature technique, it is challenging to realize the technique on commodity mobile devices due to the different focal lengths and unsynchronized frame flows of the on-device dual cameras and the heavy stereo-matching algorithm. Ju Ren 0001, Bangwen He, Ting Cao 0003, Yuanchun Li 0003, Yaoxue Zhang, Yunxin Liu 0001 |
MobiCom | 7 |
| 2022 | FedBalancer: data and pace control for efficient federated learning on heterogeneous clientsabstractFederated Learning (FL) trains a machine learning model on distributed clients without exposing individual data. Unlike centralized training that is usually based on carefully-organized data, FL deals with on-device data that are often unfiltered and imbalanced. As a result, conventional FL training protocol that treats all data equally leads to a waste of local computational resources and slows down the global learning process. To this end, we propose FedBalancer, a systematic FL framework that actively selects clients' training samples. Our sample selection strategy prioritizes more "informative" data while respecting privacy and computational capabilities of clients. To better utilize the sample selection to speed up global training, we further introduce an adaptive deadline control scheme that predicts the optimal deadline for each round with varying client training data. Compared with existing FL algorithms with deadline configuration methods, our evaluation on five datasets from three different domains shows that FedBalancer improves the time-to-accuracy performance by 1.20~4.48× while improving the model accuracy by 1.1~5.0%. We also show that FedBalancer is readily applicable to other FL approaches by demonstrating that FedBalancer improves the convergence speed and accuracy when operating jointly with three different FL algorithms. Jaemin Shin 0005, Yuanchun Li 0003, Yunxin Liu 0001, Sung-Ju Lee 0001 |
MobiSys | 2 |
| 2021 | DeepPayload: Black-box Backdoor Attack on Deep Learning Models through Neural Payload InjectionabstractDeep learning models are increasingly used in mobile applications as critical components. Unlike the program bytecode whose vulnerabilities and threats have been widely-discussed, whether and how the deep learning models deployed in the applications can be compromised are not well-understood since Neural Networks are usually viewed as a black box. In this paper, we introduce a highly practical backdoor attack achieved with a set of reverse-engineering techniques over compiled deep learning models. The core of the attack is a neural conditional branch constructed with a trigger detector and several operators and injected into the victim model as a malicious payload. The attack is effective as the conditional logic can be flexibly customized by the attacker, and scalable as it does not require any prior knowledge from the original model. We evaluated the attack effectiveness using 5 state-of-the-art deep learning models and real-world samples collected from 30 users. The results demonstrated that the injected backdoor can be triggered with a success rate of 93.5%, while only brought less than 2ms latency overhead and no more than 1.4% accuracy decrease. We further conducted an empirical study on real-world mobile deep learning apps collected from Google Play. We found 54 apps that were vulnerable to our attack, including popular and security-critical ones. The results call for the awareness of deep learning application developers and auditors to enhance the protection of deployed models. Yuanchun Li 0003, Jiayi Hua, Haoyu Wang 0001, Chunyang Chen 0001, Yunxin Liu 0001 |
ICSE | 1 |
| 2021 | Dependency-aware Form UnderstandingabstractForm understanding is an important task in many fields such as software testing, AI assistants, and improving accessibility. One key goal of understanding a complex set of forms is to identify the dependencies between form elements. However, it remains a challenge to capture the dependencies accurately due to the diversity of UI design patterns and the variety in development experiences. In this paper, we propose a deep-learning-based approach called DependEX, which integrates convolutional neural networks (CNNs) and transformers to help understand dependencies within forms. DependEX extracts semantic features from UI images using CNN-based models, captures contextual patterns using a multilayer transformer encoder module, and models dependencies between form elements using two embedding layers. We evaluate DependEX with a large-scale dataset from mobile Web applications. Experimental results show that our proposed model achieves over 92% accuracy in identifying dependencies between UI elements, which significantly outperforms other competitive methods, especially for heuristic-based methods. We also conduct case studies on automatic form filling and test case generation from natural language (NL) instructions, which demonstrates the applicability of our approach. Yuanchun Li 0003, Weixiang Yan, Yao Guo 0001, Xiangqun Chen |
ISSRE | 2 |
| 2021 | ModelDiff: testing-based DNN similarity comparison for model reuse detectionabstractThe knowledge of a deep learning model may be transferred to a student model, leading to intellectual property infringement or vulnerability propagation. Detecting such knowledge reuse is nontrivial because the suspect models may not be white-box accessible and/or may serve different tasks. In this paper, we propose ModelDiff, a testing-based approach to deep learning model similarity comparison. Instead of directly comparing the weights, activations, or outputs of two models, we compare their behavioral patterns on the same set of test inputs. Specifically, the behavioral pattern of a model is represented as a decision distance vector (DDV), in which each element is the distance between the model's reactions to a pair of inputs. The knowledge similarity between two models is measured with the cosine similarity between their DDVs. To evaluate ModelDiff, we created a benchmark that contains 144 pairs of models that cover most popular model reuse methods, including transfer learning, model compression, and model stealing. Our method achieved 91.7% correctness on the benchmark, which demonstrates the effectiveness of using ModelDiff for model reuse detection. A study on mobile deep learning apps has shown the feasibility of ModelDiff on real-world models. Yuanchun Li 0003, Ziqi Zhang 0017, Yunxin Liu 0001 |
ISSTA | 1 |
| 2021 | Flexible high-resolution object detection on edge devices with tunable latencyabstractObject detection is a fundamental building block of video analytics applications. While Neural Networks (NNs)-based object detection models have shown excellent accuracy on benchmark datasets, they are not well positioned for high-resolution images inference on resource-constrained edge devices. Common approaches, including down-sampling inputs and scaling up neural networks, fall short of adapting to video content changes and various latency requirements. This paper presents Remix, a flexible framework for high-resolution object detection on edge devices. Remix takes as input a latency budget, and come up with an image partition and model execution plan which runs off-the-shelf neural networks on non-uniformly partitioned image blocks. As a result, it maximizes the overall detection accuracy by allocating various amount of compute power onto different areas of an image. We evaluate Remix on public dataset as well as real-world videos collected by ourselves. Experimental results show that Remix can either improve the detection accuracy by 18%-120% for a given latency budget, or achieve up to 8.1× inference speedup with accuracy on par with the state-of-the-art NNs. Shiqi Jiang 0002, Yuanchun Li 0003, Yuanchao Shu, Yunxin Liu 0001 |
MobiCom | 3 |
| 2021 | Glider: A Reinforcement Learning Approach to Extract UI Scripts from WebsitesabstractWeb automation scripts (tasklets) are used by personal AI assistants to carry out human tasks such as reserving a car or buying movie tickets. Generating tasklets today is a tedious job which requires much manual effort. We propose Glider, an automated and scalable approach to generate tasklets from a natural language task query and a website URL. A major advantage of Glider is that it does not require any pre-training. Glider models tasklet extraction as a state space search, where agents can explore a website's UI and get rewarded when making progress towards task completion. The reward is computed based on the agent's navigating pattern and the similarity between its trajectory and the task query. A hierarchical reinforcement learning policy is used to efficiently find the action sequences that maximize the reward. To evaluate Glider, we used it to extract tasklets for tasks in various categories (shopping, real-estate, flights, etc.); in 79% of cases a correct tasklet was generated. Yuanchun Li 0003, Oriana Riva |
SIGIR | 1 |
| 2021 | TaintStream: fine-grained taint tracking for big data platforms through dynamic code translationabstractBig data has become valuable property for enterprises and enabled various intelligent applications. Today, it is common to host data in big data platforms (e.g., Spark), where developers can submit scripts to process the original and intermediate data tables. Meanwhile, it is highly desirable to manage the data to comply with various privacy requirements. To enable flexible and automated privacy policy enforcement, we propose TaintStream, a fine-grained taint tracking framework for Spark-like big data platforms. TaintStream works by automatically injecting taint tracking logic into the data processing scripts, and the injected scripts are dynamically translated to maintain a taint tag for each cell during execution. The dynamic translation rules are carefully designed to guarantee non-interference in the original data operation. By defining different semantics of taint tags, TaintStream can enable various data management applications such as access control, data retention, and user data erasure. Our experiments on a self-crafted benchmarksuite show that TaintStream is able to achieve accurate cell-level taint tracking with a precision of 93.0% and less than 15% overhead. We also demonstrate the usefulness of TaintStream through several real-world use cases of privacy policy enforcement. Chengxu Yang, Yuanchun Li 0003, Mengwei Xu 0001, Zhenpeng Chen 0001, Yunxin Liu 0001, Gang Huang 0001, Xuanzhe Liu |
ESEC/SIGSOFT FSE | 2 |
| 2021 | Beyond the virus: a first look at coronavirus-themed Android malware
Liu Wang 0002, Haoyu Wang 0001, Pengcheng Xia 0001, Yuanchun Li 0003, Lei Wu 0012, Yajin Zhou, Xiapu Luo, Yulei Sui, Yao Guo 0001, Guoai Xu |
Empir. Softw. Eng. | 5 |
| 2020 | Dynamic slicing for deep neural networksabstractProgram slicing has been widely applied in a variety of software engineering tasks. However, existing program slicing techniques only deal with traditional programs that are constructed with instructions and variables, rather than neural networks that are composed of neurons and synapses. In this paper, we introduce NNSlicer, the first approach for slicing deep neural networks based on data-flow analysis. Our method understands the reaction of each neuron to an input based on the difference between its behavior activated by the input and the average behavior over the whole dataset. Then we quantify the neuron contributions to the slicing criterion by recursively backtracking from the output neurons, and calculate the slice as the neurons and the synapses with larger contributions. We demonstrate the usefulness and effectiveness of NNSlicer with three applications, including adversarial input detection, model pruning, and selective model protection. In all applications, NNSlicer significantly outperforms other baselines that do not rely on data flow analysis. Ziqi Zhang 0017, Yuanchun Li 0003, Yao Guo 0001, Xiangqun Chen, Yunxin Liu 0001 |
ESEC/SIGSOFT FSE | 2 |
| 2019 | Humanoid: A Deep Learning-Based Approach to Automated Black-box Android App TestingabstractAutomated input generators must constantly choose which UI element to interact with and how to interact with it, in order to achieve high coverage with a limited time budget. Currently, most black-box input generators adopt pseudo-random or brute-force searching strategies, which may take very long to find the correct combination of inputs that can drive the app into new and important states. We propose Humanoid, an automated black-box Android app testing tool based on deep learning. The key technique behind Humanoid is a deep neural network model that can learn how human users choose actions based on an app's GUI from human interaction traces. The learned model can then be used to guide test input generation to achieve higher coverage. Experiments on both open-source apps and market apps demonstrate that Humanoid is able to reach higher coverage, and faster as well, than the state-of-the-art test input generators. Humanoid is open-sourced at https://github.com/yzygitzh/Humanoid and a demo video can be found at https://youtu.be/PDRxDrkyORs. Yuanchun Li 0003, Yao Guo 0001, Xiangqun Chen |
ASE | 1 |
| 2018 | Automated Extraction of Personal Knowledge from Smartphone Push NotificationsabstractPersonalized services are in need of a rich and powerful personal knowledge base, i.e. a knowledge base containing information about the user. This paper proposes an approach to extracting personal knowledge from smartphone push notifications, which are used by mobile systems and apps to inform users of a rich range of information. Our solution is based on the insight that most notifications are formatted using templates, while knowledge entities can be usually found within the parameters to the templates. As defining all the notification templates and their semantic rules are impractical due to the huge number of notification templates used by potentially millions of apps, we propose an automated approach for personal knowledge extraction from push notifications. We first discover notification templates through pattern mining, then use machine learning to understand the template semantics. Based on the templates and their semantics, we are able to translate notification text into knowledge facts automatically. Users' privacy is preserved as we only need to upload the templates to the server for model training, which do not contain any personal information. According to experiments with about 120 million push notifications from 100,000 smartphone users, our system is able to extract personal knowledge accurately and efficiently. Yuanchun Li 0003, Yao Guo 0001, Xiangqun Chen, Yuvraj Agarwal, Jason I. Hong |
IEEE BigData | 1 |
| 2018 | What's inside my app?: understanding feature redundancy in mobile appsabstractAs the number of mobile apps increases rapidly, many users may install dozens of, or even hundreds of, apps on a single smartphone. However, many apps on the same phone may contain similar or even the same feature, resulting in feature redundancy. For example, multiple apps may check weather forecast for the user periodically. Feature redundancy may cause many undesirable side-effects such as consuming extra CPU resources and network traffic. This paper proposes a method to identify common features within an app, and evaluated it on over four thousand popular apps. Experiments on a list of apps installed on actual smartphones show that the extent of feature redundancy is very high. We found that more than 85% of user smartphones contain redundant features, while in extreme cases, some smartphones may contain dozens of apps with the same feature. In addition, our user surveys found out that about half of the redundant features are undesirable from the end users' perspective, which indicates that feature redundancy has become an important issue that needs to be investigated further. Yao Guo 0001, Yuanchun Li 0003, Xiangqun Chen |
ICPC | 2 |
| 2017 | Understanding the Purpose of Permission Use in Mobile AppsabstractMobile apps frequently request access to sensitive data, such as location and contacts. Understanding the purpose of why sensitive data is accessed could help improve privacy as well as enable new kinds of access control. In this article, we propose a text mining based method to infer the purpose of sensitive data access by Android apps. The key idea we propose is to extract multiple features from app code and then use those features to train a machine learning classifier for purpose inference. We present the design, implementation, and evaluation of two complementary approaches to infer the purpose of permission use, first using purely static analysis, and then using primarily dynamic analysis. We also discuss the pros and cons of both approaches and the trade-offs involved. Haoyu Wang 0001, Yuanchun Li 0003, Yao Guo 0001, Yuvraj Agarwal, Jason I. Hong |
ACM Trans. Inf. Syst. | 2 |
| 2016 | PERUIM: understanding mobile application privacy with permission-UI mappingabstractCurrent mobile operating systems such as Android employ the permission-based access control mechanism, but it is difficult for users to understand how and why the permissions are used within a particular application. This paper introduces permission-UI mapping as an easy-to-understand representation to illustrate how permissions are used by different UI components within a given application. Connecting UI components to permissions helps users to understand the purpose of permission requests and also makes it possible to illustrate permission requests in a fine-grained manner. We propose PERUIM to extract the permission-UI mapping from an application based on both dynamic and static analysis, and represent the analysis results with a graphical representation. Experiments on popular mobile applications demonstrate the accuracy and applicability of the proposed approach. Yuanchun Li 0003, Yao Guo 0001, Xiangqun Chen |
UbiComp | 1 |
| 2015 | Fixing sensor-related energy bugs through automated sensing policy instrumentationabstractAs mobile applications (apps) become more and more complex, many apps contain various energy bugs, which may cause energy wastes that might reduce the battery life to as short as several hours. Among them, sensor-related bugs such as sensor data underutilization is one of the most common energy bugs. Instead of trying to detect these energy bugs, this paper proposes a method to fix sensor data underutilization automatically through instrumentation of existing apps. App-specific energy-aware sensing policies can be written to the apps via an automated instrumentation process, which can also be customized by users if needed. The proposed technique is easy to apply as it does not need to modify the operating system or the apps. At the same time, it also works for existing legacy apps, which makes it practical and feasible for a wide-range of mobile apps. Experimental results on popular Android apps show that we are able to achieve significant energy savings through automated instrumentation and rebuilding the targeted apps. Yuanchun Li 0003, Yao Guo 0001, Junjun Kong, Xiangqun Chen |
ISLPED | 1 |